Fast modular multiplication of large integers
The method of using a binary adder and single modular correction term for modular addition of large integers addresses the inefficiencies in existing technologies, resulting in a 5 times faster processing speed for modular operations in elliptic curve cryptography.
Patent Information
- Application Number
- JP2024573780
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-07
- Filing Date
- 2023-07-06
- Publication Date
- 2025-07-03
AI Technical Summary
Existing technologies face challenges in performing fast and efficient modular addition and multiplication operations on large integer values, particularly in the context of elliptic curve cryptography, due to the computational intensity and slow processing times associated with traditional methods.
A method and system that utilize a binary adder and a single modular correction term to perform modular addition of multiple large integer values, allowing for parallel binary additions followed by a single modular reduction, thereby reducing the number of required operations and increasing processing speed.
This approach significantly enhances the execution speed of modular operations, achieving up to 5 times faster performance compared to conventional techniques, with reduced hardware overhead and the ability to handle large bit sizes efficiently.
Smart Images

Figure 2025520506000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to modular addition of operands, and more specifically, to fast modular addition of a plurality of operands of large integer values.
Background Art
[0002] In modern information technology (IT), improvements in hardware development and software requirements reinforce each other. In some cases, there are faster advancements in hardware that enable the development of new software concepts, and in some cases, there are faster software advancements that drive new hardware performance requirements. One of these advancements on the software side is blockchain technology, which enables a completely decentralized ledger system without a central point of control. Such technology can be used for the workload of electronic business transactions. One of the basic and fundamental technologies used is cryptography, for example, based on elliptic curve cryptography (ECC) which is based on operations such as point addition and point doubling; operations such as "sign" for adding new transactions or "verify" for proving transactions to other clients always involve such ECC operations (for example, known HyperLedger software spends about 60% on ECC at a ratio of 6:1 for verification vs. signature operations).
[0003] However, ECC is used not only in blockchain projects but also outside of that specific area. Specifically, ECC support has already been integrated into hardware devices.
[0004] At a higher level, ECC requires operations such as signature, verification, and scalar multiplication. These operations are based on functions such as pointAdd, where each step is a mod-P operation, and where P is a large prime number. pointAdd involves simpler operations such as addition, subtraction, multiplication, and half-subtraction using very wide integer values that range from, for example, 256 bits to 521 bits, and in some cases even larger. Typical curves used in this context include NIST P256, P384, P521 (NIST: US National Institute of Standards and Technology), and Edwards curves 448, 255-19. Thus, a bit size of 521 (in words: five-hundred-twenty-one) bits is not uncommon, and in some cases, operands of up to 8k bits are used today, which could even be extended in the future.
[0005] On the other hand, there are hardware processors that have the option of dealing with double-word integers and whose registers for addition and multiplication operations are the size of a normal memory word (e.g., 32 or 64 bits). However, the large number of bits mentioned above for the operands usually exceeds the performance of a standard processor for performing addition and multiplication operations or modular addition and multiplication operations as a single instruction in a fully pipelined manner. In particular, modular multiplication for those wide numbers is computationally intensive and difficult to verify. Therefore, modular multiplication of two wide integers A and B is usually performed as a binary multiplication to generate the binary product M = A * B, followed by a modulo (or remainder) operation to generate the result R = M mod P. Note that M can be expressed as the sum of multiples of R and P, i.e., M = R + k * P, where k is an integer. In a naive implementation, modular reduction of M can be performed by subtracting multiple Ps from M until the difference is between 0 and P. Each of these subtractions can be formally verified even for integers having hundreds of bits. However, this scheme is too slow. Therefore, more sophisticated algorithms and hardware implementations are used to reduce the binary product in just a few steps. Usually, such algorithms apply coarse-grained and fine-grained corrections to the binary product.
[0006] One type of operation required in such an approach depends on modular addition operations for more terms such as within a sequence or on modular subtraction, which can be considered as modular addition with negative terms. Further, these operations are also required for multiplication of bit-wise wide operands as mentioned above.
[0007] Meyer (US2021 / 0243006 A1) (hereinafter, Meyer) describes that "an integrated circuit for modular multiplication of two integers for an encryption method represents an integer to be multiplied in Montgomery representation using specified Montgomery representation parameters and a specified modulus, and has a processor that repeatedly calculates, from the least significant word to the most significant word, the result of the modular multiplication of the integer to be multiplied in Montgomery representation." (Meyer, abstract).
[0008] Sinardi et al. (WO2018 / 019788 A1) (hereinafter, Sinardi) also disclose "implementing modular multiplication" and "Montgomery multiplication by leveraging the properties of Montgomery multiplication and Fermat's little theorem to avoid the costly pre-computation of the Montgomery constant R 2 in an elliptic curve digital signature algorithm (ECDSA) using Montgomery multiplication... verification and signature processes." (Sinardi, abstract).
[0009] However, there remains a need for addition and multiplication operations on large integers, reduced cycle counts, and even faster hardware implementations for other elliptic curves. SUMMARY OF THE INVENTION
[0010] According to one aspect of the present invention, a computer-implemented method may be provided. The method may include receiving a plurality of first operand values, where the first operand values are integer values, adding the plurality of first operand values using binary addition to yield a total value S, and determining a single combined modular correction term D for the binary sum of all operand values based on the leading bit of the total value S. Further, the method may include performing a modular addition of S and D to yield a modular sum of the plurality of first operand values.
[0011] According to another aspect of the present invention, a system may be provided. The system may include a receiving unit for a plurality of first operand values, where the operand values are integer values, and a binary adder unit adapted to add the plurality of first operand values using binary addition to yield a total value S. Further, the high-speed modular addition system may include a determination unit adapted to determine a single combined modular correction term D for the binary sum of all operand values based on the total value S, and a modular adder unit adapted to modularly add S and D to yield the modular sum of the plurality of the first operand values.
[0012] According to another aspect of the present invention, a computer-implemented method may be provided. The computer-implemented method may include receiving, using a receiving unit, a plurality of first operand values, where the first operand values are integer values, adding, using a binary adder unit, the plurality of first operand values using binary addition to yield a total value S, and determining, using a determination unit, a single combined modular correction term D for the binary sum of all operand values based on the total value S. Further, the computer-implemented method may include performing, using a modular adder unit, a modular addition of S and D to yield the modular sum of the plurality of the first operand values.
[0013] The proposed computer-implemented method for modular addition of a plurality of first large integer values may offer a number of advantages, technical effects, contributions, and / or improvements.
[0014] A very fast modular adder may be implemented using a plurality of known components in a newly invented form. The known components include a binary adder unit and a multiplexer selection unit. The novel units proposed herein as various alternatives similarly utilize the proposed methods. In particular, a modular multiplier may be based on the novel concept of a modular adder in an advantageous manner.
[0015] One concept of the proposed approach may use a binary adder instead of a modular adder. Thus, when adding n operands, instead of performing n binary additions followed by a modular reduction for each, the solution proposes performing n binary additions followed by a single modular reduction, which can significantly increase the speed of binary execution. This can be possible because multiple binary additions can be performed in parallel if the hardware used permits it. However, multiple modular additions can be completely avoided.
[0016] The proposed concept can significantly increase the execution speed of modular operations, such as addition and multiplication, when compared to existing techniques. In one implementation, a multiplexer, a binary adder, a selector unit, and some latches can be used to realize a 4-cycle process for a modular adder within the hardware. Further, when using a "carry save adder" unit, some input latches, a binary adder, a selector, and an output latch, only 3 machine cycles may be sufficient to perform a multi-input modular addition operation for wide-size integer values.
[0017] The proposed concept can also overcome the drawbacks of modular multiplication operations in the Montgomery domain, which is also known as a fast approach for modular multiplication operations. However, it is known that the switching to and from the Montgomery domain is slow. In particular, for Mersenne primes, the reduction operation is much faster. Currently, the cross-over point compared to a pure binary design is considered to be beyond 512 bits. However, for reasons of consistency, the proposed technique may be for integer values with fewer bits starting, for example, from 256 (or even lower) bits.
[0018] Furthermore, due to the execution of reduction in binary arithmetic, multiple Tis (note: Tis can be a collection of terms of an intermediate product organized as large integers. Multiple Ts for different is in Tis must be added together; compare with the following), without adding additional terms, can also be scaled by two powers. On the other hand, when performing reduction using modular addition, 2*Ti needs to be either treated as for the term Ti+Ti or a modular multiplication operation needs to be performed. Therefore, the binary scheme using scaled Tis may require fewer terms in reduction. For example, for P256, only 9 terms may be required instead of 11 terms. For P384, only 10 terms may be required instead of 11 terms. Thereby, the latency of P256 can be: (i) one machine cycle to construct the Tis and send them to a reduction tree, which can be shared by the multiplier; (ii) one machine cycle to reduce 9 terms to 2 terms in the form of a sum / carry vector; (iii) one machine cycle for the determination of D based on the sum / carry vector approach; and (iv) three machine cycles for modular addition, which can be only 6 machine cycles.
[0019] Compared with a conventional modular arithmetic design that typically has 30 cycles, the proposed solution can be 5 times faster. Furthermore, the hardware overhead can be small because the reduction tree of the multiplier can be reused. If the reduction of 11 terms is done with a 2-cycle adder instead, binary addition is interleaved and can be done in 10 cycles (instead of 2). Using 4 to 6 cycles for steps (iii) and (iv), it becomes 14 or 16 cycles, and when compared with the original 30 machine cycles, there is still a speedup of 1.75 to 2.1 times.
[0020] Therefore, it can be concluded that embodiments of the concepts proposed herein can be significantly faster when implemented in hardware compared to conventional techniques. Further, multiple different implementation designs are possible, (i) with respect to the modular addition concept, and (ii) with respect to the modular multiplication concept. This can be particularly applicable to integers with large bit sizes as mentioned.
[0021] Here, other things that may be advantageous in some application areas can be added. Usually, when there is only one operation to be performed, no special number representation is required. Typically, binary operations are exact and do not require correction terms as often discussed herein. However, modular addition may require that the result be less than the modulus value. If the binary addition violates that constraint, thus, correction terms as described may need to be added. Theoretically, it is possible to stop after a binary adder, but in that case, the data value width will increase after each operation, and ECC can easily have hundreds to thousands of such operations. Therefore, it may be desirable to keep the data width under control. One way is to bring it back to the range of [0, P), and another is to confine it to a range where, speaking inaccurately, a few additional bits, for example 1 to 3 bits, are required to be sufficiently accurate.
[0022] The proposed concept can function seamlessly with respect to negative Ti numbers: in this case, instead of a binary addition operation following the two's complement, a binary addition operation follows the one's complement. The number of negative numbers added to the sequence of input arguments can be regarded as an algorithmic constant; thus, the discrepancies introduced by using the one's complement can be corrected in a single step.
[0023] In the following, additional embodiments of the inventive concepts applicable to methods and systems will be described.
[0024] According to a preferred embodiment of the method, the sum value S is represented in redundant number format. When a plurality of large integer value operands, properly formatted in redundant number format, are added, it is avoided to propagate carry across the words of the sum S. This can further help to improve the arithmetic efficiency. In a related embodiment, the redundant number format can be a reduced radix format. This can make the calculation even more efficient.
[0025] According to another preferred embodiment, the method may also comprise the step of determining a correction value -k*P, where k is a value representing the leading bit of the sum value S. This can support the efficient concept that the operation R = M mod P can be converted to M = R + k*P, where k is an integer and P is a prime number.
[0026] According to an embodiment of the method, the correction value can be determined using a look-up table. Since a look-up table usually requires fewer CPU cycles than the actual calculation, the determination speed can be further increased. In a given case, the look-up table can give the modulus of the number represented by the high-order bit D.
[0027] According to one embodiment, the method may also comprise the step of performing a binary multiplication of a second integer value A and B, e.g., R = A*B, which results in a binary product M, and the binary product M is represented by a plurality of adjacent words of a predetermined number of bits. The integer values A and B can also be in the range from 256 to 521 bits here, but are not limited thereto.
[0028] This embodiment or related embodiments may also include the step of determining a plurality of coarse-grained modular correction terms Ti of the binary product M, resulting in a plurality of adjacent words of a predetermined bit size for each Ti. The predetermined bit size may be the same as the starting value of the operation, but may usually be of a smaller size. This can even be as small as the word size of the CPU or a memory word. Next, this can generate the result of a fast modular multiplication operation, which may be the motivation for this embodiment. It may also be mentioned that there are several C++ libraries that support the R = A * B operation based on Solinas' theory. Therefore, existing concepts can be advantageously used or combined for the inventive concept.
[0029] According to another embodiment of the method, the second integer value is represented in redundant number format. And in a further embodiment, the redundant number format can be a reduced-radix format. Here, the same advantages as those for the first large integer value may apply.
[0030] According to one embodiment, the method may also include the step of converting the received second integer value into a redundant number format in response to receiving the second integer value in a non-redundant format. This format does not have to be the same as that mentioned above and may be another redundant format or representation.
[0031] According to one embodiment of the method, the plurality of first operand values may be derived from a Solinas reduction operation or a Barrett reduction operation. These reduction techniques have been found to be particularly useful in the context of the novel approach embodiments.
[0032] According to one embodiment of the method, operands A and B can each be integer values having a number of bits between 255 and 521, or up to 2^13 at most. This may also include the prime number 521 used for Edwards curve implementation. The number of bits per integer may be even higher. Smaller numbers of bits per operand would be technically possible, but from a performance improvement perspective, it may be technically meaningless.
[0033] According to another embodiment of the method, the received integer value is an operand value for elliptic curve operations. In elliptic curve cryptography (ECC), point operations, i.e., addition of two points on the curve or scalar multiplication, are ultimately reduced to modular addition and modular multiplication of large integers.
[0034] Subsequently, additional embodiments of a high-speed modular addition system are described:
[0035] According to one embodiment of the high-speed modular addition system, the determination unit may be adapted to determine a single combined modular correction term D for the binary sum of all operand values based on the leading bit of the sum value S. Thus, instead of performing an expensive modular addition for each received large integer value (except that n - 1 additions are required to add n values), a cheaper binary addition and one final modular operation are performed to keep the sum generated by the binary addition within the range of the modulus.
[0036] According to one embodiment of the high-speed modular addition system, the binary adder unit may have a reduction tree generator module adapted to receive a plurality of integer values from a receiving unit, for example in the form of a carry sum adder, and an internal binary adder unit that may be adapted to receive the output value of the reduction tree generator module. This option will be discussed in more detail in the context of Figure 6.
[0037] According to a further embodiment of the high-speed modular addition system, the modular adder may have a binary adder adapted to add the total values S and D using binary addition, and a select logic adapted to determine the modular sum of a plurality of first operand values. This option will be discussed in more detail in the context of FIG. 7.
[0038] According to yet another interesting embodiment of the high-speed modular addition system, the binary adder unit may be a reduction tree generator module adapted to receive a plurality of large integer values from a receiving unit, and the reduction tree generator module may be adapted to generate an output as a sum / carry vector pair (compare with FIG. 7, 704), where the binary sum includes the same bit width as the received large integer values, and where the carry includes a predetermined number of carry bits. Clearly, this embodiment will use a redundant number format or representation. It may also be referred to as a more complete figure in FIG. 7.
[0039] According to another advanced embodiment of the high-speed modular addition system, the modular adder may have a binary adder for binary adding a sum / carry vector pair and a value D, where the value D is determined based on the sum / carry vector pair (compare with reference numeral 406 in FIG. 7), and a select logic (compare with reference numeral 708 in FIG. 7) adapted to determine the modular sum of a plurality of first operand values. This embodiment may also be shown as Variant 1.
[0040] According to a further useful embodiment, the high-speed modular addition system may further comprise a receiving unit for two second operand values (A, B), where the operand values are integer values, a binary multiplier adapted to receive the two second operand values from the receiving unit (compare with reference numeral 502 in FIG. 5), and a correction unit adapted to determine a correction term of the binary multiplier and to provide a plurality of first large integer values (T1, ..., Tn) (compare with reference numeral 504 in FIG. 5). This may implement a modular multiplier based on the initially proposed modular adder. Details can be understood in connection with FIG. 5.
[0041] Furthermore, an embodiment may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For the purposes of this specification, a computer-usable or computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
Brief Description of the Drawings
[0042] It should be noted that embodiments of the present invention are described with reference to different subjects. In particular, some embodiments are described with reference to method-type claims, while other embodiments are described with reference to apparatus-type claims. However, those skilled in the art will presume from the above and the following description that, unless otherwise stated, any combination of features belonging to one type of subject, in addition to any combination of features between different subjects, in particular, between the features of method-type claims and the features of apparatus-type claims, is also considered to be disclosed herein.
[0043] The aspects defined above and further aspects of the present invention will be apparent from the examples of embodiments to be described hereinafter and will be described with reference to the examples of embodiments, but the present invention is not limited thereto.
[0044] Preferred embodiments of the present invention will be described by way of example only, with reference to the following drawings:
[0045]
Figure 1
[0046]
Figure 2
[0047]
Figure 3
[0048]
Figure 4
[0049]
Figure 5
[0050]
Figure 6
[0051]
Figure 7
[0052]
Figure 8
[0053]
Figure 9
[0054]
Figure 10
[0055] In the context of this description, the following technical idiomatic expressions, terms, and / or expressions may be used.
[0056] The term "modular addition" may denote a mathematical operation in modular arithmetic and is expressed as result = (a + b) mod c. In some cases, the "%" symbol is used to denote the remainder operator symbolically. For example, assuming 0 ≦ A, B < P and P is a prime number, modular addition R = (A + B) mod P is easy to solve because the binary sum is 0 ≦ A + B < 2*P. A hardware implementation with a two's complement adder will result in R1 = A + B, R2 = R1 - P, and the remainder of R2 if R2 ≧ 0, R1 if R2 < 0.
[0057] The term "adding binary" may indicate the normal addition of at least two binary values, in contrast to modular addition. Throughout this specification, a clear distinction should be made between two types of addition operations: binary addition versus modular addition. For example, modular multiplication R = (A * B) mod P, which is 0 ≦ A + B < P * P, is relatively more difficult to solve than the modular addition described above. Modular correction is more complex than conditional subtraction by 1 * P. For most prime curves, modular correction is computationally intensive, computationally intensive, and mathematically demanding. This is especially true when the operands A, B have values with a very large number of bits, for example, more than 255 bits.
[0058] The term "large integer value" may indicate integer values having a bit size from 256 to 521 bits, up to a maximum of 8k bits, and even more bits.
[0059] The term "operand value" may indicate input values for mathematical operations such as modular addition, binary addition, binary multiplication, and modular multiplication.
[0060] The term "single combined modular correction term" may indicate a single numerical value used as an operand for an addition operation, where a sequence of binary addition and binary multiplication constructs a second operand.
[0061] The term "redundant number form" or "redundant number representation" may refer to a number system that represents a single binary digit using more bits than necessary, such that most numbers can have multiple representations. Redundant number representations differ from ordinary binary number systems, which include two's complement where each bit has a unique value depending on its position. Many of the characteristics of redundant binary representation (BR) are different from those of regular binary representation systems. Most importantly, redundant number representations allow addition without using typical carries. Compared to non-redundant representations, redundant number representations may slow down bitwise logic operations, but mathematical operations can be faster when a larger bit width is used. Usually, each digit has its own sign, which is not necessarily the same as the sign of the number being represented. When the digits have signs, the redundant number representation is also a signed-digit representation.
[0062] The term "reduced-radix form" or "reduced-radix representation" may refer to a known number format in computer science that enables faster algorithmic operations on binary numbers. Two numbers X and Y are considered each other's reduced-radix complements if X + Y = 10bn-1, where n is the number of digits in X and Y, and where b is the radix of X and Y.
[0063] The term "binary multiplication" may refer to the simple multiplication of binary numbers. This should not be confused with modular multiplication. Care should be taken to clearly distinguish between the two different operations. The binary multiplication operation can simply be expressed as r = a × b, where a and b are binary values represented in binary format or representation. In contrast, modular multiplication is expressed as (a × b) mod c. As a reminder, modular arithmetic is a system of arithmetic for integers where numbers "wrap around" when they reach a particular value, referred to as the modulus c in the previous sentence.
[0064] The term "coarse-grained modular correction term" may refer to a value or values used to correct a first approximation of a mathematical operation, such as a modular addition operation. In a second correction stage, a fine-grained modular correction term may be used to further correct the first correction stage.
[0065] The term "Solinas reduction operation" is used here to denote, for example, a special operation that is used to generate a coarse-grained modular correction term for modular multiplication of very long integers, as a building block for keys in fully homomorphic encryption and elliptic curve cryptography. Basically, such an operation reduces the complexity of the associated operations and, in turn, makes hardware implementations more efficient. In mathematics, a Solinas prime or generalized Mersenne prime is a prime number of the form f(2 m ), where f(x) is a low-degree polynomial with small integer coefficients. These primes can enable fast modular reduction algorithms. This class of numbers includes a few other categories of primes, such as Mersenne primes of the form 2 k - 1, or Crandall primes or pseudo-Mersenne primes of the form 2 k - c for small odd values of c.
[0066] The term "Barrett reduction operation" may denote another operation in this context. The naive way to compute c = a mod n is to use a fast approximation division by n, followed by the application of a correction term. Barrett reduction is an algorithm designed to optimize this operation by replacing the division with a multiplication, assuming that n is a constant and a < n 2 .
[0067] The term "elliptic curve operation" (ECC) can denote a public-key cryptography technique based on the algebraic structure of elliptic curves over finite fields. ECC enables smaller keys to provide equivalent security compared to non-EC cryptography (e.g., based on plain finite fields). Typically, elliptic curves are applicable to key sharing, digital signatures, pseudorandom number generators, blockchains, and other tasks. Indirectly, elliptic curves can be used for encryption by combining key sharing with symmetric encryption schemes. In the context of this specification, elliptic curves are mainly used as problems that are mathematically difficult to solve in the absence of some secret information, which has many applications in cryptography.
[0068] The term "reduction tree generator module" can denote a unit optimized for processing large input data sets. Assume that there is no order among the processing elements within the data set (i.e., it is associative or commutative). The data set is partitioned into smaller chunks, and parallel threads can be used to process the chunks. Next, a reduction tree can be used to summarize the results from each chunk into a final result.
[0069] The term "sum / carry vector pair" can denote a number representation discussed in the context of FIG. 9.
[0070] In the following, a detailed description of the figures will be provided. All instructions in the figures are schematic. First, a block diagram of an embodiment of a computer-implemented method of the present invention for modular addition of a plurality of operands of a first large integer value will be described. Thereafter, further embodiments, and embodiments of a high-speed modular addition system for high-speed modular addition of a first large integer value will be described.
[0071] FIG. 1 shows a block diagram of one embodiment of a computer-implemented method 100 for modular addition of a plurality of first large integer values. The operands can be positive or negative values. Thus, an explicit subtraction method is not required.
[0072] The method comprises a step (102) of receiving a plurality of first operand values, later denoted as Ti, where the operand values are very large integer values, e.g., > 255 bit-width.
[0073] The method also comprises a step (104) of adding the plurality of operand values using binary addition to yield a sum value S, a step of determining a single combined modular correction term D for the binary sum of all the operand values based on the leading bit of the sum value S, and a step (106) of determining a single combined modular correction term D for the binary sum of all the operand values based on the leading bit of the sum value S.
[0074] In step 108, the method comprises a step of performing a modular addition of S and D to yield a modular sum of the plurality of first operand values.
[0075] FIG. 2 shows a flowchart of one embodiment 200 for modular multiplication using the concept of a fast modular adder as a core unit. The elements of the steps of the flowchart for performing modular multiplication comprise: for example, performing a binary multiplication of second large integer values A, B of 256 to 521 bit-width or even much larger bit numbers, i.e., M = A * B, to yield a binary product M (202). This can be represented by a plurality of adjacent words of a predetermined number of bits, e.g., 128-bit words, or other bit sizes, and is in fact the case.
[0076] Next, a step (204) of determining a plurality of coarse-grained modular correction terms Ti of the binary product and obtaining a plurality of adjacent words of a predetermined bit size for each Ti. These may have the same number of bits as the adjacent words mentioned at the end of the previous paragraph. However, the number of bits of these adjacent words is typically lower. The determination of the coarse-grained modular correction terms Ti is typically based on Solinas' theory, i.e., Solinas prime numbers, and is also shown as prime numbers of the form of generalized Mersenne primes, i.e., f(2 m ), where f(x) is a low-degree polynomial with relatively small integer coefficients. These prime numbers enable fast modular reduction algorithms and are often used in cryptographic applications.
[0077] As a follow-up step, the technique uses a plurality of terms Ti as a plurality of received operand values (206) (compare with the description of FIG. 1). Thereby, the result of the fast modular multiplication operation is generated.
[0078] It should also be mentioned that which bits are used exactly depends on its implementation. For example, if P is of q-bit width and Sum is of q + k-bit width, the k leading bits may be referenced. When using a redundant number system, usually slightly more bits, e.g., k + 2 bits, are referenced.
[0079] FIG. 3 shows a block diagram 300 of an embodiment of modular multiplication using correction terms. The modular addition operation R = (A + B) % P, where "%" is the remainder operator and P is a prime number, can be executed relatively easily since the binary sum is 0 ≦ A + B < 2*P. In hardware, this can be realized by a two's complement adder. R1 = A + B R2 = R1 - P If R2 ≧ 0, then R = R2, or if R2 < 0, then R = R1
[0080] On the other hand, the modular multiplication operation R = (A * B) % P is a more difficult problem because 0 ≦ A * B < P * P. Therefore, modular correction is not just a simple conditional subtraction of 1 * P. Thus, the most important prime curves for modular correction are computationally intensive and require mathematical attention. This is especially true when the operands A and B have values with a very large number of bits, for example, greater than 255.
[0081] Regarding generalized Mersenne primes, the modular multiplication operation R = (A * B) % P can be determined by A * B = bin(Cn...CO), where each Ci is a 32-bit word. Next, a plurality of correction terms Ti are formed from those of Ci. The modular product is obtained by modularly adding the terms Ti such that [Number] where [Number] is an add-mod-P operation.
[0082] A naive hardware implementation for this is relatively easy to test as a repetition of modular addition. However, high-speed hardware implementations apply coarse-grained and fine-grained corrections. This is shown in Figure 2 for the modular multiplication operation.
[0083] Two operands A, B302 first undergo a binary product A * B304 operation, resulting in a product M. Next, coarse-grained correction terms T0, T1, T2,...Tk306 are determined and added in parallel by a binary adder 308, resulting in an intermediate result I. The fine-grained correction term 310 is, of course, determined by the value of the intermediate result I and its numerical representation. To apply the fine-grained correction term 310, a mod P adder 312 is used to determine the final result R.
[0084] The structure of FIG. 3 also applies to Barrett reduction, which can perform modular reduction with respect to any arbitrary prime number. Let k be the bit width, then N = floor(2^(2k) / P). The Barrett reduction for the product M = A * B then proceeds in the following steps: 1. Q = (M>>(k - 1)) * N 2. R1 = (Q>>(k - 1)) * P 3. R2 = M[k:0] - R1[k:0] (lower k + 1 bits)
[0085] The Barrett reduction then conditionally corrects R2 by adding 2^(k + 1) if R2 is negative, or subtracting P or 2P to obtain R = (A * B) % P. Here, steps 1 and 2 determine the coarse correction (206), and step 3 applies this to the product. The subsequent correction of R2 is fine-grained correction such as (310) and (312).
[0086] That is, the core operation that needs to be repeated is the modular adder; this also applies to the case of modular multiplication. Therefore, a good implementation of the modular multiplier arithmetic unit depends on the hardware implementation of the modular adder circuit as the basic arithmetic unit. This is shown in the following figure.
[0087] FIG. 4 shows a modular adder 400 for the concept of a fast modular multiplier for a plurality of terms Ti (402) such as those generated when executing the sequence of steps described in FIG. 3. The input terms 402 T1, T2,..., Tn are fed into an n-way binary adder 404, resulting in a sum S. The result D of a determination unit 406 adapted to determine a single combined correction term D for k additions / subtractions is fed, together with the sum S, into a modular adder 408 that accepts two variables as inputs, namely S and D. Next, the output of the 2-way modular adder 408 is the result R 410 that is the modular sum of a plurality of operands for the modular addition operation.
[0088] As already explained, FIG. 5 shows the details of the implementation of a modular multiplier unit 500 for modular multiplication of integer terms A and B, each having a large number of bits. Operands A and B are fed to a binary multiplier unit 502 and then require correction terms that are determined using a determiner 504. The concepts used are further explained in FIG. 2. The output correction terms 402 T1, T2, ..., Tn (compare with FIG. 4) are then used as inputs for a modular n-way adder 400 (compare with FIG. 4). Thus, the result R506 is a modular multiplication operation for A and B.
[0089] FIG. 6 shows another implementation option for a modular adder 600 that details the n-way binary adder 404 of FIG. 4 in additional detail. Inside the n-way binary adder 404, a reduction tree unit 602 is operative and its result is input to a binary adder 604. Next, the output S of the n-way binary adder 404 is fed to a 2-way binary adder 606.
[0090] At this point, the determiner 406 for the total S and the combined modular correction term D functions as explained in the context of FIG. 4.
[0091] The 2-way modular adder 408 (compare also with FIG. 4) operates internally with a 2-way binary adder 606 and selection logic 608 that may use digital components such as multiplexers and / or look-up tables. The binary adder 606 receives as inputs the total S of the binary 2-way adder 604 and the output D of the determiner 406 for the combined modular correction term D. As a final result R410, a modular total for the operands T1, T2, ..., Tn is generated.
[0092] Figure 7 shows a block diagram of an alternative implementation 700 of a modular adder. As inputs, again a plurality of operands T1, T2, ..., Tn402 are used, which are fed into an n-way binary adder 404, which may have internally a reduction tree unit 702 that outputs the result of the binary addition operation as a carry / sum vector pair 704. This suggests that a redundant format, e.g., S0, S1, is used as the data format. One example is the reduced radix number format that is frequently used in cryptographic libraries such as OpenSSL X255-19 using the prime number 521. Additionally, a full radix format, and delay reduction in [0.88P) may also be used. The modular adder 408 may comprise a binary adder 706 and respective selection logic 708 in such an embodiment to yield a modular sum R410.
[0093] Figure 9 gives an indication 900 of how a redundant format having a carry / sum vector pair 704 may function. In scientific form, the reduced radix representation is defined as follows:
[0094] Assuming that P is an n-bit number and W is the word size of the processor; such that 0 < P < W
Number
Number
Number
Number
[0095] FIG. 8 shows a block diagram of an alternative implementation of a modular adder 800 that uses reduced number representation. Input 402, binary adder 404, and reduction tree unit 702 are compatible with the corresponding elements in FIG. 7. This also applies to the output 806 of the binary adder 404, that is, the carry / sum vector pair.
[0096] However, in such a case, the feedback loop through the selector (or decision maker) for the determination of the combined correction term D is wired differently. First, the sum vector S812 is directly fed back as an additional input to the modular adder 408 together with the combined correction term D814 to generate the modular sum R410. The vector S is also fed to the selector for the combined modular correction 808 that gives rise to the term D814, which is also fed to the modular adder 408. This can be implemented again using the binary adder 810 and the select logic 812. Various other implementation forms are possible.
[0097] Embodiments of the present invention can be implemented with substantially any type of computer, regardless of whether the platform is suitable for storing and / or executing program code. FIG. 10 shows, as an example, a computing system 1000 suitable for executing program code related to the proposed method.
[0098] Computing system 1000 is merely one example of a suitable computer system, and is not intended to present any limitation as to the scope of use or functionality of the embodiments of the present invention described herein, whether or not computing system 1000 is capable of implementing and / or performing any of the functions described above. There are a vast number of other general-purpose or special-purpose computing system environments or configurations in which components of computing system 1000 operate. Examples of well-known computing systems, environments and / or configurations that may be suitable for use with computing system / server 1000 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices and the like. Computing system / server 1000 can be described in the general context of computer system-executable instructions, such as program modules, being executed by computing system 1000. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computing system / server 1000 can be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media including memory storage devices.
[0099] As shown in the figure, computer system / server 1000 is shown in the form of a general-purpose computing device. The components of computer system / server 1000 may include, but are not limited to, one or more processors or processing units 1002, system memory 1004, and a bus 1006 that couples various system components including system memory 1004 to processor 1002. Bus 1006 represents one or more of any of a plurality of types of bus structures including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus that uses any of a variety of bus architectures. By way of example and without limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus. Computer system / server 1000 typically includes a variety of computer system readable media. Such media can be any available media that is accessible by computer system / server 1000, and it includes both volatile and nonvolatile media as well as removable and non-removable media.
[0100] System memory 1004 may include a computer system readable medium in the form of volatile memory such as random access memory (RAM) 1008 and / or cache memory 1010. The computer system / server 1000 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 1012 can be provided for reading from and writing to a non-removable non-volatile magnetic medium (not shown and typically referred to as a “hard drive”) or to a non-removable non-volatile magnetic medium. Although not shown, a magnetic disk drive for reading from and writing to a removable non-volatile magnetic disk (e.g., a “floppy disk”), and a removable non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media, and an optical disk drive for reading from and writing to the removable non-volatile optical disk can be provided. In such instances, each can be connected to bus 1006 by one or more data media interfaces. As further depicted and described below, memory 1004 can include at least one program product having a set (e.g., at least one) of program modules configured to execute the functions of embodiments of the present invention.
[0101] A program / utility having a set (at least one) of program modules 1016 can be stored in memory 1004, by way of example and not limitation, as an operating system, one or more application programs, other program modules, and program data. Each of these operating systems, one or more application programs, other program modules, and program data, or some combination thereof, can include an implementation of a networking environment. Program modules 1016 generally execute the functions and / or methodologies of embodiments of the present invention as described herein.
[0102] In addition, the computer system / server 1000 can communicate with one or more external devices 1018, such as a keyboard, a pointing device, a display 1020, etc.; one or more devices that enable a user to interact with the computer system / server 1000; and / or any device that enables the computer system / server 1000 to communicate with one or more other computing devices (for example, it can also communicate with a network card, a modem, etc.). Such communication can occur via an input / output (I / O) interface 1014. Further, the computer system / server 1000 can communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (for example, the Internet), via a network adapter 1022. As depicted, the network adapter 1022 can communicate with other components of the computer system / server 1000 via a bus 1006. Although not shown, it should be understood that other hardware and / or software components can be used with the computer system / server 1000. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0103] Furthermore, a high-speed modular addition system for high-speed modular addition of a first large integer value, or any of embodiments 500, 600, 700, or 800 (compare with FIGS. 4, 5, 6, 7, and 8), can be attached to the bus system 1006. Also, the other adder and multiplier units discussed above can be integrated into a normal computer design. Further, the implementation can also include integrating a completed modular adder and / or modular multiplier into the hardware of a processor chip, or in the form of a coprocessor using, for example, an FPGA (field programmable gate array), or in a direct CMOS or bipolar form.
[0104] The descriptions of the various embodiments of the present invention are presented for purposes of illustration and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application to technologies found in the marketplace, or the technical improvements, or to enable other skilled artisans to understand the embodiments disclosed herein.
[0105] The inventive concept can be summarized by the following items.
[0106] 1. A computer-implemented method comprising: receiving a plurality of first operand values, wherein the first operand values are integer values; adding the plurality of first operand values using binary addition to yield a total value S; determining a single combined modular correction term D for the binary sum of all operand values based on the leading bit of the total value S; and performing a modular addition of S and D to yield a modular sum of the plurality of the first operand values.
[0107] 2. The method according to item 1, wherein the total value S is represented in a redundant number format.
[0108] 3. The method according to item 2, wherein the redundant number format is a reduced radix format.
[0109] 4. The method further comprising the step of determining a correction value -k*P, where k is a value representing the leading bit of the total value S. The method according to any one of the preceding items.
[0110] 5. The method according to item 4, wherein the correction value is determined using a look-up table.
[0111] 6. The method according to any one of the preceding items, further comprising: performing a binary multiplication of a second integer value A and B, the binary multiplication resulting in a binary product M, the binary product M being represented by a plurality of adjacent words of a predetermined number of bits; determining a plurality of coarse-grained modular correction terms Ti for the binary product M, resulting in a plurality of adjacent words of a predetermined bit size for each Ti; and using the plurality of terms Ti as the plurality of received first operand values to generate a result of a fast modular multiplication operation.
[0112] 7. The method according to item 6, wherein the second integer value is represented in a redundant number format.
[0113] 8. The method according to item 7, wherein the redundant number format is a reduced radix format.
[0114] 9. The method according to item 6, further comprising the step of converting the received second integer value to a redundant number format in response to receiving the second integer value in a non-redundant format.
[0115] 10. The method according to any one of the preceding items, wherein the plurality of first operand values are derived from a selection from the group consisting of a Solinas reduction operation and a Barrett reduction operation.
[0116] 11. The method according to any one of the preceding items, wherein A and B are each integer values having a number of bits between 255 and 521.
[0117] 12. The method according to any one of the preceding items, wherein the received integer value is an operand value for elliptic curve operations.
[0118] 13. A system (400) comprising: a receiving unit for a plurality of first operand values, wherein the first operand values are integer values; a binary adder unit (404) adapted to add the plurality of first operand values using binary addition to yield a total value S; a determination unit (406) adapted to determine a single combined modular correction term D for the binary total of all operand values based on the total value S; and a modular adder unit (408) adapted to perform modular addition of S and D to yield the modular total of the plurality of the first operand values.
[0119] 14. The system according to item 13, wherein the determination unit (406) is adapted to determine the single combined modular correction term D for the binary total of all operand values based on the leading bit of the total value S.
[0120] 15. The system according to item 13 or 14, wherein the binary adder (404) unit comprises: a reduction tree generator module (602) adapted to receive the plurality of integer values from the receiving unit; and an internal binary adder unit (604) adapted to receive the output value of the reduction tree generator module.
[0121] 16. The system according to any one of items 13 to 15, wherein the modular adder (408) comprises: a binary adder (706) adapted to add the total value S and D using binary addition; and a select logic (708) adapted to determine the modular total of the plurality of the first operand values.
[0122] 17. The binary adder (404) unit is: a reduction tree generator module (702) adapted to receive the plurality of large integer values from the receiving unit, the reduction tree generator module being adapted to generate an output as a sum / carry vector pair (704), where the binary sum includes the same bit width as the received large integer value, and where the carry includes a predetermined number of carry bits, the system according to any one of items 13 to 16.
[0123] 18. The modular adder (408) is: a binary adder (706) for binary adding the sum / carry vector pair (704) and the value D, where the value D is determined based on the sum / carry vector pair (704); and a select logic (708) adapted to determine the modular sum of the plurality of the first operand values, the system according to item 17.
[0124] 19. A receiving unit for two second operand values (A, B), where the operand values are integer values; a binary multiplier (502) adapted to receive the two second operand values from the receiving unit; and a correction unit (504) adapted to determine a correction term of the binary multiplier (502) and to provide the plurality of first large integer values (T1,..., Tn), the system according to any one of items 13 to 18.
[0125] 20. A computer program product comprising a computer-readable storage medium embodying program instructions, the program instructions being executable by one or more computing systems or a controller to cause the one or more computing systems to: receive a plurality of first operand values, where the first operand values are integer values; add the plurality of first operand values using binary addition to yield a sum value S; determine a single combined modular correction term D for the binary sum of all operand values based on the leading bit of the sum value S; and perform a modular addition of S and D to yield the modular sum of the plurality of the first operand values.
[0126] The present invention may be a system, method, and / or computer program product at any possible technical detail level. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions to cause a processor to execute aspects of the present invention.
[0127] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0128] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium in each respective computing / processing device.
[0129] The computer-readable program instructions for carrying out the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer as a stand-alone software package, partly on the user's computer, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to carry out aspects of the present invention.
[0130] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0131] These computer-readable program instructions may be provided to a processor of a computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0132] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0133] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may be performed in an order different from that noted in the drawings. For example, two blocks shown in succession may, in fact, be implemented as one step, executed at the same time, substantially simultaneously, in a partially or wholly overlapping manner in time, or the blocks may, in some cases, be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0134] The descriptions of the various embodiments of the present invention have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the present invention. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or technical improvements made to the technology found in the marketplace, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method, comprising: receiving a plurality of first operand values, wherein the first operand values are integer values; adding the plurality of first operand values using binary addition to yield a sum value S; determining a single combined modular correction term D for the binary sum of all operand values based on the leading bit of the sum value S; and performing a modular addition of S and D to yield a modular sum of the plurality of first operand values .
2. The method of claim 1, wherein the sum value S is represented in redundant number format.
3. The method of claim 2, wherein the redundant number format is a reduced radix format.
4. The method according to any one of the preceding claims, further comprising determining a correction value -k*P, where k is a value representing the leading bit of the sum value S.
5. The method of claim 4, wherein the correction value is determined using a look-up table.
6. The method according to any one of the preceding claims, further comprising: performing a binary multiplication of a second integer value A and B, the binary multiplication yielding a binary product M, the binary product M being represented by a plurality of adjacent words of a predetermined number of bits; determining a plurality of coarse-grained modular correction terms Ti for the binary product M, yielding a plurality of adjacent words of a predetermined bit size for each Ti; and using the plurality of terms Ti as the plurality of received first operand values to generate a result of a fast modular multiplication operation.
7. The method of claim 6, wherein the second integer value is represented in redundant number format.
8. The method of claim 7, wherein the redundant number format is a reduced radix format.
9. The method of claim 6, further comprising converting the received second integer value to redundant number format in response to receiving the second integer value in non-redundant format.
10. The method according to any one of the preceding claims, wherein the plurality of first operand values are derived from a selection from the group consisting of a Solinas reduction operation and a Barrett reduction operation.
11. The method according to any one of the preceding claims, wherein A and B are each integer values having a number of bits between 255 and 521.
12. The method according to any one of the preceding claims, wherein the received integer value is an operand value for elliptic curve operations. **Claim 13** A receiving unit for a plurality of first operand values, wherein the first operand values are integer values; A binary adder unit adapted to add the plurality of first operand values using binary addition to yield a sum value S; A determination unit adapted to determine a single combined modular correction term D for the binary sum of all operand values based on the sum value S; and A modular adder unit adapted to perform modular addition of S and D to yield the modular sum of the plurality of the first operand values A system comprising. **Claim 14** The system according to claim 13, wherein the determination unit is adapted to determine the single combined modular correction term D for the binary sum of all operand values based on the leading bit of the sum value S. **Claim 15** The binary adder unit comprises: A reduction tree generator module adapted to receive the plurality of integer values from the receiving unit; and An internal binary adder unit adapted to receive the output value of the reduction tree generator module The system according to claim 13 or 14, having. **Claim 16** The modular adder comprises: A binary adder adapted to add the sum value S and D using binary addition; and Select logic adapted to determine the modular sum of the plurality of the first operand values The system according to any one of the preceding claims 13 to 15, having. **Claim 17** The binary adder unit comprises: A reduction tree generator module adapted to receive the plurality of large integer values from the receiving unit, wherein the reduction tree generator module is: Adapted to generate an output as a sum / carry vector pair, wherein the binary sum includes the same bit width as the received large integer values, and wherein the carry includes a predetermined number of carry bits, The system according to any one of the preceding claims 13 to 16. **Claim 18** The modular adder comprises: A binary adder for binary adding the total / carry vector pair and the value D, where the value D is determined based on the total / carry vector pair; and Select logic adapted to determine the modular total of the plurality of the first operand values The system according to claim 17, comprising the same. **Claim 19** A receiving unit for two second operand values (A, B), where the operand values are integer values; A binary multiplier adapted to receive the two second operand values from the receiving unit; A correction unit adapted to determine a correction term of the binary multiplier and yield the plurality of first large integer values (T1,..., Tn) The system according to any one of the preceding claims 13 to 18, further comprising the same. **Claim 20** A computer-readable storage medium having program instructions embodied thereon, the program instructions causing one or more computing systems to: Receive a plurality of first operand values, where the first operand values are integer values; Add the plurality of first operand values using binary addition to yield a total value S; Determine a single combined modular correction term D for the binary total of all operand values based on the leading bit of the total value S; and Perform a modular addition of S and D to yield the modular total of the plurality of the first operand values A computer program product executable by the one or more computing systems or controllers. **Claim 21** Receiving, using a receiving unit, a plurality of first operand values, where the first operand values are integer values; Adding, using a binary adder unit, the plurality of first operand values using binary addition to yield a total value S; Determining, using a determination unit, a single combined modular correction term D for the binary total of all operand values based on the total value S; and Performing, using a modular adder unit, a modular addition of S and D to yield the modular total of the plurality of the first operand values A computer-implemented method comprising the same.
Citation Information
Patent Citations
Methods and apparatus for incomplete modular arithmetic
US20020059353A1