Easily scalable large integer multiplication hardware implementation circuit
Patent Information
- Application Number
- CN202111423067.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-11-26
AI Technical Summary
[0005]针对以上技术上的不足,本发明提出了一种易于扩展的大整数乘法硬件实现电路,解决了定制化的流控逻辑开发周期长、灵活性低和成本高等缺点,使得设计者能够快速的搭建各种规模的大整数乘法器
[0026] In the technical solution provided by this invention, a large-bit-width multiplier A is divided into M terms of integer A with a bit width of J. x Divide the large-bit-width multiplicand B into N integers B of bit width K. y The multiplier is fed into a hardware implementation circuit consisting of M multipliers, M-1 first-level K-bit adders, a first register, a second register, M+1 second-level K-bit adders, and a carry-lookahead adder, according to a specific timing sequence. The product result of each stage is split into K bits as the first basic unit, and the product results of different stages are accumulated through a cascaded structure, thus realizing a pipelined, parallel, and easily expandable large integer multiplier.
Smart Images

Figure CN116185337B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hardware circuit design, and in particular to a hardware implementation circuit for large integer multiplication that is easily expandable. Background Technology
[0002] Large integers refer to integers that exceed the representational capabilities of computer data types, typically ranging from hundreds to thousands of digits. In applications such as numerical computation, some data can have hundreds of millions or even more decimal digits, sometimes referred to as ultra-large integers. Large integers have a wide range of applications; for example, in number theory, they are frequently used for factorization and prime number detection. In cryptography, public-key cryptography algorithms such as RSA, EIGamal, and elliptic curve cryptography are all based on large integer arithmetic, with modular multiplication, modular exponentiation, and modular reduction operations being extensively used. Among these, large integer multiplication is the most crucial and time-consuming operation.
[0003] To improve the computational efficiency of large integer multiplication, many optimization schemes have been proposed. For extremely large integers with tens of thousands of digits, schemes based on Fast Fourier Transform (FFT) or Number Theory Transform (NTT) are particularly effective. The Strassen algorithm (SSA) and Fürer's algorithm can significantly reduce the number of multiplication operations. However, due to their high algorithm complexity, they lack a clear advantage for large integers of several hundred to several thousand digits.
[0004] In the application of public-key cryptography algorithms, the most commonly used multiplication of large integers ranging from hundreds to thousands of bits still largely employs a divide-and-conquer approach. This involves splitting the large integer into smaller integers and then multiplying and accumulating them multiple times to obtain the final result. For hardware computing platforms such as Field-Programmable Gate Arrays (FPGAs) or Application-Specific Integrated Circuits (ASICs), multiple multiplier units can be used for parallel or pipelining processing. Because reasonable scheduling, shifting, and accumulation based on the timing relationship between input data and intermediate results are necessary, flow control logic often requires customization and optimization for different input modes and data bit widths. This approach suffers from drawbacks such as long development cycles, low flexibility, and high costs. Summary of the Invention
[0005] To address the above-mentioned technical shortcomings, this invention proposes an easily scalable hardware implementation circuit for large integer multiplication, which solves the problems of long development cycles, low flexibility, and high costs associated with customized flow control logic, enabling designers to quickly build large integer multipliers of various sizes.
[0006] This invention provides an easily expandable hardware implementation circuit for large integer multiplication, comprising: M multipliers, M-1 first-level K-bit adders, a first register, a second register, M+1 second-level K-bit adders, and a carry-lookahead adder;
[0007] The M multipliers are respectively connected to the M-1 first-level K-bit adders, the first register is connected to the least significant bit multiplier, and the second register is connected to the most significant bit multiplier;
[0008] M-1 of the first-level K-bit adders are respectively connected to M+1 of the second-level K-bit adders, the first register is connected to the first second-level K-bit adder, and the second register is connected to the (M+1)th second-level K-bit adder;
[0009] M+1 of the second-level K-bit adders are respectively connected to the carry-lookahead adder;
[0010] Among them, the M multipliers are arranged from least significant bit to most significant bit, the M-1 first-level K-bit adders are arranged from least significant bit to most significant bit, and the M+1 second-level K-bit adders are arranged from least significant bit to most significant bit.
[0011] Optionally, the multiplier A and multiplicand B are integers, and the multiplier A is divided into M integers A with a bit width of J. x The multiplier B is divided into N integers B with a bit width of K. y Where x∈{0,1,…,M-1}, y∈{0,1,…,N-1};
[0012] The multiplier A, the multiplicand B, and the integer A x and the integer B y All numbers are hexadecimal integers, and J <= K;
[0013] The multiplier terminals of the M multipliers receive integers A0, A1, ..., A... M-1 And keep N system clocks unchanged; the data received by the multiplicands of the M multipliers are the same integer sequence {B0, B1, ..., B...} N-1}, and the integer B0 is the first system clock, the integer B1 is the second system clock, and so on, until the integer B1 is the Nth system clock. N-1 .
[0014] Optionally, the M multipliers output M integers A on the (S+1)th system clock cycle. x The first-level product result A with the integer B0 x *B0, at the (S+2)th system clock cycle, outputs M integers A. x The second-level product result A with the integer B1 x *B1, sequentially, output M integers A at the S+Nth system clock cycle. x With the integer B N-1 The Nth level product result A x *B N-1 ;
[0015] Wherein, S is the operation time of the multiplier.
[0016] Optionally, for the first-level product result A, respectively x *B0, the second-level product result A x *B1, and so on until the Nth level product result A. x *B N-1 The process is divided into K bits as the first basic unit and fed into M-1 first-level K-bit adders for parallel summation operations.
[0017] At this point, for adjacent high-order and low-order multipliers, the data received at the addend of each of the first-stage K-bit adders is the high J bits of the output of the low-order multiplier, denoted as D. xy The data received by the addend of each of the first-stage K-bit adders is the lower K bits of the output of the higher-order multiplier, denoted as C. (x+1)y .
[0018] Optionally, the first register is connected to the low K bits of the least significant multiplier output, and the second register is connected to the high J bits of the most significant multiplier output, in order to maintain the synchronization timing of the product results at each stage.
[0019] Optionally, the first-level K-bit adder for D xy and C (x+1)y During summation, at most one bit carry is generated, and the output of the first register is denoted as E. 0y The output of the second register is denoted as E. My The sum of the outputs of the M-1 first-stage K-bit adders is denoted as E. 1y E 2y , and so on up to E (M-1)y The carry outputs of the M-1 first-stage K-bit adders are denoted as F. 1y F 2y , and so on up to F (M-1)y .
[0020] Optionally, the results output by the M-1 first-level K-bit adders, the first register, and the second register are used as the second basic unit, and then fed into the M+1 second-level K-bit adders respectively to obtain the cumulative result of the first-level product, the second-level product, and so on up to the Nth-level product.
[0021] Among them, the M+1 secondary K-bit adders are in a cascaded structure.
[0022] Optionally, each of the M+1 second-stage K-bit adders has a first input terminal, a second input terminal, a third input terminal, and a fourth input terminal; the data received at the first input terminal is the carry F output by the lower-level first-stage K-bit adder. (x-1)y The data received at the second input terminal is the sum E output by the first-level K-bit adder. xy The data received at the third input terminal is the sum G output by the high-order two-stage adder. (x+1)(y-1) The data received by the fourth input terminal is the carry H of its own output. x(y-1) .
[0023] Optionally, the least significant bit secondary K-bit adder outputs the N lower K bits of the multiplication result, denoted as G. 0y The lowest K bits of data G are output during the (S+3)th system clock cycle. 00 The (S+4)th system clock outputs the second lowest K bits of data G. 01 This continues until the Nth low-K bit data G is output at the S+N+2th system clock cycle. 0(N-1) .
[0024] Optionally, the carry H output by the M+1 secondary K-bit adders is... xy The sum of the outputs of the M high-order secondary K-bit adders is connected to the carry-lookahead adder. At the (S+N+2)th system clock cycle, G... 1(N-1) G 2(N-1) , and so on up to G M(N-1) Parallel carry processing is performed, and the M high-K bits of the multiplication result, i.e., G, are output synchronously at the S+N+3rd system clock. 0N G 0(N+1) , and so on up to G 0(N+M-1) .
[0025] Optionally, the multiplier, the first-level K-bit adder, and the second-level K-bit adder are all operated in parallel.
[0026] In the technical solution provided by this invention, a large-bit-width multiplier A is divided into M terms of integer A with a bit width of J. x Divide the large-bit-width multiplicand B into N integers B of bit width K. y The multiplier is fed into a hardware implementation circuit consisting of M multipliers, M-1 first-level K-bit adders, a first register, a second register, M+1 second-level K-bit adders, and a carry-lookahead adder, according to a specific timing sequence. The product result of each stage is split into K bits as the first basic unit, and the product results of different stages are accumulated through a cascaded structure, thus realizing a pipelined, parallel, and easily expandable large integer multiplier.
[0027] This invention provides an easily expandable hardware implementation circuit for large integer multiplication. It has the advantages of flexible configuration, pipelined processing, low bandwidth requirements, and resource saving. It is very suitable for packaging into an independent IP to provide the ability to perform large integer multiplication operations for cryptographic algorithm modules, numerical computing applications, etc., thereby quickly building large-scale hardware circuit systems. Attached Figure Description
[0028] Figure 1 This is a partial schematic diagram of the structure of an embodiment of the hardware implementation circuit for easily expandable large integer multiplication according to the present invention;
[0029] Figure 2 This is a schematic diagram of the vertical calculation principle upon which the easily expandable hardware implementation circuit for large integer multiplication of this invention is based;
[0030] Figure 3 This is a timing waveform diagram of an embodiment of the hardware implementation circuit for easily expandable large integer multiplication of the present invention. Detailed Implementation
[0031] To facilitate understanding of the present invention, a more detailed description is provided below with reference to the accompanying drawings and specific embodiments. It should be noted that when an element is described as being "fixed to" another element, it can be directly on the other element, or one or more intermediate elements may exist between them. When an element is described as being "connected to" another element, it can be directly connected to the other element, or one or more intermediate elements may exist between them. The terms "vertical," "horizontal," "left," "right," "inner," "outer," and similar expressions used in this specification are for illustrative purposes only. In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating relative importance or implying the number of indicated technical features. Thus, unless otherwise stated, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature; "multiple" means two or more. The term "comprising" and any variations thereof mean non-exclusive inclusion, where one or more other features, integers, steps, operations, units, components, and / or combinations thereof may be present or added.
[0032] Furthermore, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections via an intermediate medium, or internal communication between two components. All technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.
[0033] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0034] This invention proposes an easily scalable hardware implementation circuit for large integer multiplication, such as... Figure 1 As shown, it includes: M multipliers, M-1 first-level K-bit adders connected to the M multipliers respectively, M+1 second-level K-bit adders connected to the M-1 first-level K-bit adders respectively, a first register, a second register, and carry-lookahead adders. The M multipliers, M-1 first-level K-bit adders, and M+1 second-level K-bit adders are arranged from least significant bit to most significant bit. One end of the first register is connected to the least significant bit multiplier, and the other end is connected to the first second-level K-bit adder. One end of the second register is connected to the most significant bit multiplier, and the other end is connected to the (M+1)th second-level K-bit adder. The M+1 second-level K-bit adders are each connected to a carry-lookahead adder.
[0035] Furthermore, since the multiplier A and multiplicand B are large-bit-width integers, the multiplier A is divided into M terms of integer A with a bit width of J. x The multiplicand B is divided into N terms of integers B with a bit width of K. y Where x∈{0,1,…,M-1}, y∈{0,1,…,N-1};
[0036] Furthermore, in this embodiment, the multiplier A, the multiplicand B, and the integer A x and integer B y All numbers are hexadecimal integers, J <= K; specifically, the multiplier A and multiplicand B are represented in hexadecimal, and the multiplier A is divided into M integers A with a width of J. x That is, A0, A1, ..., A M-1 , then A=A M-1 *2 (M-1)J +…+A1*2 J +A0; Similarly, the multiplicand B is divided into N terms of integers B with a bit width of K. yThat is, B0, B1, ..., B N-1 Then B = B N-1 *2 (N-1)K +…+B1*2 K +B0; The selection of J and K can be determined according to the characteristics required by the target circuit. For example, in Xilinx ultrascale devices, the input bit width of the DSP unit is 27x18, so J = 18 and K = 27 are set. In Intel Stratix 10 devices, the input bit width of the DSP unit is 27x27, so J = 27 and K = 27 are set. In ASIC design, these values can be set according to application requirements. Here, for ease of optimization, it is constrained that J <= K.
[0037] The product of the large integer A and the large integer B is:
[0038]
[0039] Among them, the number of multipliers is the integer A divided by the large integer A. x The number is consistent, M for each multiplier. During the calculation, the multiplier input of each multiplier receives data consisting of a large integer A divided into smaller bit-width integers A0 and A1. x That is, the data received at the multiplier input of the first multiplier is A0, the data received at the multiplier input of the second multiplier is A1, and so on until the data received at the multiplier input of the Mth multiplier is A0. M-1 Furthermore, during the calculation process, N system clocks are kept constant.
[0040] Furthermore, the multiplicand inputs of the M multipliers receive the same data, namely B0, B1, ..., B... N-1 Sequence, and each B y Maintain a single system clock. That is, the first system clock is B0, the second is B1, and so on, until the Nth system clock is B... N-1 ;
[0041] Furthermore, the M multipliers output M integers A on the (S+1)th system clock cycle. x The product A with integer B0 x *B0; Output M integers A at the (S+2)th system clock cycle. x The product of A and integer B1 x *B1; Output M integers A at the S+Nth system clock cycle. x With integer B N-1 The product result A x *B N-1 Where S is the multiplier's operation time, in clock cycles.
[0042] Specifically, such as Figure 2As shown, A x The product of A0 and B0 is denoted as the first-level product result, including A0 B0, A1 B0, A2B0…A M-1 B0; A x The product of A0 and B1 is denoted as the second-level product result, including A0 B1, A1 B1, A2 B1…A M-1 B1; in turn, A x With B N-1 The product of these products is denoted as the Nth level product result, including A0 and B. N-1 A1 B N-1 A2 B N-1 …A M-1 B N-1 .
[0043] Furthermore, in this embodiment, the first-level product result, the second-level product result, and so on up to the Nth-level product result are summed respectively;
[0044] Since any product number A x *B y The bit width of each product is (J+K) bits, and the relative displacement of adjacent product numbers in the same level product result is K bits. In order to process M product results of the same level in parallel, each product number can be split into K bits as the first basic unit, and the overlapping parts can be added in pairs.
[0045] Specifically, according to formula (1), the adjacent product numbers in the same level of product results, such as integer A x+1 *B y With integer A x *B y The relative displacement is K bits. Since when two hexadecimal integers with bit widths of J and K are multiplied, the bit width of the product is (J+K) bits.
[0046] Then we can obtain: A x *B y =D xy *2 K +C xy (2)
[0047] A x+1 *B y =D (x+1)y *2 K +C (x+1)y (3)
[0048] Among them, D xy A represents x With B y The high J bits of the product, C xy A represents x With B y The lower K bits of the product; D(x+1)y A represents x+1 With B y The high J bits of the product, C (x+1)y A represents x+1 With B y The lower K bits of the product. In order to process M products of the same level in parallel, each product number can be split into K bits as the first basic unit, and the overlapping parts can be added pairwise.
[0049] At this point, the outputs of the M multipliers are connected to M-1 first-stage K-bit adders. The data received at the addend input of each first-stage K-bit adder is the high J bits of the output of the lower-order multiplier, denoted as D. xy The data received at the addend input is the lower K bits of the output of the high-order multiplier, denoted as C. (x+1)y .
[0050] Meanwhile, in order to maintain the synchronous timing of the product results at each stage, the first register is connected to the low K bits of the least significant multiplier output, and the second register is connected to the high J bits of the most significant multiplier output.
[0051] Specifically, the low K bits of the least significant multiplier are output through the first register, and the high J bits of the most significant multiplier are output through the second register. Furthermore, the operation time of the first-stage K-bit adder is one system clock cycle.
[0052] Furthermore, the first-level K-bit adder for D xy and C (x+1)y Summation produces at most one carry-over bit. Let the output of the first register be denoted as E. 0y The output of the second register is denoted as E. My The sum of the outputs of M-1 first-level K-bit adders is denoted as E. 1y E 2y , and so on up to E (M-1)y The carry outputs of M-1 first-level K-bit adders are denoted as F. 1y F 2y , and so on up to F (M-1)y .
[0053] Specifically, according to formulas (2) and (3), the calculation process for merging the product numbers of the y-th level product results is as follows:
[0054] A M-1 *B y *2 (M-1)K +…+A1*B y *2 K +A0*B y
[0055] =(D (M-1)y *2K +C (M-1)y )2 (M-1)K +…+(D 1y *2 K +C 1y )2 K +D 0y *2 K +C 0y
[0056] =D (M-1)y *2 MK +(C (M-1)y +D (M-2)y )2 (M-1)K +…+(C 1y +D 0y )2 K +C 0y
[0057] =E My *2 MK +(F (M-1)y *2 K +E (M-1)y )*2 (M-1)K +…+(F 1y *2 K +E 1y )*2 K +E 0y (4)
[0058] As can be seen from the above, when summing all the product numbers in the same level of product results, K can be used as the first basic unit for processing, and parallel processing can be used to improve the operational efficiency of the multiplier and the first-level K-bit adder.
[0059] Furthermore, in this embodiment, the outputs E of the first register and the second register are... 0y E My The sum of the outputs of M-1 first-level K-bit adders, E 1y E 2y , and so on up to E My The carry F output by M-1 first-level K-bit adders 1y F 2y , and so on up to F (M-1)y When summing, K bits are used as the second basic unit, and the accumulation operation of the first-level product result, the second-level product result, and so on up to the Nth-level product result is realized through the cascaded structure of M+1 second-level K-bit adders.
[0060] Specifically, since the product results directly output by M multipliers are M, according to formula (4), after merging and carrying within the same level, there are M+1 K-bit data. When the M+1 data of the first level are added to the M+1 data of the second level, since the product result of the second level is shifted left by K bits compared to the first level, the lowest K bits of the first level do not need to be processed and are already the final result of their respective bits. At this time, it is only necessary to add the M+1 data of the second level to the M data of the first level, and the intermediate result obtained is still M+1. Therefore, it can be seen that only M+1 second-level K-bit adders are needed to accumulate the product results of different levels.
[0061] Furthermore, in this embodiment, the cumulative result of the (y-1)th level is represented by M+1 data points as follows:
[0062] (H M(y-1) *2 K +G M(y-1) )2 MK +…+(H 1(y-1) *2 K +G 1(y-1) )2 K +H 0(y-1) *2 K +G 0(y-1) (5)
[0063] Where y>=1, G x(y-1) H represents the cumulative sum at position x in level y-1. x(y-1) This represents the accumulated carry result at position x in the (y-1)th level;
[0064] The product results of the y-th level are further combined by formula (4) to form:
[0065] (E My +F (M-1)y )*2 MK +(E (M-1)y +F (M-2)y )*2 (M-1)K +…+(E 2y +F 1y )*2 2K +E 1y *2 K +E 0y (6)
[0066] When adding the accumulated result of level (y-1) to the combined product result of level y, it is important to note that the combined product result of level y should be shifted left by K positions. Following the principle of aligning and adding with K positions as the second basic unit, G... 0(y-1) The output of the least significant K-bit adder will be used as the final result for the current K bits. At this point, the coefficients of the least significant K bits are determined by E. 0y H 0(y-1)and G 1(y-1) The coefficients are obtained by adding them together; the coefficient of the highest K digit is then determined by E. My F (M-1)y and H M(y-1) Adding them together gives the result. For any x position in the middle (x>=1, and x<=M-1), then:
[0067] H xy *2 K +G xy =G (x+1)(y-1) +H x(y-1) +E xy +F (x-1)y (7)
[0068] The sum output by the second-level K-bit adder is denoted as G. xy Let H represent the cumulative sum at position x in the y-th stage, and let H be the carry output of the second-stage K-bit adder. xy , represents the accumulated carry result at position x of level y; by cascading the output of the high-level second-level K-bit adder to the input of the low-level second-level K-bit adder, and cascading the carry output of the second-level K-bit adder to its own input, the accumulation operation of the first-level product result, the second-level product result, and so on up to the Nth-level product result is realized.
[0069] Therefore, the summation result at position x in the y-th stage is the sum E of the output of the first-stage K-bit adder at position x. xy The carry F output by the first-stage K-bit adder located at position x-1 (x-1)y The cumulative sum G of the previous level located at position x+1 (x+1)(y-1) and the carry result H of the previous level accumulation located at position x x(y-1) The result is obtained by adding the four data points together.
[0070] Furthermore, in this embodiment, each second-stage K-bit adder has four input terminals. The data received at the first input terminal is the carry F output by the lower-stage first-stage K-bit adder. (x-1)y The data received at the second input is the output E of the first-level K-bit adder. xy The data received at the third input terminal is the output result G of the high-order two-stage K-bit adder. (x+1)(y-1) The data received at the fourth input terminal is the carry H of its own output. x(y-1) .
[0071] Specifically, the outputs of the first register, the second register, and the M-1 first-level K-bit adders are connected to the M+1 second-level K-bit adders. Each second-level K-bit adder has four inputs. By cascading the sum of the outputs of the higher-order second-level K-bit adders to the inputs of the lower-order second-level K-bit adders, and cascading the carry outputs of the second-level K-bit adders to their own inputs, the accumulation operation of the first-level product result, the second-level product result, and so on up to the Nth-level product result is realized.
[0072] The operation time of the two-stage K-bit adder is one system clock cycle.
[0073] Furthermore, the N lower K bits of the multiplication result are output by a second-level K-bit adder of the least significant bit, denoted as G. 0y The lowest K bits of data G are output during the (S+3)th system clock cycle. 00 The (S+4)th system clock outputs the second lowest K bits of data G. 01 This continues until the Nth low-K bit data G is output at the S+N+2th system clock cycle. 0(N-1) .
[0074] Furthermore, the carry outputs of the M+1 second-level K-bit adders and the sum outputs of the M high-order second-level K-bit adders are all connected to the carry-lookahead adder. At the S+N+2th system clock cycle, the carry is applied to G... 1(N-1) G 2(N-1) , and so on up to G M(N-1) Parallel carry processing is performed, and the M high-K bits of the multiplication result, i.e., G, are output synchronously at the S+N+3rd system clock. 0N G 0(N+1) , and so on up to G 0(N+M-1) .
[0075] Specifically, at the (S+N+2)th system clock cycle, the second-level K-bit adder outputs the sum of the first-level product, the second-level product, and so on up to the Nth-level product, G. x(N-1) With carry H x(N-1) G 0(N-1) The final result of the Nth low-K bits is output through a two-stage K-bit adder with the least significant bit. And G... 1(N-1) G 2(N-1) Until G M(N-1) This requires the carry H from the output of the lower-order second-level K-bit adder. x(N-1) The final product is obtained by adding the results together. To improve computational efficiency, a carry-lookahead adder is used to accumulate the result G. 1(N-1) G 2(N-1) Until G M(N-1) , and carry H 0(N-1) H 1(N-1) Until HM(N-1) Parallel computation is performed, and the M high-K bits of the multiplication result, i.e., G, are output in synchronization with the S+N+3 system clock. 0N G 0(N+1) , and so on up to G 0(N+M-1) .
[0076] Furthermore, in this embodiment, the multiplier, the first-level K-bit adder, and the second-level K-bit adder all operate in parallel.
[0077] Specifically, all multipliers and adders operate in parallel and pipelined manner, which not only improves computational efficiency but also saves resources.
[0078] Furthermore, the hardware circuit proposed in this invention does not require additional cache resources. By rationally scheduling the input data, it saves input bandwidth as much as possible, and the computational resources are all implemented using multipliers and adders with small bit widths, which is more conducive to the convergence of synthesis timing. On the other hand, when the bit width of the large integer B used as the multiplicand increases, there is no need to change the hardware structure; only the computation time will increase slightly. If the bit width of the large integer A used as the multiplier increases, it is only necessary to increase the number of parallel multipliers, first-level adders, and second-level adders proportionally according to the bit width. It is even possible to adopt a parameterized parallel processing unit structure from the initial design stage, making the hardware circuit easy to expand for diverse application scenarios.
[0079] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above. For the sake of brevity, they are not provided in detail. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hardware implementation circuit for easily scalable large integer multiplication, characterized in that, The easily expandable large integer multiplication hardware implementation circuit includes: M multipliers, M-1 first-level K-bit adders, a first register, a second register, M+1 second-level K-bit adders, and a carry-lookahead adder; The M multipliers are respectively connected to the M-1 first-level K-bit adders, the first register is connected to the least significant bit multiplier, and the second register is connected to the most significant bit multiplier; M-1 of the first-level K-bit adders are respectively connected to M+1 of the second-level K-bit adders, the first register is connected to the first second-level K-bit adder, and the second register is connected to the (M+1)th second-level K-bit adder; M+1 of the second-level K-bit adders are respectively connected to the carry-lookahead adder; Among them, the M multipliers are arranged from least bit to most bit, the M-1 first-level K-bit adders are arranged from least bit to most bit, and the M+1 second-level K-bit adders are arranged from least bit to most bit. Multiplier A and multiplicand B are integers. Multiplier A is divided into M integer terms with a width of J. The multiplier B is divided into N integer terms with a bit width of K. Where x∈{0,1, …, M-1}, y∈{0,1, …, N-1}; the multiplier A, the multiplicand B, and the integer and the integer All numbers are hexadecimal integers, and J <= K; The data received by the multiplier terminals of the M multipliers are respectively integers. , , ..., And keep N system clocks unchanged; the data received by the multiplicand terminals of the M multipliers are the same integer sequence { , , ..., }, and the first system clock is the integer The second system clock is the integer. In sequence, the Nth system clock is the integer. .
2. The easily expandable large integer multiplication hardware implementation circuit according to claim 1, characterized in that, The M multipliers output M integers at the (S+1)th system clock cycle. With the integer First-order product result * The system outputs M integers at the (S+2)th system clock cycle. With the integer Second-order product result * In sequence, M integers are output on the S+Nth system clock cycle. With the integer The Nth level product result * ; Wherein, S is the operation time of the multiplier.
3. The easily expandable large integer multiplication hardware implementation circuit according to claim 2, characterized in that, For the first-level product results respectively * The second-level product result * This continues until the Nth level product result. * The process is divided into K bits as the first basic unit and fed into M-1 first-level K-bit adders for parallel summation operations. At this point, for adjacent high-order and low-order multipliers, the data received at the addend of each of the first-stage K-bit adders is the high J bits of the output of the low-order multiplier, denoted as... The data received at the addend of each of the first-stage K-bit adders is the lower K bits of the output of the higher-order multiplier, denoted as... .
4. The easily expandable large integer multiplication hardware implementation circuit according to claim 3, characterized in that, The first register is connected to the low K bits of the least significant multiplier output, and the second register is connected to the high J bits of the most significant multiplier output, in order to maintain the synchronization timing of the product results at each stage.
5. The easily expandable large integer multiplication hardware implementation circuit according to claim 3, characterized in that, The first-level K-bit adder pairs and During summation, at most one bit carry is generated, and the output of the first register is recorded as... The output of the second register is denoted as The sum of the outputs of the M-1 first-stage K-bit adders is denoted as follows: , , and so on until The carry outputs of the M-1 first-stage K-bit adders are respectively denoted as... , , and so on until .
6. The easily expandable large integer multiplication hardware implementation circuit according to claim 5, characterized in that, The results output by the M-1 first-level K-bit adders, the first register, and the second register are used as the second basic unit and fed into the M+1 second-level K-bit adders respectively to obtain the cumulative result of the first-level product, the second-level product, and so on up to the Nth-level product. Among them, the M+1 secondary K-bit adders are in a cascaded structure.
7. The easily expandable large integer multiplication hardware implementation circuit according to claim 6, characterized in that, Each of the M+1 second-level K-bit adders has a first input terminal, a second input terminal, a third input terminal, and a fourth input terminal; the data received at the first input terminal is the carry output by the lower-level K-bit adder. The data received at the second input terminal is the sum of the outputs of the same-level K-bit adder. The data received at the third input terminal is the sum of the outputs of the high-order two-stage adder. The data received by the fourth input terminal is the carry from its own output. .
8. The easily expandable large integer multiplication hardware implementation circuit according to claim 6, characterized in that, The least significant bit of the second-level K-bit adder outputs the N least significant bits of the multiplication result, denoted as... The lowest K bits of data are output on the (S+3)th system clock cycle. The (S+4)th system clock outputs the second lowest K bits of data. This continues until the Nth low-K bit data G is output at the S+N+2th system clock cycle. 0(N-1) .
9. The easily expandable large integer multiplication hardware implementation circuit according to claim 7, characterized in that, Carry output from M+1 of the second-level K-bit adders The sum of the outputs of the M high-order secondary K-bit adders is connected to the carry-lookahead adder. At the (S+N+2)th system clock cycle, G... 1(N-1) G 2(N-1) , and so on up to G M(N-1) Parallel carry processing is performed, and the M high-K bits of the multiplication result, i.e., G, are output synchronously at the S+N+3rd system clock. 0N G 0(N+1) , and so on up to G 0(N+M-1) .
10. The easily expandable large integer multiplication hardware implementation circuit according to claim 1, characterized in that, The multiplier, the first-level K-bit adder, and the second-level K-bit adder all operate in parallel.
Citation Information
Patent Citations
Large-number modular multiplier circuit
CN102117195A
Nonvolatile 8-bit Booth multiplier based on RRAM
CN110196709A