Low-delay Montgomery modular multiplier

By introducing shift and flow rate calculations into the Montgomery mode multiplier, the critical path extension problem caused by calculating the quotient value under the large-position wide modulus is solved, and the low latency and high efficiency of the Montgomery mode multiplier is achieved, which improves the computing speed and reduces the area overhead.

CN120085833APending Publication Date: 2025-06-03BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411843198.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing Montgomery module multiplier has a critical path extension caused by calculating the quotient value under the large-bit width modulus, resulting in a low computing speed and the implementation architecture is sensitive to addition and multiplication, making it difficult to maintain the original computing performance while the modulus length increases.

Method used

By introducing shift and flowing quotient value operations, the dependence of the parallel part on quotient value calculation is eliminated, the intermediate value representation in iteration operations is changed, the compression series and period length of each iteration is reduced, and the calculation speed of the module multiplier is improved.

Benefits of technology

The low latency and high efficiency of Montgomery die multiplier are achieved, with at least 35% higher speed than the fastest implementation speed, and less additional area overhead and a 30% lower product of the delay area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085833A_ABST
    Figure CN120085833A_ABST
Patent Text Reader

Abstract

The invention discloses a low-delay Montgomery modular multiplier. The modular multiplier comprises an iteration control unit, a parallel part and generation unit, a quotient operation unit of a multi-stage pipeline, an output shifting unit, a part and compression unit and a final reduction unit. According to the modular multiplier, an intermediate result of each round of iteration is expressed in a triple form, and the number of compression stages is reduced while long-bit-width addition is avoided. According to the method, shift operation is introduced, the quotient operation process is dispersed to a plurality of cycles for pipeline implementation, at the cost of introducing a plurality of additional cycles, the key path of each round of iteration is greatly shortened, the overall output delay of the modular multiplier is greatly reduced, the method is suitable for scenes with high-speed operation requirements, and the longer the modulus length is, the more obvious the acceleration effect is.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of information security and hardware design, and involves a modulo multiplier. Background Art

[0002] Modulo multiplication is widely used in the field of cryptography because it can better conceal the information of the two input multipliers, and it is one of the most important primitives in public-key cryptography. The complexity of modulo multiplication comes from the modulo operation. The ordinary modulo process needs to calculate the quotient of the product and the modulus, which involves quite complex division operations. An efficient modulo multiplication scheme needs to avoid division operations. For general moduli, Barrett reduction and Montgomery reduction are two effective general modulo reduction methods.

[0003] Barrett reduction calculates R / 2 through a group of numbers R and K K to approximate the reciprocal of the modulus and thus complete the modulo process by using multiplication and truncation; while Montgomery reduction adds the number to be modulo and several multiples of the modulus so that the obtained sum can be divisible by a certain power of 2, thus avoiding complex division operations and converting them into shift and multiplication operations with lower overhead. Analyzed from the computational complexity, Montgomery reduction has lower computational complexity, but the input needs to be converted into Montgomery form, which is more suitable for operations such as modular exponentiation that continuously call modulo multiplication in large quantities; while the computational complexity of Barrett modulo multiplication is slightly higher, but it does not require form conversion and has a wider range of application scenarios.

[0004] Currently, Montgomery modulo multiplication is widely used in RSA cryptosystems, elliptic curve cryptosystems, cryptosystems based on discrete logarithms over finite fields, and cryptosystems based on bilinear pairings as the most critical and complex basic primitive operation. In order to improve the performance of the cryptosystem, a hardware platform with a higher degree of parallelism is often selected to accelerate the implementation of complex primitives. Therefore, the operation efficiency of the Montgomery modulo multiplier directly determines the running efficiency of the top-level cryptosystem.

[0005] With the continuous increase of security requirements, the modulus length of modulo multiplication operations has also increased rapidly. In cryptosystems such as RSA or Paillier, the modulus length to meet the current security requirements is 4096 - 8192 bits. At large bit-width moduli, the overhead of addition and multiplication operations increases rapidly, and the demand for a Montgomery modulo multiplier that can calculate efficiently further increases. Currently, the existing Montgomery modulo multipliers generally have the phenomenon that the critical path is extended due to calculating the quotient value, resulting in a low running speed of the modulo multiplier; at the same time, the existing implementation architectures of Montgomery modulo multipliers are strongly dependent on the implementation of addition and multiplication, are sensitive to the modulus length, and it is difficult to maintain the original operation performance when the modulus length increases. Summary of the Invention

[0006] The object of the present invention is to provide a Montgomery multiplier, which eliminates the phenomenon of the extended critical path commonly existing in the existing Montgomery multipliers due to the calculation of the quotient value, and breaks through the speed bottleneck of the implementation architecture of the existing Montgomery multipliers. Specifically, the present invention eliminates the dependence of the remaining parallel parts on the quotient calculation by introducing shift and pipelined quotient operations, and eliminates the negative effect of the quotient calculation on the critical path in the original architecture; by changing the representation form of the intermediate value in the iterative operation, the compression level of each round of iteration is reduced, the cycle length of each round of iteration is further shortened, and the operation speed of the multiplier is greatly improved.

[0007] The present invention realizes the object through the following technical solutions: a low-latency Montgomery multiplier. The multiplier includes an iterative control unit, a parallel part and a generation unit, a multi-stage pipelined quotient operation unit, an output shift unit, a partial sum compression unit, and a final reduction unit.

[0008] Among them, the iterative control unit includes an operand circular shift register and an iteration round counter. The operand circular shift register inputs one of the multipliers at the start of the calculation, and in the subsequent calculation process, it circularly shifts right by k bits in each iteration. This register is responsible for generating a specific domain segment of the input multiplier required for each round of iteration. The iteration round counter is responsible for counting the number of iteration rounds and generating an output enable signal after reaching the target round; the partial sum generation unit calculates the generation of the partial sum of the full length - word length, which can be used to calculate the multiplication related to the input multiplier and the multiplication related to the modulus; the multi-stage pipelined quotient operation unit calculates the quotient value in a multi-stage pipelined form according to the input result of the iterative intermediate value register, and this process will cover several cycles; the output shift unit calculates the low-order carry of the iterative intermediate value represented by multiple elements, deletes the low-order part, and sends the intermediate value of this round to the partial sum compression unit; the partial sum compression unit calculates the sum of two groups of partial sums XY i , qM and the sum of a group of intermediate values, which is realized by a series of juxtaposed Dadda adders or CSA adders, and outputs the intermediate value represented by three elements and sends it to the iterative intermediate value register; the final reduction unit calculates a sum in multiple cycles, corresponding to the final reduction part of the Montgomery multiplication algorithm.

[0009] The inputs of the iterative control unit are a multiplier Y of full length, a clock signal, and a working enable signal. The global parameters include the modulus length MOD_LEN, the radix k, and the counter length; the outputs are the multiplier domain segment Y of length k i and the output enable signal. When the counter value is i, the multiplier domain segment Y i = Y[ik + 2k - 1:ik + k], and the output enable signal satisfies that it is 1 when the counter value reaches MOD_LEN / k + t - 1 and sets the counter to 0, and is 0 at other times.

[0010] The input of the partial sum generation unit is a multiplier T of a full length m and a multiplier T of a word length w , and the output is a series of partial sums P generated according to the partial sum generation algorithm i , P i satisfying ∑P i = T m T w That's all

[0011] The input of the multi-stage pipelined quotient operation unit is the low k(t + 1) bits of three iterative intermediate value registers. Here, the three iterative intermediate value registers are denoted as Z 0 , Z 1 , Z 2 , and after t cycles of calculation, the unit outputs the quotient value q, q = ((Z 0 [kt + k - 1:0] + z 1 [kt + k - 1:0] + Z 2 [kt + k - 1:0]) × M′[kt - 1:0] mod 2 kt ). The output of each cycle of the pipeline can be customized according to the cycle duration requirements

[0012] The input of the output shift unit is the low k bits of three iterative intermediate value registers. The output is two k-bit bits generated by calculating Z 0 [k - 1:0] + Z 1 [k - 1:0] + Z 2 [k - 1:0]. These bits can be calculated by the look-up table method or the logic function method. If the input signals are Z 2 [k - 3], Z 1 [k - 2], Z 0 [k - 1], Z 2 [k - 4], Z 1 [k - 3], Z 0 [k - 2] from high to low in sequence, then the 64-bit truth table values of the corresponding two carry bits are 0xfffffffffffffffe and 0xfffefe80fe808000. Here, the highest bit corresponds to the table value of all 1s of the input signal, and the lowest bit corresponds to the table value of all 0s of the input signal

[0013] The input of the partial sum compression unit includes three parts, a series of multiplication partial sums XY related to the multiplier i , a series of multiplication partial sums qM related to the modulus, and the shifted iterative intermediate value Z 2 [:k - 2], Z 1 [:k - 1], Z 0[:k], the output is the ternary compression result p after a series of Dadda adders or CSA adders arranged in a tree structure 0 , p 1 , p 2 , and the weights of the ternary compression results are 0, 1, and 2 respectively.

[0014] The input of the final reduction unit is the clock signal, the shifted iterative intermediate value, and the modulus M, and the output is the modular multiplication result Z ∈ [0, M). The final reduction sums the three iterative intermediate values with M and -M respectively, and finally selects the positive one as the modular multiplication result.

[0015] The outstanding advantage of the present invention is that: the modular multiplier in the present invention eliminates the dependence of the remaining parallel parts on the quotient calculation by introducing shift and pipelined quotient operations, and eliminates the negative effect of the quotient calculation on the critical path in the original architecture; by changing the representation form of the intermediate value in the iterative operation, the compression level of each round of iteration is reduced, and the cycle length of each round of iteration is further shortened. The modular multiplier implemented on the FPGA platform has at least a 35% improvement in the fastest implementation speed compared to the existing ones, and has less additional area overhead. The delay-area product is 30% lower than the lowest existing implementation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0017] Figure 1 shows the overall architecture of the modular multiplier in the present invention;

[0018] Figure 2 shows two basic structures of the compression unit in the present invention;

[0019] Figure 3 shows the variable names and their meanings in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0020] The implementation manner of the present invention will be introduced below in conjunction with the drawings. The content described here is only used to illustrate and explain the present invention and is not used to limit the present invention.

[0021] As Figure 1 shown, a low-latency Montgomery modular multiplier according to the present invention includes an iterative control unit, two parallel partial sum generation units, a multi-stage pipelined quotient operation unit, an output shift unit, a partial sum compression unit, and a final reduction unit.

[0022] The iterative control unit includes an operand circular shift register and an iteration round counter. The operand circular shift register inputs one of the multipliers at the start of the calculation and, in subsequent calculation processes, circularly shifts right by k bits in each iteration. This register is responsible for generating specific domain segments of the input multiplier required for each round of iteration. The iteration round counter is responsible for counting the number of iteration rounds and generating an output enable signal after reaching the target round; the partial sum generation unit calculates the generation of partial sums of full length - word length, which can be used to calculate the multiplication related to the input multiplier and the multiplication related to the modulus; the multi - stage pipelined quotient operation unit calculates the quotient value in a multi - stage pipelined form based on the input result of the iterative intermediate value register, and this process covers several cycles; the output shift unit calculates the low - order carry of the iterative intermediate value in multiple representations, deletes the low - order part, and sends the intermediate value of this round to the partial sum compression unit; the partial sum compression unit calculates the sum of two sets of partial sums XY i , qM, and the sum of a set of intermediate values, which is implemented through a series of juxtaposed Dadda adders or CSA adders, and outputs the intermediate value in ternary representation, which is sent to the iterative intermediate value register; finally, the final reduction unit calculates a sum in multiple cycles, corresponding to the final reduction part of the Montgomery modular multiplication algorithm.

[0023] The inputs of the iterative control unit are a multiplier Y of full length, a clock signal, and a work enable signal. The global parameters include the modulus length MOD_LEN, the radix k, and the counter length; the outputs are the multiplier domain segment Y of length k i and the output enable signal. When the counter value is i, the multiplier domain segment Y i = Y[ik + 2k - 1:ik + k]. The output enable signal is 1 when the counter value reaches MOD_LEN / k + t - 1 and resets the counter to 0, and is 0 at other times.

[0024] The inputs of the partial sum generation unit are a multiplier T of full length and a multiplier t of word length, and the output is a series of partial sums P i , P i satisfying ∑P i = Tt. There are two parallel partial sum generation units in this architecture, which are respectively used to calculate two different full length - word length multiplications.

[0025] The inputs of the multi - stage pipelined quotient operation unit are the low k(t + 1) bits of three iterative intermediate value registers. Here, the three iterative intermediate value registers are respectively denoted as Z 0 , Z 1 , Z 2 . After t cycles of calculation, this unit outputs the quotient value q, q = ((Z 0 [kt + k - 1:0]+Z 0 [kt + k - 1:0]+Z0 [kt + k - 1:0]) × M′[kt - 1:0] mod 2 kt ) >> k(t - 1), the output of each cycle of the pipeline can be customized according to the cycle duration requirements.

[0026] The input of the output shift unit is the lower k bits of three iterative intermediate value registers. The output is to calculate Z 0 [k - 1:0] + Z 1 [k - 1:0] + Z 2 The two k - th bit generated by [k - 1:0]. This bit can be calculated by the look - up table method or the logic function method. If the input signals are Z 2 [k - 3], Z 1 [k - 2], Z 0 [k - 1], Z 2 [k - 4], Z 1 [k - 3], Z 0 [k - 2], then the 64 - bit truth table values of the corresponding two carry bits are 0xfffffffffffffffe and 0xfffefe80fe808000. Here, the highest bit corresponds to the table value of all - 1 input signals, and the lowest bit corresponds to the table value of all - 0 input signals.

[0027] The input of the partial - sum compression unit includes three parts, a series of multiplication partial - sums XY related to the multiplier i 、a series of multiplication partial - sums qM related to the modulus, and the shifted iterative intermediate value Z 2 [:k - 2], Z 1 [:k - 1], Z 0 [:k], and the output is the ternary compression result p after a series of Dadda adders or CSA adders arranged in a tree structure 0 , p 1 , p 2 , and the weights of the ternary compression result are 0, 1, 2 respectively.

[0028] The input of the final reduction unit is the clock signal, the shifted iterative intermediate value, and the modulus M. The output is the modular multiplication result Z ∈ [0, M). The final reduction sums the three iterative intermediate values with M and - M respectively, and finally selects the positive - valued item as the modular multiplication result.

[0029] Denote the modulus length in the present invention as MOD_LEN, the data length processed in each round as k, and the number of pipeline stages in the quotient operation unit with multi - stage pipelining as t. Then the registers included in the present invention are:

[0030] Three iterative intermediate value registers for storing iterative intermediate values, and the length of each register is MOD_LEN + k(t + 1);

[0031] A count register that requires a count value covering a length of MOD_LEN / k + t;

[0032] A working status register that marks whether the modular multiplier is in the working phase of the round function operation;

[0033] An operand circular shift register that avoids the use of a large number of selectors through circular shifting and is used to generate the required multiplier field segment with a length of MOD_LEN;

[0034] A quotient register used to store the quotient value with a length of k.

[0035] A pipeline register used to store the data generated at each stage of the multi-stage pipelined quotient operation unit. The number and length of the registers are both related to the implementation method.

[0036]

[0037] Combined with the above description, the specific operation process of the Montgomery modular multiplier of the present invention is as follows:

[0038] Steps 4 and 5 represent the structure of the round function. Given the modulus length MOD_LEN and the iteration base k, the low-latency Montgomery modular multiplier of the present invention can be gradually built up according to the following steps, which will be described in detail with examples below.

[0039] Step 1: Calculate the number of pipeline stages t. The logical delay of each stage in the pipeline should be less than the logical delay of the round function. The total logical delay of the round function in the architecture is the partial sum generation logic plus the partial sum compression logic. The fixed-bit-width shift has no overhead.

[0040] Example of Step 1: Here, an example meeting the above requirements is given. When k = 16, t = 4 can be selected.

[0041] Step 2: Build a multi-stage pipelined quotient operation unit. The quotient operation unit calculates This process includes three processes: first calculate as a three-term summation of kt bits, then calculate a kt × kt multiplication once, and finally complete the whole process by shifting.

[0042] Example Step 2: When t = 4, calculate q through a three-stage pipeline. In the first stage, calculate the sum of the lower 80 bits of the three iterative intermediate value registers. The lower 16 bits of this value are all 0, so the higher 64 bits are output and denoted as z_sum[63:0]. In the second stage, use eight 16×16 parallel sub-multipliers to calculate z_sum[15:0]×M’[15:0], z_sum[31:16]×M’[15:0], z_sum[47:32]×M’[15:0], z_sum[15:0]×M’[31:16], z_sum[31:16]×M’[31:16], z_sum[47:32]×M’[31:16], z_sum[15:0]×M’[47:32], z_sum[31:16]×M’[47:32], respectively. After concatenating them into seven partial sums, input them into the Figure 2 shown tree compressor and compress them into four terms. In the third stage, sum the four partial sums and then shift to obtain the quotient value q.

[0043] Step 3: Generate partial sums XY in parallel i r t , It can be implemented in various ways such as sub-multipliers or logic gates. The multiplication related to r can be achieved through shift without overhead.

[0044] Example Step 3: A 16×16 sub-multiplier array can be used to generate 4 partial sums, or the partial sums of the array multiplier can be generated through logic gates, generating a total of 32 partial sums. Taking the logic gate-based method with XY i as an example, |Y i | = 16, then pp i = {MOD_LEN{Y i [i]}}&X.

[0045] Step 4: Use a parallel compression tree to compress the partial sums. The structure of the parallel compression tree is similar to the Figure 2 balanced compression tree shown in

[0046] Denote the input as pp i , i ∈ [0, 5]. Then the logical expressions of the CSA are S = pp 0 ⊕pp 1 ⊕pp 2 , C = pp 0 &pp 1 |pp 0 &pp 2 |pp 1 &pp 2 .

[0047] Example Step 4: Here, a 16×16 sub-multiplier partial sum generation is adopted. While generating the partial sums, three iterative intermediate values are compressed into two through 1-level CSA. After the partial sum generation, the 4-term partial sums and the two intermediate values are input into a 6-3 compressor to be compressed into 3 terms, and then input into the iterative intermediate value register.

[0048] Step 5: When the counter reaches MOD_LEN / k + t - 1, the operation enable is turned off and enters the final reduction element to calculate Z 0 +Z 1 +Z 2 -M and Z 0 +Z 1 +Z 2 , and select the positive one of the two for output.

[0049] Example Step 5: Perform 256-bit addition in each round. After performing MOD_LEN / 256 times, select the one with the highest bit being 0 (positive value) as the modular multiplication result for output.

[0050] The modular multiplier disclosed in the present invention includes an iterative control unit, a parallel partial sum generation unit, a multi-stage pipelined quotient operation unit, an output shift unit, a partial sum compression unit, and a final reduction element. By introducing shift and pipelined quotient value operations, the dependence of the rest of the parallel part on the quotient value calculation is eliminated, and the negative effect of the quotient value calculation on the critical path in the original architecture is eliminated; by changing the representation form of the intermediate values in the iterative operation, the compression level in each round of iteration is reduced, the cycle length of each round of iteration is further shortened, and the operation speed of the modular multiplier is greatly improved.

Claims

1. A low-latency Montgomery modular multiplier, characterized in that: The modular multiplier includes an iterative control unit, a parallel partial sum generation unit, a multi-stage pipeline quotient calculation unit, an output shift unit, a partial sum compression unit and a final reduction simple unit. The iteration control unit obtains the input data domain required for each round of iteration by cyclic shifting, and controls the output enable signal value according to the counter result; The partial sum generation unit calculates the partial sum generation of the full length-word length for calculating the multiplication with respect to the input multiplier and the multiplication with respect to the modulus; The quotient calculation unit of the multi-stage pipeline calculates the quotient value in the form of a multi-stage pipeline according to the input result of the iterative intermediate value register; After the output shift unit calculates the low-order carry of the iterative intermediate value represented by the multivariate representation, the low-order part is deleted and the intermediate value of the round is sent to the partial sum compression unit; The partial sum compression unit calculates two sets of partial sums XY i The accumulation of , qM and a group of intermediate values ​​is realized by a parallel Dadda adder or a CSA adder, and the intermediate value represented by the ternary is output and sent to the iterative intermediate value register; The final reduction element calculates a cumulative sum in multiple cycles to implement the final reduction part of the Montgomery modular multiplication algorithm.

2. The low-latency Montgomery modular multiplier according to claim 1, characterized in that: The iteration control unit includes an operand circular shift register and an iteration round counter; The operand circular shift register inputs one of the multipliers at the beginning of the calculation, and in the subsequent calculation process, it is circularly shifted right by k bits each time it is iterated. The register is responsible for generating a specific domain segment of the input multiplier required for each round of iteration; The iteration round counter is responsible for counting iteration rounds and generating an output enable signal after reaching a target round.

3. The low-latency Montgomery modular multiplier according to claim 1, characterized in that: Two parallel partial and generation units, calculating XY simultaneously i r t and 4. The low-latency Montgomery modular multiplier according to claim 3, characterized in that: The partial sum generation unit is for a full length multiplier input T m and a word-length multiplier input T w The output is a series of partial sums P generated by the partial sum generation algorithm. i , P i Satisfy ∑P i =T m T w .

5. The low-latency Montgomery modular multiplier according to claim 1, characterized in that: When the input triplet Z0[kt+k-1:0], Z1[kt+k-1:0], Z2[kt+k-1:0] is taken, the quotient operation unit of the multi-stage pipeline completes the tuple weight correction, tuple value accumulation, kt×kt multiplication and k(t-1)-bit right shift operation after t-1 cycles, and then outputs the quotient value to the quotient value register to participate in the partial sum generation.

6. The low-latency Montgomery modular multiplier according to claim 1, characterized in that: When the lower k bits of the three iterative intermediate value registers are input, the output shift unit calculates the two k-th bits generated by Z0[k-1:0]+Z1[k-1:0]+Z2[k-1:0] and appends them to the two left shift parts and field segment a respectively. i Br t The lowest blank bit in the iterative intermediate value register is discarded and sent to the partial and compression unit.

7. The low-latency Montgomery modular multiplier according to claim 6, characterized in that: The output shift unit may calculate the carry bit by a table lookup method or a logic function method.

8. The low-latency Montgomery modular multiplier according to claim 1, characterized in that: After inputting the partial sums generated by the two partial sum generation units and the iterative intermediate values ​​processed by the CSA, the partial sum compression unit compresses the partial sums into triplets through a balanced compression tree with the CSA or Dadda adder as a leaf node, and sends the triplets to the iterative intermediate value register.

9. The low-latency Montgomery modular multiplier according to claim 8, characterized in that: The iteration intermediate values ​​are represented in the form of triplets and stored in three iteration intermediate value registers of length MOD_LEN+k(t+1).

10. The low-latency Montgomery modular multiplier according to claim 1, characterized in that: When the counter reaches MOD_LEN / k+t-1, the final approximation simple element turns off the work enable, iteratively calculates the addition with a total length of MOD_LEN+k bits in multiple cycles, calculates Z0+Z1+Z2-M and Z0+Z1+Z2 respectively, and selects the positive output of the two as the final modular multiplication result.

Citation Information

Cited By

  • Method and device for realizing Montgomery multiplication and reduction method

    CN121523645A