A high-performance Montgomery modular multiplication method based on NLP representation

By adopting a high-performance Montgomery modular multiplication method based on NLP characterization in the modular multiplication circuit, the problem of large resource overhead and low efficiency of the modular multiplication performance is solved, and more efficient and low power consumption is achieved.

CN115016765BActive Publication Date: 2025-06-06SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210774431.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-01
Publication Date
2025-06-06
Estimated Expiration
2042-07-01

AI Technical Summary

Technical Problem

The modular multiplier circuit design resource overhead and low operating efficiency, making it difficult to meet the high performance needs of the elliptic curve public key cryptography system.

Method used

The high-performance Montgomery modular multiplication method based on NLP characterization is adopted, and the modular multiplication algorithm is optimized to reduce calculation delay and hardware resource overhead through second-order NLP multiplication calculation and carry compensation mechanism.

Benefits of technology

It effectively reduces the circuit complexity and power consumption of the modular multiplication circuit, improves the working frequency and performance of the modular multiplication circuit, and achieves lower area and power consumption costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115016765B_ABST
    Figure CN115016765B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-performance Montgomery modular multiplication method based on NLP representation, inputting the product T of a and b; where T H is the high 256 bits of T, and T L is the low 257-bit part of T with a carry signal; calculating the product M of TL and p, and the multiplication of the second-order NLP representation only calculates the low 256 bits ML of M; calculating the product Q of ML and invp, where Q H is the high 256 bits of Q, and Q L is the low 257-bit part of Q with a carry signal; whether the logical OR of Q L [255] and T L [255] is 1, if it is 1, then C = T H + Q H + 1, if it is 0, then C = T H + Q H ; determining whether C is greater than p, if it is greater than p, then output C - p, otherwise output C. The present invention ignores the high digits of partial multiplication, and at the same time proposes carry-save accumulation splitting, that is, it meets the function and improves the advantages of the modular multiplier in terms of performance, power consumption and area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of cryptographic circuits and information security, and in particular relates to a high-performance Montgomery modular multiplication method based on NLP characterization. Background Art

[0002] The ECC cryptographic system has been widely used in high-performance data-intensive application scenarios based on TLS. It is a security protocol that provides security and data integrity for Internet communications and has become the industrial standard for Internet confidential communications. The transport layer security protocol is based on the elliptic curve public key cryptography system and uses key algorithms to provide endpoint identity authentication and communication confidentiality on the Internet.

[0003] Elliptic curve public key cryptography algorithms, including classic ECC encryption and decryption algorithms such as ECDH key agreement algorithm and ECDSA digital signature verification. The key to the hardware design of the ECC protocol layer lies first in whether the encryption and decryption algorithm itself has high performance advantages, and secondly in the performance of the called elliptic curve arithmetic unit, and the performance of the elliptic curve arithmetic unit is ultimately determined by the finite field arithmetic unit. Since the protocol layer is at the top level of data encryption and decryption operations, directly optimizing the hardware mapping of typical ECC algorithm protocols such as ECDSA can achieve the most obvious performance improvement and reduction in hardware resource overhead.

[0004] The finite field operation layer is located at the bottom of the ECC hierarchical structure, and its hardware resource overhead directly determines the area and power consumption of the entire ECC public key cryptography system. Since all calculations of the ECC encryption and decryption algorithm will eventually be decomposed into basic arithmetic operations on the finite field, including modular addition, modular multiplication, modular inverse and modular exponentiation. Among them, the computational complexity of modular exponentiation and modular inverse is relatively high, and is usually implemented by scheduling modular multiplication calculation units multiple times. The modular multiplication algorithm is the core of the elliptic curve public key cryptography system. From a time perspective, modular multiplication operations occupy most of the time in a single ECC encryption and decryption operation; from a space perspective, the modular multiplication module in the ECC basic calculation unit requires the most hardware resources, and more complex operations such as modular exponentiation and modular inverse are completed by calling the modular multiplication algorithm.

[0005] In summary, the implementation cost of modular multiplication is directly related to the overall implementation cost of the system. Therefore, the computational efficiency and hardware resource overhead of the modular multiplication algorithm directly determine the superiority of the entire elliptic curve public key cryptography system. Summary of the invention

[0006] The present invention aims to provide a high-performance Montgomery modular multiplication method based on NLP characterization to solve the technical problems of large resource overhead and low operating efficiency in modular multiplier circuit design.

[0007] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:

[0008] A high-performance Montgomery modular multiplication method based on NLP representation. The modular length of modular multiplication is 256. The input parameters include 256-bit input operands A and B and 256-bit prime field parameters p and invp, where p is the modulus and invp is the value of p in 2 256 The inverse on the field satisfies invp×pmod2 N =1, the output is the modular multiplication result R = (A*B) mod p, and the calculation steps are:

[0009] Step 1: Calculate the product of input A and B and use the second-order non-minimum positive number NLP method to characterize the calculation result:

[0010] (T H ,T L ) = NLPmult(A,B)

[0011] Assume that T is the standard product of input operands A and B, that is, T = A*B, with a bit width of 512 bits. T is represented using the second-order NLP method: where T H The high 256 bits of T, T L The lower 256 bits of T also carry a 1-bit carry signal, a total of 257 bits, and satisfy T = A * B = (T H *2 256 )+T L ;

[0012] Step 2: Calculate T L , p and use the second-order NLP method to represent the calculation result M, and only calculate the lower 256 bits of M L :

[0013] (M H ,M L )=NLPmult(T L ,p)

[0014] Let M be the output of step 1. L The standard product result of the modulus p is 512 bits wide. M is represented by the second-order NLP method: H The high 256 bits of M. L The lower 256 bits of M also carry a carry signal, totaling 257 bits, satisfying M = T L *p=(M H *2 256 )+M L In this step, M H No calculation is performed, that is, M H Constantly 0, only calculate M L ;

[0015] Step 3: Calculate M L The product of invp and invp is represented by the second-order NLP method.

[0016] (Q H ,Q L )=NLPmult(M L ,invp)

[0017] Where Q H The high 256 bits of Q. L The low-order part of Q with the carry signal is 257 bits:

[0018] Let Q be the M calculated in step 2 L The standard product result with the elliptic curve encryption parameter invp has a bit width of 512 bits. Q is represented by NLP: H The high 256 bits of Q. L The lower 256 bits of Q and a carry signal totaling 257 bits, and satisfying: Q = M L *invp=(Q H *2 256 )+Q L ;

[0019] Step 4: Q calculated in step 3 L

[255] and T calculated in step 1 L

[255] Perform a logical OR operation and determine whether the result of the logical OR operation is 1. If it is 1, calculate C = T H +Q H +1, if it is 0, calculate C = T H +Q H ; where Q L

[255] is Q L The second highest bit is Q L No. 256; T L

[255] is T L The second highest bit, that is, T L 256th bit; C is the intermediate output value of the Montgomery modular multiplication, which is used to calculate the final output result R, which is C=((T H +Q H )*2 N +T L +Q L ) / 2 N , using ((T H +Q H )*2 N +T L +Q L) The lower N bits of the expression result are all 0, so replace C with C=T under NLP representation. H +Q H +e, converting 512-bit large number addition into 256-bit addition, without accumulation (T L +Q L ) to obtain the direction (T H +Q H ), where e is a single-bit high-order compensation value of 1 or 0 to replace (T L +Q L ) / 2 N , the value corresponding to e is Q calculated in step 3 L

[255] and T calculated in step 1 L

[255] Perform a logical OR operation;

[0020] Step 5: Determine whether C calculated in step 4 is greater than p. If it is greater than p, the final output result R=Cp is obtained. If it is not greater than p, the final output result R=C is obtained.

[0021] Furthermore, the 256-bit multiplication calculation in step 1, step 2, and step 3 is replaced by second-order NLP multiplication calculation, and the serial calculation of the traditional large number multiplication Z=X*Y is converted into a two-step parallel calculation. The corresponding two-step parallel second-order NLP multiplication calculation is expressed as:

[0022] (Z H ,Z L )=NLPmult(X,Y)=(X H *2 128 +X L )*(Y H *2 128 +Y L )

[0023] =(X H *Y H )2 256 +[(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 +X L *Y L

[0024] =(X H *Y H )2 256 +PP H +PP L+(X L *Y L )

[0025] =(X H *Y H +PP H ,PP L +X L *Y L )

[0026] Where X and Y are the input 256-bit multipliers, X H The high 128 bits of X. L is the lower 128 bits of X, Y H The high 128 bits of Y. L The lower 128 bits of Y, PP H is [(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 The high 256 bits of PP L is [(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 The lower 256 bits of Z are the output 512-bit large number multiplication results. H The high 256 bits of Z. L The lower 256 bits of Z with a carry signal, totaling 257 bits; satisfying Z = X * Y = (Z H *2 256 )+Z L , two-step parallel second-order NLP multiplication to obtain two output results Z H and Z L It exists independently to characterize the result Z of large number multiplication.

[0027] Furthermore, in step 2, only the lower 256 bits of M, i.e., M L :

[0028] M L =T LL *p L

[0029] Where T LLis the 257-bit T calculated in step 1 L The lower 128 bits of p L is the lower 128-bit part of the input modulus p; let T HL is the 257-bit T calculated in step 1 L The high 129 bits of p H The high 128 bits of the input modulus p satisfy: M = M H *2 256 +M L =(T LH *2 128 +T LL )(p H *2 128 +p L ).

[0030] Furthermore, in step 1 and step 3, the second-order NLP multiplication is used to represent the result Z of the large number multiplication Z=X*Y as (Z H , Z L ) high and low parts, decomposing the 256-bit multiplication into two 128-bit multiplications, and the carry chain path in the large number multiplication calculation delay is truncated by half.

[0031] Further, the high-bit compensation algorithm of the Montgomery modular multiplication algorithm based on the NLP representation method is used to represent the minimum positive number LP in step 4 as C=((T H +Q H )*2 N +T L +Q L ) / 2 N Replace with NLP representation C=T H +Q H +e converts 512-bit large number addition into 256-bit addition, without accumulation (T L +Q L ) to obtain the direction (T H +Q H ), where e is the high-order compensation value used to replace (T L +Q L ) / 2 N The high-order compensation value e is calculated by subtracting the Q calculated in step 3 from L

[255] and T calculated in step 1 L

[255] Perform a logical OR operation, that is, e = Q L

[255] ||T L

[255] , where Q L

[255] is Q L The second highest bit is Q L No. 256; T L

[255] is T L The second highest bit is T L 256th position.

[0032] Furthermore, in the partial product compression summation of large number multiplication in steps 1, 2, and 3, under the condition of two-step parallel second-order NLP multiplication calculation, for Z H and Z L The calculation uses the Karatsuba large number multiplication algorithm with a base of 16 and uses the accumulation operation based on the NLP representation method for calculation. The corresponding calculation process is:

[0033] For integer Z, the NLP representation is expressed as Where b = 2 u is the multiple corresponding to the shift of the partial product, u is the base used for multiplication, and u is 16 using the KO-16 large number multiplication rule, k is the number of partial products, and Z i is the specific value of the ith partial product, Z k-1 b k-1 ,…,Z 1 b,Z 0 Then it forms a complete partial product array; the partial product is divided into B parts, each of which has a bit width of u. Under the NLP representation, the partial accumulation and addition link adopts a carry-save architecture, (c i ,z i ) no longer depends on (c i-1 ,z i-1 ), the calculation process of the addition calculation part of the parallel calculation of carry and sum is as follows:

[0034]

[0035] where c i For each carry output of the parallel calculation, c d ,...,c 1 ,c 2 Splice it into a carry signal for large number multiplication, and z d ,...,z 1 ,z 0 Spliced ​​into the sum signal in large number multiplication, all addition calculations used to generate sum and carry signals can be performed in parallel at the same time.

[0036] Furthermore, the accumulation and optimization of steps 1 and 3 are partially carried out by splitting the accumulation operation of the Montgomery modular multiplication algorithm based on the NLP representation method. The corresponding calculation process is as follows:

[0037] Step a: The result T calculated in step 1 H 、T L Stored in registers;

[0038] After step b and step 2 are completed, T H Keep it in the register and set T L The 256th bit is T L

[255] stored in a register;

[0039] Step c: T H , T L

[255] is compressed together with the partial product generated in step 3 in the compression tree of the multiplication circuit in step 3; since the lower 256 bits of the final result C are all "0", the value of the 257th bit is determined by the high-order compensation value e, so only T H and T L The 256th bit is T L

[255] is fed into the compression tree of step 3;

[0040] Step d: After completing the partial product compression tree, the 256th bit of the sum signal sum and the carry signal carry is judged, and the calculation result of step 4 is generated: C = T H +Q H +e

[0041] The high-performance Montgomery modular multiplication method based on NLP characterization of the present invention has the following advantages:

[0042] The present invention effectively reduces the computational delay of the three multiplication operations. Based on the NLP representation method, a carry compensation mechanism is added in step 4, that is, C=((T H +Q H )*2 N +T L +Q L ) / 2 N , replaced by C=T under NLP representation H +Q H +e, where e is the high-order compensation value of the corresponding carry compensation mechanism, which reduces the calculation delay of the addition operation and reduces the area and power consumption cost of the accumulation implementation. Since the calculation result is represented by NLP, the calculation result is divided into the high-order part and the low-order part. Based on the optimization idea of ​​operand isolation, the carry chain of each large number multiplication and addition calculation in steps 1, 2 and 3 is shortened by half, and nearly half of the data in step 2 does not participate in the calculation, which greatly reduces power consumption.

[0043] Compared with the existing technology, the circuit complexity is reduced, the operating frequency of the analog multiplication circuit is effectively increased, the power consumption and area cost are reduced, and the data throughput of the circuit is not affected. In addition to the advantages in area and power consumption, the propagation delay on the critical path is also greatly reduced, optimizing the performance of the analog multiplier. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 The hardware structure diagram of the NLP-based Montgomery modular multiplication circuit of the present invention;

[0045] Figure 2 It is a schematic diagram of the implementation of the accumulation operation circuit based on LP of the present invention;

[0046] Figure 3 It is a schematic diagram of the implementation of the NLP-based accumulation operation circuit of the present invention;

[0047] FIG4( a ) is a schematic diagram of the NLP representation addition structure of the present invention;

[0048] FIG4( b ) is a truth table of high-bit compensation logic of the present invention;

[0049] FIG4( c ) is a circuit diagram of the high-bit compensation logic of the present invention;

[0050] Figure 5 A schematic diagram of a circuit structure for optimizing the splitting of the NLP algorithm accumulation operation of the present invention. DETAILED DESCRIPTION

[0051] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of a high-performance Montgomery modular multiplication method based on NLP representation of the present invention in conjunction with the accompanying drawings.

[0052] A high-performance Montgomery modular multiplication method based on NLP representation. The modular length of modular multiplication is 256. The input parameters include 256-bit input operands A and B and 256-bit prime field parameters p and invp, where p is the modulus and invp is the value of p in 2 256 The inverse on the field satisfies invp×pmod2 N =1, the output is the modular multiplication result R = (A*B) mod p, such as Figure 1 As shown, it calculates the output result through the following calculation steps:

[0053] Step 1: Input operands A and B, parameters p and invp, and perform NLP multiplication to obtain (T H , T L ) where T H is the 256-bit high-order part of the calculation result; T L The lower 257 bits with a carry signal:

[0054] (T H ,T L ) = NLPmult(A,B)

[0055] Step 2: NLP multiplication calculation NLPmult(TL ,p), we get (M H ,M L ). Does not affect M L The results are accurate, but different from M H The relevant calculations are not enabled, and the final calculation result is actually (0,M L )=NLPmult(T L ,p).

[0056] Step 3: Calculate M L The product of invp and invp is represented by the second-order NLP method.

[0057] (Q H ,Q L )=NLPmult(M L ,invp)

[0058] Where Q H The high 256 bits of Q. L The low-order part of Q with the carry signal is 257 bits:

[0059] Let Q be the M calculated in step 2 L The standard product result with the elliptic curve encryption parameter invp has a bit width of 512 bits. Q is represented by NLP: H The high 256 bits of Q. L The lower 256 bits of Q and a carry signal totaling 257 bits, and satisfying: Q = M L *invp=(Q H *2 256 )+Q L Since the Karatsuba large number multiplication algorithm with a base of 16 is used, 31 partial products will be generated, among which those that do not affect C H and C L

[255] The results are accurate, but different from C L [254:0] All related calculations are disabled. C is the intermediate output value of the Montgomery modular multiplication, which is used to calculate the final output result R. In NLP representation, C = T H +Q H +e, C H The high 256 bits of C. L

[255] is the 256th bit of C, C L [254:0] is the lower 255 bits of C, combined with (T H +Q H )*2 N +T L +Q L The lower 256 bits are all "0", which is consistent with C L[254:0] The related calculation is not enabled in this step, which means that Q is not calculated. L The lower 255 bits of , reduce the number of partial products in the compression tree and thus reduce carry chain propagation.

[0060] Step 4: Q calculated in step 3 L

[255] and T calculated in step 1 L

[255] Perform a logical OR operation and determine whether the result of the logical OR operation is 1. If it is 1, calculate C = T H +Q H +1, if it is 0, calculate C = T H +Q H ; where Q L

[255] is Q L The second highest bit is Q L No. 256; T L

[255] is T L The second highest bit, that is, T L 256th bit; C is the intermediate output value of the Montgomery modular multiplication, which is used to calculate the final output result R, which is C=((T H +Q H )*2 N +T L +Q L ) / 2 N , using ((T H +Q H )*2 N +T L +Q L ) The lower N bits of the expression result are all 0, so replace C with C=T under NLP representation. H +Q H +e, converting 512-bit large number addition into 256-bit addition, without accumulation (T L +Q L ) to obtain the direction (T H +Q H ), where e is a single-bit high-order compensation value of 1 or 0 to replace (T L +Q L ) / 2 N , the value corresponding to e is Q calculated in step 3 L

[255] and T calculated in step 1 L

[255] Perform a logical OR operation; this step is combined with the NLP multiplication in step 3 as a nested operation, that is, the partial product generated in step 3 is compressed at the same time as this step, and the input T of the compressed tree is also only the T represented in LP form. H and T L

[255] . Therefore, the actual Wallace compression circuit type is not 31:2, but 18:2. Finally, the 512-bit sum signal and carry signal obtained by the 18:2 compression circuit compression are obtained, and their lower 255 bits are both "0".

[0061] (carry,sum)=NLPadd(T,NLPmult(M L ,p))

[0062] The carry compensation value obtained from step 4 and step 3. Then, a 256-bit or 257-bit addition calculation is performed to obtain a preliminary Montgomery modular multiplication calculation result C.

[0063] Step 5: Determine whether C calculated in step 4 is greater than p. If it is greater than p, the final output result R=Cp is obtained. If it is not greater than p, the final output result R=C is obtained.

[0064] The 256-bit multiplication calculation in step 1, step 2, and step 3 is replaced by second-order NLP multiplication calculation, and the serial calculation of the traditional large number multiplication Z=X*Y is converted into a two-step parallel calculation. The corresponding two-step parallel second-order NLP multiplication calculation is expressed as:

[0065] (Z H ,Z L )=NLPmult(X,Y)=(X H *2 128 +X L )*(Y H *2 128 +Y L )

[0066] =(X H *Y H )2 256 +[(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 +X L *Y L

[0067] =(X H *Y H )2 256 +PP H +PP L +(X L *Y L )

[0068] =(X H *Y H +PP H ,PP L +X L *Y L )

[0069] Where X and Y are the input 256-bit multipliers, X H The high 128 bits of X. L is the lower 128 bits of X, Y H The high 128 bits of Y. L The lower 128 bits of Y, PP H is [(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 The high 256 bits of PP L is [(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 The lower 256 bits of Z are the output 512-bit large number multiplication results. H The high 256 bits of Z. L The lower 256 bits of Z with a carry signal, totaling 257 bits; satisfying Z = X * Y = (Z H *2 256 )+Z L , two-step parallel second-order NLP multiplication to obtain two output results Z H and Z L It exists independently to characterize the result Z of large number multiplication.

[0070] In step 2, only the lower 256 bits of M, i.e., M, are calculated. L :

[0071] M L =T LL *p L

[0072] Where T LL is the 257-bit T calculated in step 1 L The lower 128 bits of pL is the lower 128-bit part of the input modulus p; let T HL is the 257-bit T calculated in step 1 L The high 129 bits of p H The high 128 bits of the input modulus p satisfy: M = M H *2 256 +M L =(T LH *2 128 +T LL )(p H *2 128 +p L ).

[0073] In step 1 and step 3, the second-order NLP multiplication is used to represent the result Z of the large number multiplication Z=X*Y as (Z H , Z L ) high and low parts, decomposing the 256-bit multiplication into two 128-bit multiplications, and the carry chain path in the large number multiplication calculation delay is truncated by half.

[0074] In step 4, the minimum positive number LP is represented by C = ((T H +Q H )*2 N +T L +Q L ) / 2 N Replace with NLP representation C=T H +Q H +e converts 512-bit large number addition into 256-bit addition, without accumulation (T L +Q L ) to obtain the direction (T H +Q H ), where e is the high-order compensation value used to replace (T L +Q L ) / 2 N The high-order compensation value e is calculated by subtracting the Q calculated in step 3 from L

[255] and T calculated in step 1 L

[255] Perform a logical OR operation, that is, e = Q L

[255] ||T L

[255] , where Q L

[255] is Q L The second highest bit is Q L No. 256; T L

[255] is T L The second highest bit is T L 256th position.

[0075] Through the design of the present invention, the high performance, low area and low energy consumption of the modular multiplication circuit in the actual application scenario can be met. Theoretically, a fast modular multiplication can be completed in 4 clock cycles while ensuring the advantages in area and power consumption.

[0076] like Figure 2 As shown, the calculation process of the accumulation operation based on the NLP representation method for the partial product compression summation of large number multiplication in steps 1, 2 and 3 is:

[0077] (1) For an integer Z, the NLP representation is expressed as Where b = 2 u is the multiple corresponding to the shift of the partial product, u is the base used for multiplication, k is the number of partial products, Z i is the specific value of the ith partial product, z k-1 b k-1 ,…,z 1 b,z 0 Then it forms a complete partial product array.

[0078] (2) Divide the partial product into B parts, each with a bit width of u, and z i Split into z i0 ,z i1 ,z i2 (When B=3)

[0079] (3) Under the NLP representation, some accumulation and addition links adopt the carry-preserve architecture, (c i ,z i ) no longer depends on (c i-1 ,z i-1 ), the calculation process of the addition calculation part of the carry and sum in parallel (without mutual dependence) is as follows:

[0080]

[0081] where c i For each carry output of the parallel calculation, c d ,...,c 1 ,c 2 Spliced ​​into the carry signal in the large number multiplication, z d ,...,z 1 ,z 0 Spliced ​​into the sum signal in large number multiplication, all addition calculations used to generate sum and carry signals can be performed in parallel at the same time.

[0082] like Figure 3 As shown, step 3 and C L[254:0] All related calculations are disabled, and the high-bit compensation algorithm of the Montgomery modular multiplication algorithm based on the NLP representation method is used. The calculation process is as follows:

[0083] (1) The calculation result of C = T + Q is represented in NLP form, where the set c H The high-order part of the 256-bit (T+Q) is included, and the set c L Includes the 256-bit low-order portion and 1 bit c L To c H The carry signal e:

[0084] (c H ,c L )=NLPadd(T,Q)

[0085] (2) Obtaining the high-order compensation value e (as shown in FIG4 , a two-input logic OR gate is required to implement high-order compensation):

[0086] (3) Perform a 256-bit addition operation based on e and calculate the value of the modular multiplication result C:

[0087] C=(T H +Q H +e)

[0088] The present invention adopts addition represented by NLP form, analyzes the characteristic of 512-bit addition calculation (T+Q) with all 256 zeros, and represents C=T+Q as C=(T H +Q H +e), a 256-bit addition with carry is used to implement a low-latency large number addition using the Montgomery modular multiplier based on NLP formal representation proposed in the present invention.

[0089] like Figure 5 As shown, the accumulation and optimization of steps 1 and 3 are partially performed, and the calculation process of the accumulation operation of the Montgomery modular multiplication algorithm based on the NLP representation method is split:

[0090] (1) Substitute the result T calculated in step 1 into H 、T L Stored in register.

[0091] (2) After step 2 is completed, T H Keep it in the register and set T L The 256th bit is T L

[255] Stored in a register.

[0092] (3) T H 、T L

[255] is compressed together with the partial products produced in step 3 in the compression tree of the multiplication circuit in step 3.

[0093] Since the lower 256 bits of the final result C are all "0", the value of the 257th bit is determined by the high-order compensation value e, so only T L The 256th bit of is sent to the hybrid compression tree in step 3.

[0094] (4) After completing the partial product compression tree, the 256th bit of the sum signal (sum) and the carry signal (carry) is judged and a calculation result is generated.

[0095] The present invention adopts the NLP form to characterize the accumulation operation splitting. In the existing 256-bit low-bit and high-bit large number addition operation, the optimization scheme makes full use of this operation and completes the generation of the final calculation result with the help of this step. The entire optimized circuit architecture does not add any additional serial calculation delay except for the simple judgment operation delay.

[0096] The present invention is based on the Montgomery modular multiplication of the NLP representation method and the optimization idea of ​​operand isolation. In each multiplication and addition calculation, nearly half of the data will not participate in the calculation, reducing the calculation delay of the three multiplication operations. A carry compensation mechanism is proposed based on the NLP representation method to reduce the calculation delay of the addition operation. In the Montgomery modular multiplication circuit architecture, the result of the second multiplication calculation is compressed together with the partial product of the third multiplication calculation, reducing the calculation delay of a high-bit width adder.

[0097] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A high-performance Montgomery modular multiplication method based on NLP representation, It is characterized in that The modulus length of modular multiplication is 256, and the input parameters include 256-bit input operands A and B and 256-bit prime field parameters p and invp, where p is the modulus and invp is p in 2 256 The inverse on the field satisfies invp×pmod2 N =1, the output is the modular multiplication result R = (A*B) mod p, and the calculation steps are: Step 1: Calculate the product of input A and B and use the second-order non-minimum positive number NLP method to characterize the calculation result: (T H ,T L )=NLPmult(A,B) Assume that T is the standard product of input operands A and B, that is, T = A*B, with a bit width of 512 bits. T is represented using the second-order NLP method: where T H The high 256 bits of T, T L The lower 256 bits of T also carry a 1-bit carry signal, a total of 257 bits, and satisfy T = A * B = (T H *2 256 )+T L ; Step 2: Calculate T L , p and use the second-order NLP method to represent the calculation result M, and only calculate the lower 256 bits of M L : (M H ,M L )=NLPmult(T L ,p) Let M be the output of step 1, T L The standard product result of the modulus p is 512 bits wide. M is represented by the second-order NLP method: H The high 256 bits of M. L The lower 256 bits of M also carry a carry signal, totaling 257 bits, satisfying M = T L *p=(M H *2 256 )+M L In this step, M H No calculation is performed, that is, M H Constantly 0, only calculate M L ; Step 3: Calculate M L The product of invp and invp is represented by the second-order NLP method. (Q H ,Q L )=NLPmult(M L ,invp) Where Q H The high 256 bits of Q. L The low-order part of Q with the carry signal is 257 bits: Let Q be the M calculated in step 2 L The standard product result with the elliptic curve encryption parameter invp has a bit width of 512 bits. Q is represented by NLP: H The high 256 bits of Q. L The lower 256 bits of Q and a carry signal totaling 257 bits, and satisfying: Q = M L *invp=(Q H *2 256 )+Q L ; Step 4: Q calculated in step 3 L [255] and T calculated in step 1 L [255] Perform a logical OR operation and determine whether the result of the logical OR operation is 1. If it is 1, calculate C = T H +Q H +1, if it is 0, calculate C = T H +Q H ; where Q L [255] is Q L The second highest bit is Q L No. 256; T L [255] is T L The second highest bit, that is, T L 256th bit; C is the intermediate output value of the Montgomery modular multiplication, which is used to calculate the final output result R, which is C=((T H +Q H )*2 N +T L +Q L ) / 2 N , using ((T H +Q H )*2 N +T L +Q L ) The lower N bits of the expression result are all 0, so replace C with C=T under NLP representation. H +Q H +e, converting 512-bit large number addition into 256-bit addition, without accumulation (T L +Q L ) to obtain the direction (T H +Q H ), where e is a single-bit high-order compensation value of 1 or 0 to replace (T L +Q L ) / 2 N , the value corresponding to e is Q calculated in step 3 L [255] and T calculated in step 1 L [255] Perform a logical OR operation; Step 5: Determine whether C calculated in step 4 is greater than p. If it is greater than p, the final output result R=Cp is obtained. If it is not greater than p, the final output result R=C is obtained.

2. The high-performance Montgomery modular multiplication method based on NLP representation according to claim 1, It is characterized in that The 256-bit multiplication calculation in step 1, step 2, and step 3 is replaced by second-order NLP multiplication calculation, and the serial calculation of the traditional large number multiplication Z=X*Y is converted into a two-step parallel calculation. The corresponding two-step parallel second-order NLP multiplication calculation is expressed as: (Z H ,Z L )=NLPmult(X,Y)=(X H *2 128 +X L )*(Y H *2 128 +Y L ) (X) H *AND H )2 256 +[(X H +X L )*(AND H +Y L )-X H *AND H -X L *AND L ]*2 128 +X L *AND L =(X H *Y H )2 256 +PP H +PP L +(X L *Y L ) =(X H *Y H +PP H ,PP L +X L *Y L ) Where X and Y are the input 256-bit multipliers, X H The high 128 bits of X. L is the lower 128 bits of X, Y H The high 128 bits of Y. L The lower 128 bits of Y, PP H is [(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 The high 256 bits of PP L is [(X H +X L )*(Y H +Y L )-X H *Y H -X L *Y L ]*2 128 The lower 256 bits of Z are the output 512-bit large number multiplication results. H The high 256 bits of Z. L The lower 256 bits of Z with a carry signal, totaling 257 bits; satisfying Z = X * Y = (Z H *2 256 )+Z L , two-step parallel second-order NLP multiplication to obtain two output results Z H and Z L It exists independently to characterize the result Z of large number multiplication.

3. The high-performance modular multiplication method based on NLP representation according to claim 1, It is characterized in that In step 2, only the lower 256 bits of M, i.e., M L : M L =T LL *p L Where T LL is the 257-bit T calculated in step 1 L The lower 128 bits of p L is the lower 128-bit part of the input modulus p; let T HL is the 257-bit T calculated in step 1 L The high 129 bits of p H The high 128 bits of the input modulus p satisfy: M = M H *2 256 +M L =(T LH *2 128 +T LL )(p H *2 128 +p L ).

4. The high-performance modular multiplication method based on NLP characterization according to claim 2, It is characterized in that In step 1 and step 3, the second-order NLP multiplication is used to calculate the large number multiplication Z = X * Y result Z is represented as (Z H , Z L ) high and low parts, decomposing the 256-bit multiplication into two 128-bit multiplications, and the carry chain path in the large number multiplication calculation delay is truncated by half.

5. The high-performance modular multiplication method based on NLP representation according to claim 2, It is characterized in that The high-order compensation algorithm of the Montgomery modular multiplication algorithm based on the NLP representation method is used to represent the minimum positive number LP in step 4, C = ((T H +Q H )*2 N +T L +Q L ) / 2 N Replace with NLP representation C=T H +Q H +e converts 512-bit large number addition into 256-bit addition, without accumulation (T L +Q L ) to obtain the direction (T H +Q H ), where e is the high-order compensation value used to replace (T L +Q L ) / 2 N The high-order compensation value e is calculated by subtracting the Q calculated in step 3 from L [255] and T calculated in step 1 L [255] Perform a logical OR operation, that is, e = Q L [255]||T L [255], where Q L [255] is Q L The second highest bit is Q L No. 256; T L [255] is T L The second highest bit is T L 256th position.

6. The high-performance modular multiplication method based on NLP representation according to claim 2, It is characterized in that In the partial product compression summation of large number multiplication in steps 1, 2, and 3, under the condition of two-step parallel second-order NLP multiplication calculation, for Z H and Z L The calculation uses the Karatsuba large number multiplication algorithm with a base of 16 and uses the accumulation operation based on the NLP representation method for calculation. The corresponding calculation process is: For integer Z, the NLP representation is expressed as Where b = 2 u is the multiple corresponding to the shift of the partial product, u is the base used for multiplication, and u is 16 using the KO-16 large number multiplication rule, k is the number of partial products, and Z i is the specific value of the ith partial product, Z k-1 b k-1 ,…,Z 1 b,Z 0 Then it forms a complete partial product array; the partial product is divided into B parts, each of which has a bit width of u. Under the NLP representation, the partial accumulation and addition link adopts a carry-save architecture, (c i ,z i ) no longer depends on (c i-1 ,z i-1 ), the calculation process of the addition calculation part of the parallel calculation of carry and sum is as follows: where c i For each carry output of the parallel calculation, c d ,...,c 1 ,c 2 Splice it into a carry signal for large number multiplication, and z d ,...,z 1 ,z 0 Spliced ​​into the sum signal in large number multiplication, all addition calculations used to generate sum and carry signals can be performed in parallel at the same time.

7. The high-performance modular multiplication method based on NLP representation according to claim 1, It is characterized in that To optimize the accumulation of steps 1 and 3, the accumulation operation of the Montgomery modular multiplication algorithm based on NLP representation is split, and the corresponding calculation process is: Step a: The result T calculated in step 1 H , T L Stored in registers; After step b and step 2 are completed, T H Keep it in the register and set T L The 256th bit is T L [255] stored in a register; Step c: T H , T L [255] is compressed together with the partial product generated in step 3 in the compression tree of the multiplication circuit in step 3; since the lower 256 bits of the final result C are all "0", the value of the 257th bit is determined by the high-order compensation value e, so only T H and T L The 256th bit is T L [255] is fed into the compression tree of step 3; Step d: After completing the partial product compression tree, the 256th bit of the sum signal sum and the carry signal carry is judged, and the calculation result of step 4 is generated: C = T H +Q H +e.

Citation Information

Patent Citations

  • Extensible modular multiplier circuit based on improved Montgomery modular multiplication algorithm

    CN103914277A

  • Montgomery analog multiplication algorithm for VLSI and VLSI structure of intelligenjt card analog multiplier

    CN1392472A