Reconfigurable modular multiplier implementation circuit and method
By designing a reconfigurable model multiplier implementation circuit, the hardware resources and calculation complexity of the model multiplication operation are optimized, and the problem of Barrett model multiplication is solved. The problem of high resource consumption in FPGA design is solved, and more efficient model multiplication operation is achieved.
Patent Information
- Application Number
- CN202510523340.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-05
Smart Images

Figure CN120428952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an algorithm optimization and hardware implementation technology for post-quantum cryptographic operations, and in particular to an algorithm optimization and hardware circuit design for a high-performance, low-overhead reconfigurable modular multiplier for lattice-based post-quantum cryptographic operations. Background Art
[0002] With the rapid development of quantum computers, traditional public-key cryptography algorithms are no longer secure. Post-quantum cryptography, however, can effectively resist attacks from quantum computers, thus providing security for information systems in the future quantum computing era. In 2016, the National Institute of Standards and Technology (NIST) officially launched the standardization work for post-quantum cryptography algorithms. As of August 2023, NIST has released an initial draft of public standards for post-quantum cryptography, including the lattice-based public-key encryption / key encapsulation algorithm Crystals-Kyber and the public-key signature algorithms Crystals-Dilitium and Falcon. Because lattice-based post-quantum cryptography algorithms involve a large number of polynomial multiplication operations on modules, this limits their performance. Therefore, the industry often uses number-theory transforms (NTTs) to accelerate lattice-based post-quantum cryptography algorithms. Whether it is the NTT transformation or the point multiplication operation between two ciphertexts, it is necessary to perform the modular multiplication operation of two numbers. Therefore, the design and optimization of the modular multiplier circuit has received extensive attention and research.
[0003] The modular multiplier is the most core and time-consuming operation component in NTT transformation and polynomial multiplication. In the application scenario of any two variable inputs, the Barrett modular multiplier has a simpler operation structure than the Montgomery modular multiplier and does not require switching between Montgomery domains, so it has higher modular multiplication efficiency. and , modulus , assuming the modulus bit width is , precomputed constants , to calculate , the typical Barrett modular multiplier operation steps are as follows: 1. Calculation and Multiplication of: ; 2. Yes Perform a right shift bits, and get the high-order multiplication result: ; 3. Calculation and Multiplication and right shift bits, and get the high-order multiplication result: ; 4. Calculation and Multiplication and modular reduction , get the low-order multiplication result: ; 5. Yes Modular reduction , get the low-order multiplication result, that is ; 6. Execute conditional subtraction and judge and The size of , then execute Otherwise, execute ; 7. Execute conditional subtraction and judge and The size of , then execute Otherwise, judge and The size of , then execute Otherwise, output directly .
[0004] Although the Barrett modular multiplier can support modular multiplication of any two variable inputs, under certain special conditions, such as when performing NTT and inverse NTT transforms, where one input is a variable and the other is a twiddle factor constant, the computational performance of the Barrett modular multiplier is low. On the other hand, compared to ASIC design, FPGA design offers greater flexibility, shorter development cycles, and lower development costs. However, when building a Barrett modular multiplier circuit on an FPGA, the multiplication operations in the aforementioned algorithms consume significant DSP resources. Summary of the Invention
[0005] The purpose of the present invention is to provide a reconfigurable modular multiplier implementation circuit and method, aiming to solve the problems of large resource overhead in modular multiplication operations and low computing performance when one input is a constant in the prior art.
[0006] The present invention is achieved through the following technical solutions: In a first aspect, the present invention provides a reconfigurable modular multiplier implementation circuit, comprising: a block multiplier, a high-order multiplier, a low-order multiplier, a first conditional subtraction unit, a second conditional subtraction unit, a first data selector, a second data selector, a third data selector, a first register, a second register, a third register, and a fourth register; The block multiplier is used to receive input data A and B and perform block multiplication, output a low-order multiplication result U1 and a high-order multiplication result U2, U1 is directly assigned to the intermediate variable X through the second register and the third register and enters the first conditional subtraction unit through the fourth register, U2 is assigned to the intermediate variable V1 through the second register and the first data selector and enters the high-order multiplier, and the pre-calculated constant T is assigned to the intermediate variable V2 through the first register and the second data selector and enters the high-order multiplier; or, the block multiplier is used to receive input data A and B and perform block multiplication, output a low-order multiplication result U1, U1 is directly assigned to the intermediate variable X through the third register and enters the first conditional subtraction unit through the fourth register, A is assigned to the intermediate variable V1 through the first data selector and enters the high-order multiplier, and the pre-calculated constant BT is assigned to the intermediate variable V2 through the second data selector and enters the high-order multiplier; The high-order multiplier is used to perform high-order multiplication on V1 and V2 to obtain a high-order multiplication result W, which enters the low-order multiplier through the third register; The low-bit multiplier is used to perform low-bit multiplication on W and the modulus q to obtain a low-bit multiplication result Y, which enters the first conditional subtraction unit through the fourth register; The first conditional subtraction unit is used to perform conditional subtraction on X and Y to obtain Z1 and Z2 respectively, and assign Z1 or Z2 to ZS through the selection of the third data selector; The second conditional subtraction unit is used to ZS and Perform conditional subtraction and output the result of modular multiplication.
[0007] Preferably, the block multiplier includes DSP1, DSP2, DSP3, DSP4, a first adder, a second adder, and a third adder; A is split into high bits with the same bit width. and low bit ; B is split into high bits with the same bit width and low bit ; DSP1 is used for receiving and , and execute , get Tp0 and enter the second adder; DSP2 is used for receiving and , and execute , get Tm0 and enter the first adder; DSP3 is used for receiving and , and execute , get Tm1 and enter the first adder; DSP4 is used for receiving and , and execute , get Tm2 and enter the third adder; The first adder is used to perform an addition operation on Tm0 and Tm1, and the obtained output enters the second adder; The second adder is used to perform an addition operation on the output of the first adder and Tp0 to obtain Tp1 and juxtapose Tp0 to obtain U1; The third adder is used to perform an addition operation on Tp1 and Tm2 to obtain Tp2 as U2.
[0008] Preferably, the high-order multiplier includes DSP5, DSP6, DSP7, a fourth adder, a fifth adder, a sixth adder, a first subtractor, a left shift n / 2-bit unit, left shift n Bit unit; split V1 into high bits with the same bit width and low bit ; V2 is split into high bits with the same bit width and low bit ;in, n is the modulus bit width; DSP5 is used for receiving and , and execute , get VV0 and enter the first subtractor; DSP6 is used for receiving and , and execute , get VV2 and enter the first subtractor and enter the left shift n bit unit; The fourth adder is used to receive and , and perform addition operation, the output of which enters DSP7; The fifth adder is used to receive and , and perform addition operation, the output of which enters DSP7; DSP7 is used to perform a multiplication operation on the output of the fourth adder and the output of the fifth adder, and the obtained output enters the first subtractor; The first subtractor is used to perform subtraction operation on the output of DSP7 from VV0 and VV2 to obtain VV1 and enter the left n / 2 bit unit; Shift Leftn / 2 bit unit is used to perform left shift on VV1 n / 2 bits, the obtained output enters the sixth adder; Shift Left n The bit cell is used to perform a left shift on VV2 n The output is fed into the sixth adder; The sixth adder is used to shift the n Output and left shift of / 2 bit unit n The output of the bit unit performs an addition operation to obtain W1, and the high value of W1 is taken. n The bits are output as W.
[0009] Preferably, the low-order multiplier includes DSP8, DSP9, DSP 10 , the seventh adder, the eighth adder; split W into high bits with the same bit width and low bit ;q is split into high bits with the same bit width and low bit ; DSP8 is used for receiving and , and execute , get Fp0 and enter the eighth adder; DSP9 is used for receiving and , and execute , get Tn0 and enter the seventh adder; DSP 10 For receiving and , and execute , get Tn1 and enter the seventh adder; The seventh adder is used to perform an addition operation on Tn0 and Tn1, and the obtained output enters the eighth adder; The eighth adder is used to perform an addition operation on the output of the seventh adder and Fp0 to obtain Fp1 and concatenate it with Fp0 to obtain Y.
[0010] Preferably, the first conditional subtraction unit includes a second subtractor, a ninth adder, a tenth adder, an eleventh adder, a fourth data selector, and a fifth data selector; The second subtractor is used to receive X and Y and perform a subtraction operation to obtain ZZ and enter the ninth adder, the tenth adder, the eleventh adder, the fourth data selector and the fifth data selector; The ninth adder is used for performing an addition operation on ZZ and q to obtain FF and enter the result into the fourth data selector; The tenth adder is used to add ZZ and the constant 2 n+1Performing an addition operation, and obtaining an output that enters a fourth data selector and a fifth data selector; an eleventh adder for performing an addition operation on ZZ and q, and the obtained output enters the fourth data selector; The fourth data selector is used to determine the value of ZZ and FF concatenated into {ZZ, FF}. If {ZZ, FF} = {00}, ZZ is output as Z1. If {ZZ, FF} = {01}, the output of the tenth adder is output as Z1. If {ZZ, FF} = {10}, the output of the eleventh adder is output as Z1. The fifth data selector is used to determine whether ZZ is less than 0. If ZZ is equal to 0, ZZ is directly output as Z2; otherwise, the output of the tenth adder is output as Z2.
[0011] Preferably, the second conditional subtraction unit includes a 1-bit left shift unit, a third subtractor, a fourth subtractor, and a sixth data selector; The left shift 1-bit unit is used to shift q left by 1 bit to obtain 2q and enter the third subtractor; The third subtractor is used for performing a subtraction operation on ZS and 2q to obtain Zq2 and input the result into the sixth data selector; The fourth subtractor is used to perform a subtraction operation on ZS and q to obtain Zq1 and enter the sixth data selector; The sixth data selector is used to judge and set the value of {Zq2, Zq1}. If {Zq2, Zq1}={00}, Zq2 is output as Z; if {Zq2, Zq1}={10}, Zq1 is output as Z; if {Zq2, Zq1}={11}, ZS is directly output as Z.
[0012] Preferably, the bit width of the input data A and B and the modulus q is 32 bits.
[0013] Preferably, the reconfigurable modular multiplier implementation circuit is built on an FPGA chip.
[0014] In a second aspect, the present invention provides a reconfigurable modular multiplier, comprising the reconfigurable modular multiplier implementation circuit as described above.
[0015] In a third aspect, the present invention provides a reconfigurable modular multiplication method, characterized by comprising: Step 1: Calculate input data and Multiplication and modular reduction , get the low-order multiplication result: ; Step 2: Determine the data selection signal Is it 0? If it is 0, then Shift right bits, and get the high-order multiplication result: , and precompute constants Assign to Otherwise, just Assign to , precomputed constants Assign to ; Step 3: Calculation and Multiplication and right shift bits, and get the high-order multiplication result: ; Step 4: Calculation and Multiplication and modular reduction , get the low-order multiplication result: ; Step 5: Perform conditional subtraction and judge and The size of , then execute , ; Further judgment and The size of , then execute Otherwise, execute ; Step 6: Determine the data selection signal Is it 0? If it is 0, then Assign to Otherwise, Assign to ; Step 7: Perform conditional subtraction and judge and The size of , then execute Otherwise, judge and The size of , then execute Otherwise, output directly As .
[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention designs a reconfigurable modular multiplier circuit structure that can support two variable inputs and one variable and one constant input at the same time. When the reconfigurable modular multiplier circuit is configured as one variable and one constant input, it is not necessary to execute the input data. and The high-bit multiplication is performed, so the reconfigurable modular multiplier data path does not pass through the first register and the second register. Compared with the typical Barrett modular multiplier that occupies a full three-level register delay, the reconfigurable modular multiplier implementation circuit only needs to occupy two levels of register delay, which means that the one-level register delay is reduced, and the performance is improved by 33.2%. Therefore, it has higher computing efficiency. The present invention can be widely used in hardware acceleration of various post-quantum cryptographic algorithms, especially for providing efficient modular multiplication operation support for lattice-based post-quantum cryptographic algorithms. By constructing an efficient lattice-based public key encryption / key encapsulation algorithm and a public key signature algorithm, the present invention can further provide a high-performance and area-optimized hardware platform support for the migration of traditional public key cryptographic algorithms to post-quantum cryptographic algorithms.
[0017] Furthermore, the block multiplier of the present invention is designed based on the block multiplication concept, the high-order multiplier is designed based on the Karatsuba multiplication concept, and the low-order multiplier is also implemented using the block multiplication concept; when the reconfigurable modular multiplier implementation circuit is configured for two variable inputs, compared to the typical Barrett modular multiplier that requires 12 DSP resources (integer multiplication, high-order multiplication, and low-order multiplication each occupy 4 DSPs), the reconfigurable modular multiplier implementation circuit of the present invention only requires 10 DSP resources (block multiplication, high-order multiplication, and low-order multiplication each occupy 4, 3, and 3 DSPs, respectively), reducing key hardware resources by 16.7%.
[0018] The present invention proposes a reconfigurable modular multiplication method that can support two variable inputs and one variable and one constant input at the same time. When performing a modular multiplication operation of a variable and a constant input, the algorithm does not need to perform input and The high-order multiplication of , therefore, has lower computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 1 is a schematic diagram of a circuit structure of a reconfigurable modular multiplier according to an embodiment of the present invention; Figure 2 1 is a circuit diagram of a block multiplier according to an embodiment of the present invention; Figure 3 1 is a circuit diagram of a high-bit multiplier in an embodiment of the present invention; Figure 41 is a circuit diagram of a low-bit multiplier in an embodiment of the present invention; Figure 5 is a circuit structure diagram of a first conditional subtraction unit in an embodiment of the present invention; Figure 6 2 is a circuit structure diagram of the second conditional subtraction unit in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following describes the embodiments of the present invention through specific examples. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention.
[0022] It should be noted that the process equipment or devices not specifically specified in the following embodiments are all conventional equipment or devices in the art.
[0023] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. Furthermore, unless otherwise specified, the numbering of each method step is merely a convenient tool for identifying each method step, and is not intended to limit the order of arrangement of each method step or to define the scope of the invention. Changes or adjustments to their relative relationships, without substantially changing the technical content, should also be considered within the scope of the invention.
[0024] In addition, it should be noted that the terms "first," "second," etc., used in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.
[0025] The embodiments of the present invention address the problems of large hardware resource overhead and low computational performance of the existing Barrett modular multiplier when one input is a constant. A reconfigurable modular multiplier implementation circuit with even lower hardware resource overhead is proposed. At the same time, when one input is a constant, the algorithm complexity and modular multiplication circuit delay can be further reduced by configuring signal selection and circuit bypass.
[0026] The embodiment of the present invention provides a reconfigurable modular multiplication method, which converts the input and Integer multiplication is split into low-order multiplication and high-order multiplication Two parts, when the data selection signal When it is 0, it means that the reconfigurable modular multiplier performs the modular multiplication operation of two variable inputs, and the reconfigurable modular multiplier performs normally. and Low-order multiplication and high-order multiplication; when the data selection signal When it is 1, it means that the reconfigurable modular multiplier only performs the modular multiplication operation of a variable and a constant input. and The low-order multiplication of the constant is performed without performing the high-order multiplication, thereby reducing the computational complexity of the modular multiplication operation of a variable and a constant input.
[0027] For input data and , data selection signal , modulus , assuming the modulus bit width is , precomputed constants , , to calculate The reconfigurable modular multiplication method comprises the following specific steps: Step 1: Calculate input data and Multiplication and modular reduction , get the low-order multiplication result: ; Step 2: Determine the data selection signal Is it 0? If it is 0, it means to perform a modular multiplication operation of two variable inputs, then the take The low-order multiplication result is shifted right bits, and get the high-order multiplication result: , and Assign to Otherwise, it means to perform a modular multiplication operation of a variable and a constant input, then directly Assign to ,Will Assign to ; Step 3: Calculation and Multiplication and right shift bits, and get the high-order multiplication result: ; Step 4: Calculation and Multiplication and modular reduction , get the low-order multiplication result: ; Step 5: Perform conditional subtraction and judge and The size of , then execute , ; Further judgment and The size of , then execute Otherwise, execute ; Step 6: Determine the data selection signal Is it 0? If it is 0, it means to perform a modular multiplication operation of two variable inputs. Assign to Otherwise, it means to perform a modular multiplication operation of a variable and a constant input, then Assign to ; Step 7: Perform conditional subtraction and judge and The size of , then execute Otherwise, judge and The size of , then execute Otherwise, output directly As .
[0028] like Figure 1 As shown, an embodiment of the present invention provides a reconfigurable modular multiplier implementation circuit, comprising a block multiplier, a high-order multiplier, a low-order multiplier, a first conditional subtraction unit, a second conditional subtraction unit, a first data selector, a second data selector, a third data selector, a first register, a second register, a third register, and a fourth register. The reconfigurable modular multiplier of the present invention, including the reconfigurable modular multiplier implementation circuit, utilizes a common set of hardware circuits to support modular multiplication operations with two variable inputs, and can also support modular multiplication operations with one variable and one constant input with a shorter delay path. Furthermore, when implemented on an FPGA chip, it reduces DSP hardware resource overhead.
[0029] The following detailed description assumes that the bit width of input data, output data, and modulus q is 32 bits and each DSP is a DSP48E1 chip.
[0030] When the reconfigurable modular multiplier performs modular multiplication of two variable inputs, the data selection signal sel is configured to 0. At this time, the input data A[31:0] and B[31:0] first enter the block multiplier to perform block multiplication, and obtain the low-order multiplication result U1[31:0] and the high-order multiplication result U2[31:0] respectively, where U1[31:0] is directly assigned to the intermediate variable X[31:0] through the second register and the third register, and U2[31:0] is assigned to the intermediate variable V1[31:0] through the second register and the first data selector, and the pre-calculated constant T[31:0] is assigned to the intermediate variable V2[31:0] through the first register and the second data selector. Subsequently, V1[31:0] and V2[31:0] enter the high-order multiplier to obtain the high-order multiplication result W[31:0], which is then passed to the low-order multiplier through the third register and subjected to low-order multiplication with the modulus q[31:0] to obtain the low-order multiplication result Y[31:0]. Next, X[31:0] and Y[31:0] enter the first conditional subtraction unit through the fourth register to perform the conditional subtraction in step 5 above, and obtain Z1[31:0] and Z2[31:0] respectively. Through the selection of the third data selector, Z2[31:0] is assigned to ZS[31:0] and continues to enter the second conditional subtraction unit to perform the conditional subtraction in step 7 above, thus finally obtaining the modular multiplication result of the two variable inputs.
[0031] When the reconfigurable modular multiplier performs a modular multiplication operation on a variable and a constant input, the data selection signal sel is configured to 1. At this time, the input data A[31:0] and B[31:0] first enter the block multiplier to perform block multiplication. Only the low-order multiplication result U1[31:0] needs to be obtained, and U1[31:0] is directly assigned to the intermediate variable X[31:0] through the third register. The input data A[31:0] and the pre-calculated constant BT[31:0] are assigned to the intermediate variables V1[31:0] and V2[31:0] through the first data selector and the second data selector respectively. Subsequently, V1[31:0] and V2[31:0] enter the high-order multiplier to obtain the high-order multiplication result W[31:0]. This is passed to the low-order multiplier through the third register and performs low-order multiplication with the modulus q[31:0] to obtain the low-order multiplication result Y[31:0]. Next, X[31:0] and Y[31:0] enter the first conditional subtraction unit through the fourth register to perform the conditional subtraction in step 5 above, and obtain Z1[31:0] and Z2[31:0] respectively. Through the selection of the third data selector, Z1[31:0] will be assigned to ZS[31:0] and continue to enter the second conditional subtraction unit to perform the conditional subtraction in step 7 above, thereby finally obtaining the modular multiplication result of one variable and one constant input. Compared to performing a modular multiplication operation with two variable inputs, in the current case, the reconfigurable modular multiplier data path does not pass through the first and second registers, which reduces the delay of one register, thereby achieving higher computational efficiency.
[0032] In one embodiment, if Figure 2 As shown, the block multiplier includes DSP1, DSP2, DSP3, DSP4, a first adder, a second adder, and a third adder. In order to enable the reconfigurable modular multiplier to flexibly support modular multiplication of two variable inputs and modular multiplication of one variable input and one constant input, while meeting the bit width constraint of the DSP48E1 chip in the FPGA (i.e., 25×18 bits), the integer multiplication operation of input data A[31:0] and B[31:0] is implemented using the block multiplication idea. Specifically, A[31:0] and B[31:0] are split into the upper 16 bits ( and , i.e. A[31:16] and B[31:16]) and the lower 16 bits ( and , that is, A[15:0] and B[15:0]), then the block multiplication is as shown in formula (1): (1) When implementing the hardware circuit, and Can be taken directly and The input data A[15:0] and B[15:0] first enter DSP1 for execution. Get Tp0[31:0], A[31:16] and B[15:0] and enter DSP2 for execution Get Tm0[31:0], A[15:0] and B[31:16] and enter DSP3 for execution Get Tm1[31:0], A[31:16] and B[31:16] and enter DSP4 for execution Tm2[31:0] is obtained. Subsequently, Tm0[31:0] and Tm1[31:0] enter the first adder for addition, and the result is further added to Tp0[31:16] and enter the second adder for addition to obtain Tp1[31:0]. Next, the concatenation of Tp1[15:0] and Tp0[15:0] forms the low-order multiplication result of A[31:0] and B[31:0], that is, U1[31:0] = {Tp1[15:0], Tp0[15:0]}. Finally, Tp1[31:16] and Tm2[31:0] enter the third adder for addition to obtain Tp2[31:0], and the high-order multiplication result of A[31:0] and B[31:0] is obtained, that is, U2[31:0] = Tp2[31:0].
[0033] In one embodiment, if Figure 3 As shown, the high-bit multiplier includes DSP5, DSP6, DSP7, a fourth adder, a fifth adder, a sixth adder, a first subtractor, a 16-bit left shift unit, and a 32-bit left shift unit. In order to reduce the DSP resource occupation of the reconfigurable modular multiplier and meet the bit width constraints of the DSP48E1 chip, the high-bit multiplication operation of the input data V1[31:0] and V2[31:0] is implemented using the Karatsuba multiplication idea. Specifically, V1[31:0] and V2[31:0] are split into the high 16 bits ( and , namely V1[31:16] and V2[31:16]) and the lower 16 bits ( and , that is, V1[15:0] and V2[15:0]), then the Karatsuba multiplication is shown in formula (2): (2) Input data V1[15:0] and V2[15:0] first enter DSP5 for execution Get VV0[31:0], V1[31:16] and V2[31:16] and enter DSP6 for execution Get VV2[31:0], and V1[31:16] and V1[15:0] enter the fourth adder to perform addition operation, V2[31:16] and V2[15:0] enter the fifth adder to perform addition operation, and their addition operation results are sent to DSP7 for execution The multiplication result is further sent to the first subtractor for subtraction with VV0[31:0] and VV2[31:0] to obtain VV1[31:0]. Subsequently, VV1[31:0] enters the 16-bit left shift unit for 16-bit left shift, and VV2[31:0] enters the 32-bit left shift unit for 32-bit left shift. The respective results and VV0[31:0] enter the sixth adder for addition, obtaining W1[63:0], which is the integer multiplication result of V1[31:0] and V2[31:0]. Finally, the high 32 bits of W1[63:0] (i.e., W1[63:0] is right-shifted 32 bits) are taken as W[31:0] output.
[0034] In one embodiment, if Figure 4 As shown, the low-order multiplier includes DSP8, DSP9, DSP 10 , the seventh adder, and the eighth adder. Similar to the block multiplier, the low-order multiplier is also implemented using the idea of block multiplication. The difference is that the low-order multiplier only needs the low-order multiplication result, so it does not need to perform the last step in formula (1). Specifically, W[31:0] and q[31:0] are split into the upper 16 bits ( and , namely W[31:16] and q[31:16]) and the lower 16 bits ( and , that is, W[15:0] and q[15:0]), then the block multiplication is as shown in formula (3): (3) Specifically, the input data W[15:0] and q[15:0] first enter the DSP8 for execution Get Fp0[31:0], W[31:16] and q[15:0] to enter DSP9 for execution Get Tn0[31:0], W[15:0] and q[31:16] into DSP 10 implement Tn1[31:0] is obtained. Subsequently, Tn0[31:0] and Tn1[31:0] enter the seventh adder for addition. The result is further added to Fp0[31:16] and enters the eighth adder for addition to obtain Fp1[31:0]. Next, the concatenation of Fp1[15:0] and Fp0[15:0] forms the low-order multiplication result of W[31:0] and q[31:0], that is, Y[31:0] = {Fp1[15:0], Fp0[15:0]}.
[0035] In one embodiment, if Figure 5 As shown, the first conditional subtraction unit includes a second subtractor, a ninth adder, a tenth adder, an eleventh adder, a fourth data selector, and a fifth data selector. In order to reduce the data path delay of the conditional subtraction, the first conditional subtraction unit is implemented using a data selector structure. Specifically, the input data X[31:0] and Y[31:0] first enter the second subtractor to perform a subtraction operation to obtain ZZ[31:0], and then continue to enter the ninth adder to perform an addition operation with q[31:0] to obtain FF[31:0]. Subsequently, ZZ[31:0] and the constant 2 n+1 Enter the tenth adder to perform the addition operation, and at the same time, ZZ[31:0] and q[31:0] enter the eleventh adder to perform the addition operation. Next, determine whether ZZ[31:0] and FF[31:0] are less than 0. Since the highest bits of ZZ[31:0] and FF[31:0] are sign bits, we can directly determine the sign bits of the two and set the value of {ZZ
[31] ,FF
[31] }. If {ZZ
[31] ,FF
[31] }={00}, the fourth data selector outputs ZZ[31:0] as Z1[31:0]; if {ZZ
[31] ,FF
[31] }={01}, the fourth data selector outputs the value of the tenth adder as Z1[31:0]; if {ZZ
[31] ,FF
[31] }={10}, the fourth data selector outputs the value of the eleventh adder as Z1[31:0]. Similarly, determine whether ZZ[31:0] is less than 0, that is, determine the value of the sign bit ZZ
[31] . If it is equal to 0 (indicating that ZZ[31:0] is a value greater than or equal to 0), the fifth data selector directly outputs ZZ[31:0] as Z2[31:0]. Otherwise (indicating that ZZ[31:0] is a value less than 0), the value of the tenth adder is output as Z2[31:0].
[0036] In one embodiment, if Figure 6As shown, the second conditional subtraction unit includes a left shift 1-bit unit, a third subtractor, a fourth subtractor, and a sixth data selector. The second conditional subtraction unit is also implemented using a data selector structure to reduce the conditional subtraction data path delay. Specifically, the input data q[31:0] first enters the left shift 1-bit unit and is shifted left by 1 bit to obtain 2 times q. Subsequently, ZS[31:0] and 2q enter the third subtractor to perform a subtraction operation to obtain Zq2[31:0]. ZS[31:0] and q[31:0] enter the fourth subtractor to perform a subtraction operation to obtain Zq1[31:0]. Next, the values of {Zq2
[31] , Zq1
[31] } are judged and set. If {Zq2
[31] , Zq1[3 1]}={00}, the sixth data selector outputs the value of the third subtractor as Z[31:0]; if {Zq2
[31] ,Zq1
[31] }={10}, the sixth data selector outputs the value of the fourth subtractor as Z[31:0]; if {Zq2
[31] ,Zq1
[31] }={11}, the sixth data selector directly outputs ZS[31:0] as Z[31:0], that is, the final result of A[31:0]×B[31:0]mod q[31:0].
[0037] In combination with specific post-quantum cryptographic algorithm parameters, although the bit width of the input and output data and modulus q used in the embodiments of the present invention is 32 bits, the reconfigurable modular multiplier implementation circuit, block multiplier circuit, high-bit multiplier circuit and low-bit multiplier circuit disclosed in the present invention also support the expansion of larger coefficient bit width and modulus bit width.
Claims
1. A reconfigurable modular multiplier implementation circuit, characterized in that: include: A block multiplier, a high-order multiplier, a low-order multiplier, a first conditional subtraction unit, a second conditional subtraction unit, a first data selector, a second data selector, a third data selector, a first register, a second register, a third register, and a fourth register; The block multiplier is used to receive input data A and B and perform block multiplication, output a low-order multiplication result U1 and a high-order multiplication result U2, U1 is directly assigned to the intermediate variable X through the second register and the third register and enters the first conditional subtraction unit through the fourth register, U2 is assigned to the intermediate variable V1 through the second register and the first data selector and enters the high-order multiplier, and the pre-calculated constant T is assigned to the intermediate variable V2 through the first register and the second data selector and enters the high-order multiplier; or, the block multiplier is used to receive input data A and B and perform block multiplication, output a low-order multiplication result U1, U1 is directly assigned to the intermediate variable X through the third register and enters the first conditional subtraction unit through the fourth register, A is assigned to the intermediate variable V1 through the first data selector and enters the high-order multiplier, and the pre-calculated constant BT is assigned to the intermediate variable V2 through the second data selector and enters the high-order multiplier; The high-order multiplier is used to perform high-order multiplication on V1 and V2 to obtain a high-order multiplication result W, which enters the low-order multiplier through the third register; The low-bit multiplier is used to perform low-bit multiplication on W and the modulus q to obtain a low-bit multiplication result Y, which enters the first conditional subtraction unit through the fourth register; The first conditional subtraction unit is used to perform conditional subtraction on X and Y to obtain Z1 and Z2 respectively, and assign Z1 or Z2 to ZS through the selection of the third data selector; The second conditional subtraction unit is used to ZS and Perform conditional subtraction and output the result of modular multiplication.
2. The reconfigurable modular multiplier implementation circuit according to claim 1, characterized in that: The block multiplier includes DSP1, DSP2, DSP3, DSP4, a first adder, a second adder, and a third adder; A is split into high bits with the same bit width and low bit ; B is split into high bits with the same bit width and low bit ; DSP1 is used for receiving and , and execute , get Tp0 and enter the second adder; DSP2 is used for receiving and , and execute , get Tm0 and enter the first adder; DSP3 is used for receiving and , and execute , get Tm1 and enter the first adder; DSP4 is used for receiving and , and execute , get Tm2 and enter the third adder; The first adder is used to perform an addition operation on Tm0 and Tm1, and the obtained output enters the second adder; The second adder is used to perform an addition operation on the output of the first adder and Tp0 to obtain Tp1 and juxtapose Tp0 to obtain U1; The third adder is used to perform an addition operation on Tp1 and Tm2 to obtain Tp2 as U2.
3. The reconfigurable modular multiplier implementation circuit according to claim 1, characterized in that: The high-order multiplier includes DSP5, DSP6, DSP7, a fourth adder, a fifth adder, a sixth adder, a first subtractor, a left shift n / 2 bit unit, left shift n Bit unit; split V1 into high bits with the same bit width and low bit ; V2 is split into high bits with the same bit width and low bit ;in, n is the modulus bit width; DSP5 is used for receiving and , and execute , get VV0 and enter the first subtractor; DSP6 is used for receiving and , and execute , get VV2 and enter the first subtractor and enter the left shift n bit unit; The fourth adder is used to receive and , and perform addition operation, the output of which enters DSP7; The fifth adder is used to receive and , and perform addition operation, the output of which enters DSP7; DSP7 is used to perform a multiplication operation on the output of the fourth adder and the output of the fifth adder, and the obtained output enters the first subtractor; The first subtractor is used to perform subtraction operation on the output of DSP7 from VV0 and VV2 to obtain VV1 and enter the left n / 2 bit unit; Shift Left n / 2 bit unit is used to perform left shift on VV1 n / 2 bits, the obtained output enters the sixth adder; Shift Left n The bit cell is used to perform a left shift on VV2 n The output is fed into the sixth adder; The sixth adder is used to shift the n Output and left shift of / 2 bit unit n The output of the bit unit performs an addition operation to obtain W1, and the high value of W1 is taken. n The bits are output as W.
4. The reconfigurable modular multiplier implementation circuit according to claim 1, wherein: The low-order multiplier includes DSP8, DSP9, DSP 10 , the seventh adder, the eighth adder; split W into high bits with the same bit width and low bit ; q is split into high bits with the same bit width and low bit ; DSP8 is used for receiving and , and execute , get Fp0 and enter the eighth adder; DSP9 is used for receiving and , and execute , get Tn0 and enter the seventh adder; DSP 10 For receiving and , and execute , get Tn1 and enter the seventh adder; The seventh adder is used to perform an addition operation on Tn0 and Tn1, and the obtained output enters the eighth adder; The eighth adder is used to perform an addition operation on the output of the seventh adder and Fp0 to obtain Fp1 and concatenate it with Fp0 to obtain Y.
5. The reconfigurable modular multiplier implementation circuit according to claim 1, wherein: The first conditional subtraction unit includes a second subtractor, a ninth adder, a tenth adder, an eleventh adder, a fourth data selector, and a fifth data selector; The second subtractor is used to receive X and Y and perform a subtraction operation to obtain ZZ and enter the ninth adder, the tenth adder, the eleventh adder, the fourth data selector and the fifth data selector; The ninth adder is used for performing an addition operation on ZZ and q to obtain FF and enter the result into the fourth data selector; The tenth adder is used to add ZZ and the constant 2 n+1 Performing an addition operation, and the obtained output enters the fourth data selector and the fifth data selector; an eleventh adder for performing an addition operation on ZZ and q, and the obtained output enters the fourth data selector; The fourth data selector is used to determine the value of ZZ and FF concatenated into {ZZ, FF}. If {ZZ, FF} = {00}, ZZ is output as Z1. If {ZZ, FF} = {01}, the output of the tenth adder is output as Z1. If {ZZ, FF} = {10}, the output of the eleventh adder is output as Z1. The fifth data selector is used to determine whether ZZ is less than 0. If ZZ is equal to 0, ZZ is directly output as Z2; otherwise, the output of the tenth adder is output as Z2.
6. The reconfigurable modular multiplier implementation circuit according to claim 1, characterized in that: The second conditional subtraction unit includes a left shift 1-bit unit, a third subtractor, a fourth subtractor, and a sixth data selector; The left shift 1-bit unit is used to shift q left by 1 bit to obtain 2q and enter the third subtractor; The third subtractor is used for performing a subtraction operation on ZS and 2q to obtain Zq2 and input the result into the sixth data selector; The fourth subtractor is used to perform a subtraction operation on ZS and q to obtain Zq1 and enter the sixth data selector; The sixth data selector is used to judge and set the value of {Zq2, Zq1}. If {Zq2, Zq1}={00}, Zq2 is output as Z; if {Zq2, Zq1}={10}, Zq1 is output as Z; if {Zq2, Zq1}={11}, ZS is directly output as Z.
7. The reconfigurable modular multiplier implementation circuit according to claim 1, characterized in that: The bit width of the input data A and B and the modulus q is 32 bits.
8. The reconfigurable modular multiplier implementation circuit according to claim 1, characterized in that: The reconfigurable modular multiplier implementation circuit is built on an FPGA chip.
9. A reconfigurable modular multiplier, characterized in that: The invention comprises a reconfigurable modular multiplier implementation circuit according to any one of claims 1 to 8.
10. A reconfigurable modular multiplication method, characterized in that: include: Step 1: Calculate input data and Multiplication and modular reduction , get the low-order multiplication result: ; Step 2: Determine the data selection signal Is it 0? If it is 0, then Shift right bits, and get the high-order multiplication result: , and precompute constants Assign to ; Otherwise, just Assign to , precomputed constants Assign to ; Step 3: Calculation and Multiplication and right shift bits, and get the high-order multiplication result: ; Step 4: Calculation and Multiplication and modular reduction , get the low-order multiplication result: ; Step 5: Perform conditional subtraction and judge and The size of , then execute , ; Further judgment and The size of , then execute Otherwise, execute ; Step 6: Determine the data selection signal Is it 0? If it is 0, then Assign to Otherwise, Assign to ; Step 7: Perform conditional subtraction and judge and The size of , then execute Otherwise, judge and The size of , then execute Otherwise, output directly As .