A lattice cryptography modular multiplication method and multiplier based on an improved Plantard algorithm
The improved plantard algorithm preprocesses the rotation factor and four-stage pipeline design, optimizes the modular multiplication process, solves the calculation efficiency and resource consumption problems of the modular subtraction algorithm in the quantum computing environment, and realizes efficient modular multiplication operation.
Patent Information
- Application Number
- CN202510497007.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing modulus reduction algorithms have low computational efficiency, high resource consumption, and limited parallelism in quantum computing environments, especially when dealing with special prime modulus, there are problems with high computing overhead.
The improved plantard algorithm is used to pre-process and optimize the rotation factor, decompose the modular multiplication process into three-stage calculations, and a four-stage pipelined modular multiplication is designed, and the shift-addition substitution multiplication and dynamic bit-width truncation technology is used to optimize the utilization of hardware resources.
The throughput and energy efficiency ratio of the grid cryptographic algorithm NTT/INTT operations has been significantly improved, the logical resources have been reduced by 50.0%-60.9%, the operating frequency has been increased to 350-400 MHz, and the single-mode multiplication only takes 5 cycles, and the computing speed has been increased by 2.5-3 times.
Smart Images

Figure CN120050042B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hardware security technologies, and particularly to a lattice cryptography modular multiplication method and a modular multiplier based on an improved Plantard algorithm. Background Art
[0002] In the context of the rapid development of quantum computing technology, modern cryptography is facing an unprecedented paradigm shift. The problems of large integer factorization and elliptic curve discrete logarithm on which the currently widely used public key cryptosystems such as RSA and ECC rely are polynomially solvable when a quantum computer executes the Shor algorithm. This breakthrough has fundamentally shaken the security foundation of traditional cryptosystems.
[0003] The core operation efficiency of lattice basis algorithms directly depends on the implementation efficiency of the number-theoretic transform (NTT) on its polynomial ring. Polynomial multiplication occupies nearly 50% of the running time of lattice basis algorithm implementation. As the core acceleration algorithms for polynomial multiplication, NTT and INTT (Inverse Number Theoretic Transform, INTT), the modular multiplication operation occupies the main computational proportion and has a significant impact on the overall working frequency and the number of computational cycles. Therefore, optimizing the modular multiplication operation is crucial for improving the performance of NTT implementation. This operation consists of two key stages: integer multiplication and modular reduction. Among them, the modular reduction process needs to map the high-order product result to the finite field space, and its computational intensity accounts for 60%-80% of the modular multiplication operation. In scenarios of resource-constrained embedded devices and high-throughput servers, the optimization degree of the modular reduction algorithm directly determines key performance indicators such as signature generation / verification latency and system energy consumption.
[0004] However, existing modular reduction algorithms (such as Barrett, Montgomery, etc.) have problems such as low computational efficiency, high resource consumption, and limited parallelism in the NTT calculation of post-quantum lattice cryptography, especially the defect of high computational overhead when dealing with special prime moduli (such as q = 8380417). Summary of the Invention
[0005] The purpose of the present invention is to provide a lattice cryptography modular multiplication method and a modular multiplier based on an improved Plantard algorithm, which can solve the defects in the prior art such as long critical path of modular multiplication operation, redundant intermediate data bit width, and large number of computational cycles, and significantly improve the throughput rate and energy efficiency ratio of the NTT / INTT operation of the lattice cryptography algorithm.
[0006] To achieve the above technical purpose, the technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention discloses a lattice cryptography modular multiplication method based on an improved Plantard algorithm, and the method includes:
[0008] Preprocess and optimize the input rotation factor, and preprocess the rotation factor into a 26-bit rotation factor through the following formula : :
[0009] ;
[0010] In the formula, is the rotation factor in the NTT calculation of the lattice cryptography algorithm; is the modulus corresponding to the lattice cryptography algorithm; is the bit width corresponding to the modulus of the lattice cryptography algorithm, , ;
[0011] Decompose the modular multiplication process into three-level calculations. Among them, calculate in the first level. In the formula, represents the 12-bit number to be multiplied modulo. Correct the truncation error in the second level to get , and finally obtain the 12-bit modular multiplication result in the third level; at the same time, detect the boundary condition of . If detected, output 0 through the multiplexer.
[0012] In a second aspect, the present invention discloses a lattice cryptography modular multiplier based on an improved Plantard algorithm. The lattice cryptography modular multiplier includes a two-input multiplier, four two-input adders, and five registers, which together form a four-stage pipeline;
[0013] Among them, in the first-stage pipeline, the two input ports of the two-input multiplier are respectively used to receive the 12-bit number to be multiplied modulo and the preprocessed 26-bit rotation factor . The calculated 39-bit multiplication result is stored in the first register;
[0014] In the second-stage pipeline, the first register shifts the multiplication result right by 13 bits and then sends it to the first adder. After adding the compensation constant 1 by the first adder, the result is obtained, and the 13-bit result is stored in the second register;
[0015] In the third - stage pipeline, the two input ports of the second adder are respectively used to receive the result of shifting the lower 13 bits of the second register to the right by 3 bits and the result of shifting the lower 13 bits to the right by 2 bits. After addition operation processing, the result is stored in the third register; the two input ports of the third adder are respectively used to receive the result of the last 8 bits of the lower 13 - bit stage of the second register and the lower 13 bits. After addition operation processing, the result is stored in the fourth register;
[0016] In the fourth - stage pipeline, the two input ends of the fourth adder are respectively used to receive the calculation results of the third register and the fourth register. After addition operation processing, the result is stored in the fifth register. The fifth register performs a right - shift operation on the calculation result by 5 bits to output a 12 - bit modular multiplication result , , and at the same time, the boundary condition is detected through a hardware comparator and 0 is output through a multiplexer.
[0017] Furthermore, the pre - processed 26 - bit rotation factor is as follows:
[0018] ;
[0019] In the formula, is the rotation factor in the NTT calculation of the lattice cryptography algorithm; is the modulus corresponding to the lattice cryptography algorithm; is the bit - width corresponding to the modulus of the lattice cryptography algorithm, , .
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] First, the lattice cryptography modular multiplication method and multiplier based on the improved plantard algorithm of the present invention, in the hardware implementation process of the ML - DSA algorithm, in view of the characteristic that there is a known rotation factor constant as one of the multipliers in the calculation processes of polynomial multiplication operations NTT and INTT, optimize the original pantard algorithm. When one of the multipliers is known, the known constant is pre - processed in advance, and substituting the pre - processed result into the calculation process of the algorithm can save the post - processing of the calculation result, which can improve the operation speed of the algorithm while saving hardware resources. In the design and implementation process of the hardware circuit, a four - stage pipeline is designed by using methods such as critical path splitting and early truncation of invalid bits to help the algorithm be quickly implemented.
[0022] Second, the lattice cryptography modular multiplication method and multiplier based on the improved Plantard algorithm of the present invention use a four-stage pipeline design to split the critical path of modular multiplication into multiple parallel operations, and adopt techniques such as shift-addition instead of multiplication and dynamic truncation of bitwidth to support multiple pairs of modular multiplication operations to be completed within a single cycle. The actual measurement results based on the Xilinx Artix-7 FPGA platform show that compared with algorithms such as Barrett and K-RED, the multiplier reduces the logic resources (Slices) by 50.0% - 60.9% in the Kyber implementation, improves the hardware efficiency by 44.1% - 57.29%, has a working frequency of 350 - 400 MHz, only requires 5 cycles for a single modular multiplication, and the operation speed is 2.5 - 3 times faster than the traditional scheme.
[0023] Third, the lattice cryptography modular multiplication method and multiplier based on the improved Plantard algorithm of the present invention propose a dynamic bitwidth truncation and resource optimization technology, and adopt the method of using shift instead of multiplication to decompose the multiplication by q into three shifts and two additions and subtractions, saving at least 3 DSP units; during the transmission process of the intermediate data stream, the method of early truncation of invalid bits is adopted, and the third data path of the third-stage pipeline intercepts the lower 8 bits, reducing the adder bitwidth by 8 bits and reducing nearly 50% of the bitwidth resources. Description of the Drawings
[0024] Figure 1 It is a schematic diagram of the four-stage pipeline circuit structure of the proposed lattice cryptography multiplier based on the improved Plantard algorithm. Detailed Embodiment
[0025] The following further describes the embodiments of the present invention in detail with reference to the drawings.
[0026] The present invention proposes a lattice cryptography modular multiplication method based on the improved Plantard algorithm, which specifically includes: first, preprocess and optimize the input rotation factor. In view of the characteristic that the rotation factor B in the NTT is a known constant, B is preprocessed into 26 bits through the following formula , eliminating the post-processing step of the modular multiplication result:
[0027] ;
[0028] In the formula, is the rotation factor in the NTT calculation of the lattice cryptography algorithm; is the modulus corresponding to the lattice cryptography algorithm; is the bitwidth corresponding to the modulus of the lattice cryptography algorithm, , .
[0029] The modular multiplication process is decomposed into three-level calculations. Among them, in the first level, calculate , where represents a 12-bit multiplicand to be modulo-multiplied, and the truncation error is corrected in the second stage to obtain , and finally a 12-bit modulo-multiplication result is obtained in the third stage; meanwhile, the boundary condition of is detected, and if detected, the output is 0 through a multiplexer.
[0030] The following algorithm is the pseudocode of the original plantard modulo reduction algorithm.
[0031] Input: A, B, R, n, q with , ;
[0032] Output: C with and ;
[0033] 1: ;
[0034] 2: if C = q then;
[0035] 3: Return 0;
[0036] 4: end if;
[0037] 5: Return C.
[0038] This algorithm describes the execution process of the traditional plantard algorithm. The input parameters of this algorithm include integers , , pre-computed constant , bit width and modulus , where . Its core steps are: first calculate , shift the result to the right by bits and then add 1, then multiply by and shift to the right again by bits and round down to obtain the intermediate value ; if then output 0, otherwise output . This process converts the modulo multiplication into double-shift and multiplication operations by introducing the constant , but it depends on the pre-computation of and requires two high-precision shifts, resulting in large resource consumption and high critical path delay when implemented in hardware.
[0039] The following algorithm is the pseudocode of the improved pantard modulo reduction algorithm.
[0040] Input: A, B, n, q with , ;
[0041] Output: with ;
[0042] 1: ;
[0043] 2: ;
[0044] 3: ;
[0045] 4: if C = q then;
[0046] 5: Return 0;
[0047] 6: end if;
[0048] 7: Return C.
[0049] Compared with the original algorithm, the improved scheme eliminates the dependence on , directly preprocesses the input , calculates , and simplifies the modular multiplication operation into three steps: 1) Calculate to obtain ; 2) Right-shift by bits and add 1 to get ; 3) Calculate and handle the boundary conditions. This improvement eliminates the pre-computation step, and unifies the modular reduction process into a fixed-bitwidth operation through the preprocessing of , significantly reducing the complexity and resource occupancy of the hardware implementation.
[0050] On this basis, the present invention discloses a lattice cipher modular multiplier based on the improved plantard algorithm. According to the calculation characteristics in the NTT of the lattice cipher algorithm, the improved plantard algorithm is obtained, and a high-speed modular multiplier hardware structure is designed by means of decimal expansion. The hardware structure of the modular multiplier of the present invention specifically includes: a two-input multiplier, four two-input adders, and five registers. The total input of the modular multiplier is a 12-bit number to be modulo-multiplied and a preprocessed 26-bit rotation factor , and the total output is a 12-bit modular multiplication result . The circuit of the entire modular multiplier is divided into four-level pipelining.
[0051] In the first-level pipeline, the two input ports of multiplier 1 are respectively used to receive the 12-bit number to be modulo-multiplied and the preprocessed 26-bit rotation factor , the output port stores the 39-bit multiplication result in register 1. In the second-level pipeline, the result in register 1 is shifted right by 13 bits and then added with compensation constant 1 through adder 1 to obtain the result , and the output port stores the 13-bit result in register 2. In the third-level pipeline, the two input ports of adder 2 are respectively used to receive the result after shifting the lower 13 bits of register 2 by 3 bits and the result after shifting the lower 13 bits by 2 bits, and the output end stores the result in register 3; the two input ports of adder 3 are respectively used to receive the result of the lower 8 bits after the lower 13 bits stage of register 2 and the lower 13 bits, and the output end stores the result in register 4. In the fourth-level pipeline, the two input ends of adder 4 are respectively used to receive the results of register 3 and register 4, the output end stores the result in register 5, the output result of register 5 is shifted right by 5 bits, and then the boundary condition is detected through the hardware comparator, and 0 is output through the multiplexer (MUX). Finally, the modular multiplication results of two numbers to be modular-multiplied and the modulus q are output .
[0052] Figure 1 This is the hardware implementation scheme of the four-level pipeline modular multiplier architecture proposed by the present invention, and its structure strictly maps the calculation process of the improved plantard algorithm. The first-level pipeline uses a 12-bit × 26-bit multiplier to complete the operation, and after outputting a 39-bit result, it is compressed to 26 bits through a 13-bit right shift operation; the second-level pipeline performs a right shift on the high bits (in Kyber ) and adds 1 to correct the truncation error, generating a 4-bit intermediate value ; the third-level pipeline splits the parallel calculation path and utilizes mathematical characteristics to decompose into four shift operations and two addition operations, and at the same time intercepts the lower 8 bits to limit the bit width of the adder; the fourth-level pipeline combines the path results, outputs a 12-bit final modular multiplication value through a 5-bit right shift, and embeds a comparator to detect the boundary condition. This architecture reduces the logic resource occupancy by 50% - 60.9% through dynamic bit width truncation and shift-addition to replace the multiplier, increases the working frequency to 350 - 400 MHz, and only requires 5 cycles for a single modular multiplication, which is 2.5 - 3 times faster than the traditional scheme. The above scheme solves the bottleneck of the traditional modular reduction algorithm in terms of resource efficiency and calculation speed through algorithm optimization and hardware co-design, and is applicable to the NTT acceleration module of lattice cryptography algorithms such as Kyber and Dilithium, and has significant industrial application value.
[0053] The present invention discloses an efficient multiplier design based on an improved Plantard algorithm, which is specifically used to accelerate the modular multiplication operation in number-theoretic transform (NTT) of lattice cryptography algorithms (such as Kyber, Dilithium). Through the collaborative innovation of algorithms and hardware, this design significantly optimizes the computational efficiency and resource utilization. At the algorithm level, by preprocessing the rotation factors, the post-processing overhead of traditional algorithms is eliminated. In terms of hardware implementation, shift-addition is used to replace complex multipliers, combined with dynamic bit-width truncation technology, to compress the intermediate data bit-width from 39 bits to 21 bits, significantly reducing the consumption of logic resources. The actual measurement results based on the Xilinx Artix-7 FPGA platform show that, compared with algorithms such as Barrett and K-RED, this multiplier reduces the logic resources (Slices) by 50.0% - 60.9% in the Kyber implementation, improves the hardware efficiency by 44.1% - 57.29%, has a working frequency of 350 - 400 MHz, and only requires 5 cycles for a single modular multiplication, with the operation speed being 2.5 - 3 times higher than that of traditional solutions, and can quickly and efficiently complete the modular multiplication operation in the lattice cryptography digital signature scheme.
[0054] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0055] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications.
Claims
1. A lattice cryptography modular multiplication method based on an improved Plantard algorithm, characterized in that The method includes: Preprocess and optimize the input rotation factor, and preprocess the rotation factor into a 26-bit rotation factor through the following formula : ; In the formula, is the rotation factor in the NTT calculation of the lattice cryptography algorithm; is the modulus corresponding to the lattice cryptography algorithm; is the bit width corresponding to the modulus of the lattice cryptography algorithm, , ; Decompose the modular multiplication process into three levels of calculations. Among them, in the first level, calculate , where, represents the 12-bit multiplicand to be modulo-multiplied. In the second level, correct the truncation error to obtain . Finally, in the third level, obtain the 12-bit modular multiplication result ; at the same time, detect 's boundary condition. If detected, the output through the multiplexer is 0.
2. An lattice cryptography modular multiplier based on an improved Plantard algorithm, characterized in that, The lattice cipher modular multiplier includes a two-input multiplier, four two-input adders and five registers, which together form a four-stage pipeline; Among them, in the first-level pipeline, the two input ports of the two-input multiplier are respectively used to receive the 12-bit modulo multiplicand to be processed and the preprocessed 26-bit rotation factor , and the calculated 39-bit multiplication result is stored in the first register; In the second - stage pipeline, the first register shifts the multiplication result right by 13 bits and then sends it to the first adder. After adding the compensation constant 1 by the first adder, the result is obtained, and the 13 - bit result is stored in the second register; In the third-stage pipeline, the two input ports of the second adder are respectively used to receive the result of shifting the lower 13 bits of the second register to the right by 3 bits and the result of shifting the lower 13 bits to the right by 2 bits. After addition operation, the result is stored in the third register; the two input ports of the third adder are respectively used to receive the result of the last 8 bits of the lower 13 bits stage of the second register and the lower 13 bits. After addition operation, the result is stored in the fourth register; In the fourth - stage pipeline, the two input terminals of the fourth adder are respectively used to receive the calculation results of the third register and the fourth register. After addition operation processing, the result is stored in the fifth register, and the fifth register performs a right - shift operation on the calculation result by 5 bits to output a 12 - bit modular multiplication result , , and at the same time, the boundary conditions of are detected by the hardware comparator, and 0 is output through the multiplexer.
3. The lattice cryptographic modular multiplier based on the improved Plantard algorithm according to claim 2, wherein The preprocessed 26-bit rotation factor is as follows: ; In the formula, is the rotation factor in the NTT calculation of the lattice cryptography algorithm; is the modulus corresponding to the lattice cryptography algorithm; is the bit width corresponding to the modulus of the lattice cryptography algorithm, , .
Citation Information
Patent Citations
Fast modular multiplication operation method based on homomorphic encryption and modular multiplier
CN115268840A
Secret key packaging method based on number-theory transformation variant optimization of new modular multiplication algorithm
CN118944868A