Improved plantard algorithm-based lattice cipher modular multiplication method and modular multiplier

Through the improved plantard algorithm and the mode multiplier designed by the four-stage pipeline, the existing mode reduction algorithm has solved the problems of low computing efficiency and high resource consumption in the post-quantum grid cryptography NTT calculation, and efficient mode multiplication operation is realized, which significantly improves the computing performance and resource utilization.

CN120050042AActive Publication Date: 2025-05-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510497007.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-27
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing modulus subtraction algorithm has low computational efficiency, high resource consumption, and limited parallelism in post-quantum grid cryptography NTT calculations, especially when dealing with special prime modulus, there are defects in high computing overhead.

Method used

The grid cryptographic modular multiplication method and modular multiplier based on the improved plantard algorithm are adopted. By pre-processing and optimization of the rotation factor, the modular multiplication process is decomposed into three-level calculations, and multi-stage parallel operation is realized through the four-level pipeline design, and shift-addition substitution multiplication, dynamic truncation bit width and other technologies are adopted.

Benefits of technology

The throughput and energy efficiency ratio of the grid cryptographic algorithm NTT/INTT operations has been significantly improved, the logic resources have been reduced by 50.0%-60.9%, the hardware efficiency has been improved by 44.1%-57.29%, the working frequency has reached 350-400 MHz, and the single-mode multiplication only takes 5 cycles, and the computing speed has been increased by 2.5-3 times compared to traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050042A_ABST
    Figure CN120050042A_ABST
Patent Text Reader

Abstract

The invention discloses a lattice cryptographic modular multiplication method and a modular multiplier based on an improved plantard algorithm, which eliminate the post-processing overhead of the traditional modular multiplication method by preprocessing twiddle factors in the level of the modular multiplication method. In hardware implementation, shift-addition is adopted to replace a complex multiplier, a dynamic bit width truncation technology is combined, the bit width of intermediate data is compressed from 39 bits to 21 bits, and logic resource consumption is greatly reduced, so that modular multiplication operation in a lattice password digital signature scheme can be quickly and efficiently completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of hardware security, and in particular to a lattice cipher modular multiplication method and a modular multiplier based on an improved Plantard algorithm. Background Art

[0002] In the era of rapid development of quantum computing technology, modern cryptography is facing an unprecedented paradigm shift. The large integer factorization and elliptic curve discrete logarithm problems that are currently widely used in public key cryptography systems such as RSA and ECC are solvable in polynomial time when quantum computers execute Shor's algorithm. This breakthrough has fundamentally shaken the security foundation of traditional cryptography systems.

[0003] The core computing efficiency of the lattice algorithm directly depends on the performance of the number theoretic transform (NTT) on its polynomial ring. Polynomial multiplication accounts for nearly 50% of the running time of the lattice algorithm. NTT and INTT (Inverse Number Theoretic Transform, INTT) are the core acceleration algorithms for polynomial multiplication. The modular multiplication operation accounts for the main computing proportion and has a significant impact on the overall operating frequency and the number of computing cycles. Therefore, the optimization of the modular multiplication operation is crucial to improving the performance of the NTT implementation. This operation consists of two key stages: integer multiplication and modular reduction. The modular reduction process needs to map the high-order product result to a finite domain space, and its computational intensity accounts for 60%-80% of the modular multiplication operation. In resource-constrained embedded devices and high-throughput server scenarios, the degree of optimization of the modular reduction algorithm directly determines key performance indicators such as signature generation / verification latency and system energy consumption.

[0004] However, existing modular reduction algorithms (such as Barrett, Montgomery, etc.) have problems such as low computational efficiency, high resource consumption, and limited parallelism in post-quantum lattice cryptography NTT calculations, especially the high computational overhead when processing special prime moduli (such as q=8380417). Summary of the invention

[0005] The purpose of the present invention is to provide a lattice cipher modular multiplication method and modular multiplier based on an improved Plantard algorithm, which can solve the defects of the prior art such as long critical path of modular multiplication operation, redundant intermediate data bit width, and large number of calculation cycles, and significantly improve the throughput and energy efficiency of the NTT / INTT operation of the lattice cipher algorithm.

[0006] In order to achieve the above technical objectives, the technical solution adopted by the present invention is: In a first aspect, the present invention discloses a lattice cipher modular multiplication method based on an improved Plantard algorithm, the method comprising: The input rotation factor is preprocessed and optimized, and the rotation factor is converted into Preprocessed into 26-bit rotation factors : ; In the formula, It is the rotation factor in the NTT calculation of the lattice cryptographic algorithm; is the modulus corresponding to the lattice cryptographic algorithm; is the bit width corresponding to the modulus of the lattice cryptographic algorithm, , ; The modular multiplication process is decomposed into three levels of calculation, where in the first level, , where Represents the 12-bit modular multiplier to be obtained by correcting the truncation error in the second stage. , and finally get the 12-bit modular multiplication result at the third level ; Simultaneous detection If a boundary condition is detected, the output is 0 through the multiplexer.

[0007] In a second aspect, the present invention discloses a lattice cipher modular multiplier based on an improved Plantard algorithm, wherein the lattice cipher modular multiplier comprises a two-input multiplier, four two-input adders and five registers, which together constitute a four-stage pipeline; Among them, in the first stage pipeline, the two input ports of the two-input multiplier are used to receive the 12-bit modulus to be multiplied. And the preprocessed 26-bit rotation factor , the calculated 39-bit multiplication result is stored in the first register; In the second stage of the pipeline, the first register will multiply the result After the right shift operation of 13 bits, it is sent to the first adder, and the compensation constant 1 is added to the result by the first adder. , and convert the 13-bit result Store in the second register; In the third-stage pipeline, the two input ports of the second adder are respectively used to receive the result of the lower 13 bits of the second register after being shifted right by 3 bits and the result of the lower 13 bits after being shifted right by 2 bits, and the result is stored in the third register after the addition operation; the two input ports of the third adder are respectively used to receive the result of the lower 13 bits of the second register after the 8-bit stage and the lower 13 bits, and the result is stored in the fourth register after the addition operation; In the fourth stage pipeline, the two input ends of the fourth adder are used to receive the calculation results of the third register and the fourth register respectively. After the addition operation, the result is stored in the fifth register. The fifth register performs a 5-bit right shift operation on the calculation result to output a 12-bit modular multiplication result. , , and detected by hardware comparator The boundary condition is met and the multiplexer outputs 0.

[0008] Furthermore, the preprocessed 26-bit rotation factor for: ; In the formula, It is the rotation factor in the NTT calculation of the lattice cryptographic algorithm; is the modulus corresponding to the lattice cryptographic algorithm; is the bit width corresponding to the modulus of the lattice cryptographic algorithm, , .

[0009] Compared with the prior art, the present invention has the following beneficial effects: First, the lattice cipher modular multiplication method and modular multiplier based on the improved Plantard algorithm of the present invention, in the hardware implementation process of the ML-DSA algorithm, optimizes the original Pantard algorithm for the characteristic that there is a known rotation factor constant as one of the multipliers in the calculation process of the polynomial multiplication operations NTT and INTT. When one of the multipliers is known, the known constant is preprocessed in advance, and the result after preprocessing is substituted into the calculation process of the algorithm, which can save the post-processing of the calculation result, and can also improve the calculation speed of the algorithm while saving hardware resources. In the design and implementation process of the hardware circuit, the four-stage pipeline is designed by using methods such as critical path splitting and early truncation of invalid bits to help the algorithm to be implemented quickly.

[0010] Second, the lattice cipher modular multiplication method and modular multiplier based on the improved Plantard algorithm of the present invention splits the modular multiplication key path into multi-level parallel operations through a four-stage pipeline design, adopts shift-addition to replace multiplication, dynamically truncates the bit width and other technologies, and supports completing multiple pairs of modular multiplication operations in a single cycle. The actual measurement results based on the Xilinx Artix-7 FPGA platform show that compared with algorithms such as Barrett and K-RED, the modular multiplier reduces the logical resources (Slices) by 50.0%-60.9% in the Kyber implementation, improves the hardware efficiency by 44.1%-57.29%, and the operating frequency reaches 350-400 MHz. A single modular multiplication only requires 5 cycles, and the operation speed is 2.5-3 times higher than the traditional solution.

[0011] Third, the lattice cipher modular multiplication method and modular multiplier based on the improved Plantard algorithm of the present invention proposes dynamic bit width truncation and resource optimization technology, adopts the method of replacing multiplication by shift, decomposes the multiplication of q into three shifts and two additions and subtractions, and saves at least three DSP units; in the transmission process of the intermediate data stream, adopts the method of early truncation of invalid bits, and intercepts the third data path of the third-level pipeline The lower 8 bits reduce the adder bit width by 8 bits, reducing the bit width resources by nearly 50%. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a schematic diagram of the four-stage pipeline circuit structure of the proposed lattice cipher modular multiplier based on the improved Plantard algorithm. DETAILED DESCRIPTION

[0013] The embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings.

[0014] The present invention proposes a lattice cipher modular multiplication method based on an improved Plantard algorithm, which specifically includes: firstly, preprocessing and optimizing the input rotation factor, and according to the characteristic that the rotation factor B in NTT is a known constant, preprocessing B into 26 bits by the following formula , eliminating the post-processing step of the modular multiplication result: ; In the formula, It is the rotation factor in the NTT calculation of the lattice cryptographic algorithm; is the modulus corresponding to the lattice cryptographic algorithm; is the bit width corresponding to the modulus of the lattice cryptographic algorithm, , .

[0015] The modular multiplication process is decomposed into three levels of calculation, where in the first level, , where Represents the 12-bit modular multiplier to be obtained by correcting the truncation error in the second stage. , and finally get the 12-bit modular multiplication result at the third level ; Simultaneous detection If a boundary condition is detected, the output is 0 through the multiplexer.

[0016] The following algorithm is the pseudocode for the original Plantard modular reduction algorithm.

[0017] Input: A, B, R, n, q with , ; Output: C with and ; 1: ; 2: if C=q then; 3: return 0; 4: end if; 5: Return to C.

[0018] This algorithm describes the execution process of the traditional Plantard algorithm. The input parameters of this algorithm include integer , , precomputed constants , bit width and modulus ,in The core steps are: first calculate , shift the result right Add 1 to the last digit and then multiply by and move right again Round the bit to get the middle value ;like Then output 0, otherwise output This process is accomplished by introducing the constant Converts modular multiplication into a double shift and multiply operation, but its dependence The pre-calculation and two high-precision shifts are required, which leads to large resource consumption and high critical path delay during hardware implementation.

[0019] The following algorithm is the pseudo code of the improved Pantard modular reduction algorithm.

[0020] Input: A, B, n, q with , ; Output: with ; 1: ; 2: ; 3: ; 4: if C=q then; 5: return 0; 6: end if; 7: Return to C.

[0021] Compared with the original algorithm, the improved scheme cancels the Dependency, directly on the input Preprocessing, calculation , and simplifies the modular multiplication operation into three steps: 1) Calculate get 2) Yes Shift Right Add 1 after the digit ;3) Calculation and handle boundary conditions. This improvement eliminates the need to precompute Steps, through The preprocessing unifies the modular reduction process into a fixed-bit-width operation, significantly reducing the complexity and resource usage of hardware implementation.

[0022] On this basis, the present invention discloses a lattice cipher modular multiplier based on an improved Plantard algorithm. Aiming at the calculation characteristics of the lattice cipher algorithm NTT, an improved Plantard algorithm is obtained, and a high-speed modular multiplier hardware structure is designed by using a decimal expansion method. The hardware structure of the modular multiplier of the present invention specifically includes: a two-input multiplier, four two-input adders and five registers. The total input of the modular multiplier is a 12-bit modular multiplication number. and the preprocessed 26-bit rotation factors , the total output is a 12-bit modular multiplication result The entire analog multiplier circuit is divided into a four-stage pipeline.

[0023] In the first stage of the pipeline, the two input ports of multiplier 1 are used to receive the 12-bit modulus to be multiplied. And the preprocessed 26-bit rotation factor , the output port will be the 39-bit multiplication result Stored in register 1. In the second stage pipeline, the result in register 1 is right shifted 13 bits and then added to the compensation constant 1 by adder 1 to obtain the result. , the output port will be the 13-bit result Stored in register 2. In the third stage pipeline, the two input ports of adder 2 are used to receive the result of the lower 13 bits of register 2 after right shifting 3 bits and the result of the lower 13 bits after right shifting 2 bits, and the output end stores the result in register 3; the two input ports of adder 3 are used to receive the result of the lower 13 bits of register 2 after the 8-bit stage and the lower 13 bits, and the output end stores the result in register 4. In the fourth stage pipeline, the two input ports of adder 4 are used to receive the results of register 3 and register 4, and the output end stores the result in register 5. The output result of register 5 is right shifted by 5 bits and then detected by the hardware comparator. The boundary condition is met, and the multiplexer (MUX) outputs 0. Finally, the modular multiplication result of the two modular multiplication numbers and the modulus q is output. .

[0024] Figure 1This is the hardware implementation of the four-stage pipeline modular multiplier architecture proposed in this invention. Its structure strictly maps the calculation process of the improved Plantard algorithm. The first stage of the pipeline is completed by a 12-bit × 26-bit multiplier. The 39-bit result is then compressed to 26 bits by right shifting 13 bits. Bit (Kyber ) is shifted right and 1 is added to correct the truncation error to generate a 4-bit intermediate value ; The third-level pipeline splits the parallel computing path and uses The mathematical properties of Decomposed into four shifts and two addition operations, while intercepting The lower 8 bits are used to limit the adder bit width; the fourth-stage pipeline merges the path results, outputs the 12-bit final modular multiplication value by right shifting 5 bits, and embeds the comparator detection The architecture reduces the logic resource usage by 50%-60.9% through dynamic bit width truncation and shift-addition to replace multipliers, increases the operating frequency to 350-400 MHz, and only requires 5 cycles for a single modular multiplication, which is 2.5-3 times faster than the traditional solution. The above solution solves the bottleneck of resource efficiency and computing speed of the traditional modular reduction algorithm through algorithm optimization and hardware co-design. It is suitable for the NTT acceleration module of lattice cryptographic algorithms such as Kyber and Dilithium, and has significant industrial application value.

[0025] The present invention discloses a high-efficiency modular multiplier design based on the improved Plantard algorithm, which is specifically used to accelerate the modular multiplication operation of number theoretic transformation (NTT) in lattice cryptographic algorithms (such as Kyber and Dilithium). The design significantly optimizes computing efficiency and resource utilization through the collaborative innovation of algorithms and hardware. At the algorithm level, the post-processing overhead of traditional algorithms is eliminated by preprocessing the rotation factors. In hardware implementation, the shift-addition method is used to replace the complex multiplier, combined with the dynamic bit width truncation technology, to compress the intermediate data bit width from 39 bits to 21 bits, greatly reducing the consumption of logic resources. The measured results based on the Xilinx Artix-7 FPGA platform show that compared with algorithms such as Barrett and K-RED, the modular multiplier in the Kyber implementation reduces logic resources (Slices) by 50.0%-60.9%, improves hardware efficiency by 44.1%-57.29%, operates at a frequency of 350-400 MHz, and requires only 5 cycles for a single modular multiplication. The computing speed is 2.5-3 times faster than that of traditional schemes, and can quickly and efficiently complete the modular multiplication operations in the lattice cryptographic digital signature scheme.

[0026] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0027] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A lattice cipher modular multiplication method based on an improved Plantard algorithm, characterized in that: The method comprises: The input rotation factor is preprocessed and optimized, and the rotation factor is converted into Preprocessed into 26-bit rotation factors : ; In the formula, It is the rotation factor in the NTT calculation of the lattice cryptographic algorithm; is the modulus corresponding to the lattice cryptographic algorithm; is the bit width corresponding to the modulus of the lattice cryptographic algorithm, , ; The modular multiplication process is decomposed into three levels of calculation, where in the first level, , where Represents the 12-bit modular multiplier to be obtained by correcting the truncation error in the second stage. , and finally get the 12-bit modular multiplication result at the third level ; Simultaneous detection If a boundary condition is detected, the output is 0 through the multiplexer.

2. A lattice cipher module multiplier based on an improved Plantard algorithm, characterized in that: The lattice code module multiplier includes a two-input multiplier, four two-input adders and five registers, which together form a four-stage pipeline; Among them, in the first stage pipeline, the two input ports of the two-input multiplier are used to receive the 12-bit modulus to be multiplied. And the preprocessed 26-bit rotation factor , the calculated 39-bit multiplication result is stored in the first register; In the second stage of the pipeline, the first register will multiply the result After the right shift operation of 13 bits, it is sent to the first adder, and the compensation constant 1 is added to the result by the first adder. , and convert the 13-bit result Store in the second register; In the third-stage pipeline, the two input ports of the second adder are respectively used to receive the result of the lower 13 bits of the second register after being shifted right by 3 bits and the result of the lower 13 bits after being shifted right by 2 bits, and the result is stored in the third register after the addition operation; the two input ports of the third adder are respectively used to receive the result of the lower 13 bits of the second register after the 8-bit stage and the lower 13 bits, and the result is stored in the fourth register after the addition operation; In the fourth stage pipeline, the two input ends of the fourth adder are used to receive the calculation results of the third register and the fourth register respectively. After the addition operation, the result is stored in the fifth register. The fifth register performs a 5-bit right shift operation on the calculation result to output a 12-bit modular multiplication result. , , and detected by hardware comparator The boundary condition is met and the multiplexer outputs 0.

3. The lattice cipher module multiplier based on the improved Plantard algorithm according to claim 2, characterized in that: The preprocessed 26-bit rotation factor for: ; In the formula, It is the rotation factor in the NTT calculation of the lattice cryptographic algorithm; is the modulus corresponding to the lattice cryptographic algorithm; is the bit width corresponding to the modulus of the lattice cryptographic algorithm, , .

Citation Information

Patent Citations

  • Fast modular multiplication operation method based on homomorphic encryption and modular multiplier

    CN115268840A

  • Lattice cipher modular multiplier based on parallel K-RED modular reduction algorithm

    CN118796152A

  • Secret key packaging method based on number-theory transformation variant optimization of new modular multiplication algorithm

    CN118944868A

Cited By

  • Configurable reduction circuit for integer operation in post quantum cryptography algorithm

    CN121864307A

  • A configurable reduction circuit for integer operations in post-quantum cryptographic algorithms

    CN121864307B