NTT instruction acceleration method of post quantum cryptography algorithm Kyber based on RISC-V
By using a RISC-V-based instruction acceleration method, Kyber's NTT coefficients are grouped and the problem of NTT computational resource limitations is solved by utilizing register batch loading and dual-instruction design. This achieves high computational efficiency and resource optimization, thereby improving overall performance.
Patent Information
- Application Number
- CN202511469269.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-15
AI Technical Summary
The existing post-quantum cryptography algorithm Kyber's number-theoretic transformation (NTT) suffers from low overall computational efficiency due to the frequent data access and high computational resource requirements during large-scale computations. Furthermore, the existing methods fail to effectively combine software instruction set optimization.
A RISC-V-based instruction acceleration method is adopted, which divides the coefficients of NTT into multiple groups and uses the programmable registers of RISC-V for batch loading and dual-instruction design. By interleaving the instruction sequence and allocating registers reasonably, multiple instructions can be executed simultaneously, making full use of processor resources.
It significantly reduced the computation cycle of NTT, improved overall computational efficiency, boosted performance by 63.93%, optimized resource utilization, and improved overall energy efficiency.
Smart Images

Figure CN120934760A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of post-quantum cryptography technology, specifically relating to an NTT instruction acceleration method for the Kyber post-quantum cryptography algorithm based on RISC-V. Background Technology
[0002] Post-quantum cryptography (PQC) is a cryptographic algorithm designed to resist attacks from future quantum computers. Quantum computers can efficiently solve the mathematical problems relied upon by traditional public-key cryptography (such as RSA and ECC), rendering existing encryption algorithms insecure in the quantum computing era. To address this threat, the development of post-quantum cryptography focuses on cryptographic schemes that remain secure against quantum computers.
[0003] Kyber is a lattice-based post-quantum public-key cryptography scheme nominated as a candidate algorithm in the NIST post-quantum cryptography standardization process. In Kyber, the Number Theoretic Transform (NTT) is one of the key technologies for achieving efficient encryption and decryption. NTT is a Fast Fourier Transform (FFT) over the modular field, similar to the traditional FFT, but it operates over finite fields, especially in modular arithmetic with large integers. NTT significantly improves the algorithm's efficiency by accelerating polynomial multiplication, increasing encryption and decryption speed, and optimizing hardware implementation. This makes Kyber a promising encryption scheme in the field of post-quantum cryptography, especially in addressing the threat of quantum computing. Although NTT significantly accelerates polynomial multiplication, its speedup is affected by frequent data access, high demands on computing resources and memory, hardware resource limitations, and the difficulty of optimizing the inverse transform during large-scale computations. Existing technologies have been researched to address this issue, but existing methods rely too heavily on hardware design and neglect software optimization, failing to fully utilize the computational capabilities of the software instruction set; or they reduce the computational load of core computing but incur polynomial conversion overhead; they cannot achieve resource optimization and improve overall energy efficiency while improving instruction execution efficiency. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides an NTT instruction acceleration method based on the RISC-V post-quantum cryptography algorithm Kyber, which can effectively reduce NTT cycle consumption and improve overall NTT computation efficiency.
[0005] The present invention discloses an NTT instruction acceleration method for the Kyber post-quantum cryptography algorithm based on RISC-V, comprising the following steps: Step 1: Divide the 256 coefficients in Kyber's NTT into 16 groups, each containing 16 coefficients; use the RISC-V programmable registers to load them in batches, loading 16 coefficients and their corresponding twiddle factors each time. Step 2: Perform butterfly operations on the 8 pairs of data in each layer in sequence, including the multiplication module, the reduction module, and the addition and subtraction module executed in sequence; Step 3: When both pairs of data are multiplication module operations, processor C can only issue one multiplication module operation instruction at a time, and the second multiplication operation instruction must wait for the first multiplication operation instruction to complete; when the two pairs of data are not double multiplication module operations, processor C issues the corresponding module operation instructions for both pairs of data simultaneously. Step 4: After the butterfly operation is completed for the first five pairs of data in the first layer, insert the multiplication module of the next layer of data into the addition and subtraction module of the previous layer. The multiplication module of the second layer and the addition and subtraction module of the first layer are executed concurrently. Repeat step 3 until all groups of data have been executed in the first to third layers. Step 5: After the butterfly operation is completed for the first five pairs of data in the fourth layer, insert the multiplication module of the next layer's data pairs into the addition and subtraction module of the previous layer. The multiplication module of the fifth layer and the addition and subtraction module of the fourth layer are executed concurrently. Repeat step 3 until all groups of data have been calculated in the fourth to seventh layers.
[0006] Further, in step 1, the editable registers of the RISC-V are... It stores the 16 coefficients and corresponding rotation factors for each group; As a temporary register, it is used to store intermediate variables; The modulus is stored as a constant register. and ,register As an intermediate variable for scheduling.
[0007] Furthermore, in step 2, the multiplication module is... The reduction module is The addition and subtraction modules are , ; where a i and a j Let i and j be the two coefficients of NTT, satisfying the butterfly distance at each level; ζ is the rotation factor.
[0008] Furthermore, the butterfly distance in the first layer is 128, the butterfly distance in the second layer is 64, the butterfly distance in the third layer is 32, the butterfly distance in the fourth layer is 16, the butterfly distance in the fifth layer is 8, the butterfly distance in the sixth layer is 4, and the butterfly distance in the seventh layer is 2.
[0009] The beneficial effects of this invention are as follows: The method of this invention loads data in NTT in batches into registers, making full use of the efficient registers of RISC-V for computation, and completing data interaction in registers, thus making efficient use of register resources; The method of this invention designs the NTT with dual-instruction, and by issuing multiple instructions simultaneously, the processor can complete more tasks in the same clock cycle, thereby significantly improving instruction execution efficiency; Dual-instruction can better utilize multiple execution units of the processor (such as arithmetic logic units, load memory units, etc.), reducing resource idleness and improving overall energy efficiency; This invention interleaves modules that cannot be dual-instructed with modules that can be dual-instructed, adjusts the instruction order to ensure that there are no numerical conflicts, and the adjusted modules can not only run dual-instruction between instructions, but also run dual-instruction continuously between different modules, ensuring a high dual-instruction state for the overall NTT. Attached Figure Description
[0010] Figure 1 This is a flowchart of the method described in this invention; Figure 2 This is a design diagram of NTT dual-send instruction set based on RISC-V. Detailed Implementation
[0011] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0012] like Figure 1 As shown, the present invention provides a method for accelerating the NTT instruction of the Kyber post-quantum cryptography algorithm based on RISC-V, comprising the following steps: Step 1: Divide the 256 coefficients in Kyber's NTT into 16 groups, each containing 16 coefficients; use the RISC-V programmable registers to load them in batches, loading 16 coefficients and their corresponding twiddle factors each time. Step 2: Perform butterfly operations on the 8 pairs of data in each layer in sequence, including the multiplication module, the reduction module, and the addition and subtraction module executed in sequence; Step 3: When both pairs of data are multiplication module operations, processor C can only issue one multiplication module operation instruction at a time, and the second multiplication operation instruction must wait for the first multiplication operation instruction to complete; when the two pairs of data are not double multiplication module operations, processor C issues the corresponding module operation instructions for both pairs of data simultaneously. Step 4: After the butterfly operation is completed for the first five pairs of data in the first layer, insert the multiplication module of the next layer of data into the addition and subtraction module of the previous layer. The multiplication module of the second layer and the addition and subtraction module of the first layer are executed concurrently. Repeat step 3 until all groups of data have been executed in the first to third layers. Step 5: After the butterfly operation is completed for the first five pairs of data in the fourth layer, insert the multiplication module of the next layer's data pairs into the addition and subtraction module of the previous layer. The multiplication module of the fifth layer and the addition and subtraction module of the fourth layer are executed concurrently. Repeat step 3 until all groups of data have been calculated in the fourth to seventh layers.
[0013] In this invention, the butterfly operation and the multiplication module are... The reduction module is The addition and subtraction modules are , Where ζ is the rotation factor; a i and a j For NTT, i and j are two coefficients that need to satisfy the butterfly distance of each layer: 128 for the first layer, 64 for the second layer, 32 for the third layer, 16 for the fourth layer, 8 for the fifth layer, 4 for the sixth layer, and 2 for the seventh layer.
[0014] For Kyber's 7-layer NTT, this invention adopts a layered loading strategy. The first three layers are processed as a whole, and the last four layers are processed as a whole. That is, after the first three layers of the same batch of 8 pairs of data are processed, the remaining 15 sets of 8 pairs of data are then loaded and processed in the first three layers. After all 16 sets of 8 pairs of data are processed in the first three layers, the last four layers are then processed. The operation steps for the last four layers are the same as those for the first three layers.
[0015] Taking the first three layers of NTT with dimension 256 in Kyber as an example, the array set Loaded sequentially into registers .
[0016] In the first layer, memory is represented as: , After the two multiplication modules are completed, the reduction module is entered, that is: , Then perform the addition and subtraction module operations, that is: .
[0017] In the first layer, After the update is complete, execute the second layer. This means inserting the multiplication module into the addition and subtraction module of the previous layer.
[0018] like Figure 2As shown in the embodiment, red squares represent multiplication modules, green squares represent reduction modules, and blue squares represent addition and subtraction modules. In each column, the numbers represent instructions to be executed sequentially. In the same column, squares of the same color containing the same number indicate that the instruction can be executed twice.
[0019] from Figure 2 It can be seen that the coefficient values The changes will occur in the first instruction, thus affecting the source data for the second instruction. Therefore, by using temporary registers to store the coefficients, a shuffle operation is performed on the data of the 8 pairs of coefficients. For example, in level 1, temporary registers can be used. To save in advance The value of , that is, let (Temporary registers can be reused, such as...) In completing It then regains its freedom and can be reused in the next calculation. Therefore, the register It can replace Thus, the calculation formula of the addition and subtraction module ( There are no data conflicts. Data shuffling is performed in the addition and subtraction modules at each level to ensure dual-processing capability. Finally, the coefficients in the registers are saved to the corresponding NTT memory.
[0020] A continuous sequence of multiplication instructions cannot be executed in dual-processor mode. This invention, through instruction interleaving and with proper register allocation, allows the multiplication module of the next lower level (e.g., the second level) to be used alternately with the addition / subtraction module of the previous level (the first level). When the multiplication and addition modules appear simultaneously and there are no data conflicts, the processor can issue both modules concurrently and execute them in dual-processor mode.
[0021] For example, in the first layer, after the fifth pair of coefficients in the addition module is calculated, the number 0 in the multiplication module can be paired with the number 5 in the addition module. That is, assuming the data pair in the first layer is: (a0,a8),(a1,a9),(a2,a10),(a3,a11),(a4,a12),(a5,a13),(a6,a14),(a7,a15); The second layer is: (a0,a4),(a1,a5),(a2,a6),(a3,a7),(a8,a12),(a9,a13),(a10,a14),(a11,a15); When completing the butterfly operations (a0,a8), (a1,a9), (a2,a10), (a3,a11), (a4,a12) in the first layer, you can perform (a0,a4) in the second layer, then complete (a5,a13) in the first layer and (a1,a5) in the second layer, and so on.
[0022] In this case, the present invention can load coefficients before addition module 4 instead of multiplication module 0 (because a4 is used in addition module 4, and after it is used, multiplication module 0 continues to operate on a4, and then addition module 5 is operated on, which is equivalent to reusing a4 from addition module 4). Addition module 4 is the first layer. The fifth addition module is the first layer. The 0th element of the multiplication module is the second layer. Therefore, the deeply optimized modular operation ensures highly dual-process execution.
[0023] The data obtained through experiments in this invention were programmed and tested using a combination of C and assembly language on the Linux operating system. The hardware and software configurations used in the experiments are detailed in Table 1. The data source in the experiments was a polynomial of dimension 256 used as the input to NTT.
[0024] Table 1 Experimental Platform Configuration
[0025] The method described in this invention is compared with the NTT implementation in prior art document 1 (i.e., J. Bos et al., “CRYSTALS-Kyber: ACCA-secure module-lattice-based KEM,” in Proc. IEEE Eur. Symp. Secur. Privacy(EuroS & P), Piscataway, NJ, USA: IEEE Press, 2018, pp. 353–367.). Compared to the NTT implementation in document 1 which consumes 24525 cycles, this invention only requires 8845 cycles, resulting in a performance improvement of 63.93%. The performance of the NTT implemented in this invention is shown in Table 2.
[0026] Table 2 Performance of NTT
[0027] In the method described in this invention, all intermediate results are stored in registers throughout the NTT calculation process, requiring proper register allocation and management. In a 64-bit RISC-V architecture, storing the NTT coefficients in registers reduces the number of memory accesses and expands the overflow range of the coefficients. This invention can be applied not only to the RISC-V platform architecture but also to other dual-instruction architectures with 64-bit processors, including but not limited to x86-64 (AMD64 and Intel64), ARM64, etc.
[0028] The above description is merely a preferred embodiment of the present invention and is not intended to further limit the present invention. All equivalent changes made based on the description and drawings of the present invention are within the protection scope of the present invention.
Claims
1. A method for accelerating the NTT instruction of the Kyber post-quantum cryptography algorithm based on RISC-V, characterized in that, include: Step 1: Divide the 256 coefficients in Kyber's NTT into 16 groups, with each group containing 16 coefficients; The programmable registers of RISC-V are used to load in batches, with 16 coefficients and corresponding twitch factors loaded each time. Step 2: Perform butterfly operations on the 8 pairs of data in each layer in sequence, including the multiplication module, the reduction module, and the addition and subtraction module executed in sequence; Step 3: When both pairs of data are multiplication module operations, processor C can only issue one multiplication module operation instruction at a time, and the second multiplication operation instruction must wait for the first multiplication operation instruction to complete; when the two pairs of data are not double multiplication module operations, processor C issues the corresponding module operation instructions for both pairs of data simultaneously. Step 4: After the butterfly operation is completed for the first five pairs of data in the first layer, insert the multiplication module of the next layer of data into the addition and subtraction module of the previous layer. The multiplication module of the second layer and the addition and subtraction module of the first layer are executed concurrently. Repeat step 3 until all groups of data have been executed in the first to third layers. Step 5: After the butterfly operation is completed for the first five pairs of data in the fourth layer, insert the multiplication module of the next layer's data pairs into the addition and subtraction module of the previous layer. The multiplication module of the fifth layer and the addition and subtraction module of the fourth layer are executed concurrently. Repeat step 3 until all groups of data have been calculated in the fourth to seventh layers.
2. The NTT instruction acceleration method for the Kyber post-quantum cryptography algorithm based on RISC-V according to claim 1, characterized in that, In step 1, the editable registers of the RISC-V are: It stores the 16 coefficients and corresponding rotation factors for each group; The rest are temporary registers used to store intermediate variables.
3. The NTT instruction acceleration method for the Kyber post-quantum cryptography algorithm based on RISC-V according to claim 2, characterized in that, In step 2, the multiplication module is... The reduction module is The addition and subtraction modules are , ; where a i and a j Let i and j be the two coefficients of NTT, satisfying the butterfly distance at each layer; ζ is the rotation factor.
4. The NTT instruction acceleration method for the Kyber post-quantum cryptography algorithm based on RISC-V according to claim 3, characterized in that, The butterflies on the first layer are 128 apart, those on the second layer are 64 apart, those on the third layer are 32 apart, those on the fourth layer are 16 apart, those on the fifth layer are 8 apart, those on the sixth layer are 4 apart, and those on the seventh layer are 2 apart.
Citation Information
Patent Citations
RISC-V-based processor special for post-quantum cryptography algorithm
CN116432765A
NTT hardware implementation system based on FPGA platform
CN116545622A
NTT polynomial multiplier capable of realizing in-situ storage, constant geometric structure and no memory access conflict
CN119356638A
High level synthesis of cloud cryptography circuits
WO2024253846A1