Implementation Method and System of Post-Quantum Cryptographic Algorithm for Resource-Constrained Processors

By using the NTT method and a hybrid reduction algorithm to process polynomial product and modular reduction operations on resource-constrained processors, the problem of large resource consumption is solved and the efficient operation of the post-quantum cryptographic algorithm is achieved.

CN115801244BActive Publication Date: 2025-07-04SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211406145.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2025-07-04
Estimated Expiration
2042-11-10

AI Technical Summary

Technical Problem

The existing post-quantum cryptography algorithms consume too much computing resources and storage resources on resource-constrained processors, especially polynomial multiplication and modular calculations have become the main acceleration bottlenecks and cannot be effectively implemented.

Method used

The NTT method is used to calculate the polynomial product, and a mixed reduction algorithm of Montgomery reduction algorithm and k reduction algorithm are combined to process modular reduction operations to reduce calculation and storage resource consumption.

Benefits of technology

The effective implementation of the post-quantum cryptographic algorithm is implemented on the resource-constrained processor, and the key generation, encryption and decryption speeds are increased by 3 times, 3.2 times and 3.4 times respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115801244B_ABST
    Figure CN115801244B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and system for implementing a post-quantum cryptographic algorithm for a resource-constrained processor, which belongs to the field of information security technology. The solution includes the processing of modular reduction operations and polynomial multiplication operations. Among them, for the polynomial multiplication operation in the post-quantum cryptographic algorithm, a number-theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing; for the modular reduction operation in the post-quantum cryptographic algorithm, a hybrid reduction algorithm based on the k-reduction algorithm and the Montgomery reduction algorithm is used for processing; wherein, the hybrid reduction algorithm is specifically: for the first N-1 layers of operations in the number-theoretic transform, the k-reduction algorithm is used for reduction, and starting from the Nth layer, the Montgomery reduction algorithm is used; wherein, N is a positive integer, and the value of N is the layer where overflow first occurs during the transformation; based on the processing process of the modular reduction operation and the polynomial multiplication operation, the implementation of the post-quantum cryptographic algorithm on the resource-constrained processor is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of information security, and in particular, relates to a method and system for implementing a post-quantum cryptographic algorithm for a resource-constrained processor. Background Technique

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the development of technology, cryptographic technology is evolving from traditional public-key-based cryptographic technology towards PQC (Post-Quantum Cryptography) technology. The so-called PQC technology refers to cryptographic technology that can resist attacks by quantum computers and is therefore also known as "quantum-resistant cryptographic technology". The "post" here means that after the emergence of large-scale and stable quantum computers, the vast majority of existing public-key cryptographic algorithms (such as RSA, Diffie-Hellman, elliptic curves, etc.) will be broken, and only cryptographic algorithms that can resist such attacks can survive after entering the era of quantum computing.

[0004] In PQC technology, the post-quantum cryptographic Saber algorithm is one of them. The Saber algorithm is a cryptographic primitive based on the hard problem of MLWR (Module Learning With Rounding) on lattices. Saber.PKE (Public Key Encryption) is an IND-CPA (Indistinguishability under Chosen-Plaintext Attack) secure encryption scheme, and Saber.KEM (Key Encapsulation Mechanism) is an IND-CCA (Indistinguishability under Chosen-Ciphertext Attack) secure key encapsulation scheme. The Saber.PKE can be transformed into Saber.KEM by using the FO (Fujisaki-Okamoto) transformation. In order to achieve the original design intention of the Saber algorithm and realize the design goals of simplicity, efficiency, and flexibility, the experts who designed the Saber algorithm chose the modulo integer as a multiple of 2, -13 to avoid complex modulo reduction operations and rejection sampling processes. However, the inventor found that in the hardware implementation of the existing Saber algorithm, polynomial multiplication is still the main acceleration bottleneck in the algorithm implementation, and its calculation still consumes a large amount of computing resources and storage resources, and cannot be effectively implemented on processors with limited resources and performance. Summary of the Invention

[0005] To solve the above problems, the present disclosure provides a method and system for implementing a post-quantum cryptographic algorithm for a resource-constrained processor. In the solution, the NTT method is adopted in the resource-constrained processor to calculate the polynomial product in the post-quantum cryptographic algorithm. Meanwhile, a hybrid reduction algorithm combining the Montgomery reduction algorithm and the k-reduction algorithm is proposed to implement the modular reduction operation in the post-quantum cryptographic algorithm, effectively reducing the consumption of computing resources and storage resources, and realizing the effective implementation of the lattice-based post-quantum cryptographic algorithm on a processor with limited resources and performance.

[0006] According to the first aspect of the embodiments of the present disclosure, a method for implementing a post-quantum cryptographic algorithm for a resource-constrained processor is provided, including processing modular reduction operations and polynomial product operations, where

[0007] For the polynomial product operation in the post-quantum cryptographic algorithm, the number-theoretic transform algorithm based on the Chinese Remainder Theorem is adopted for processing;

[0008] For the modular reduction operation in the post-quantum cryptographic algorithm, a hybrid reduction algorithm based on the k-reduction algorithm and the Montgomery reduction algorithm is adopted for processing; where the hybrid reduction algorithm is specifically: the k-reduction algorithm is adopted for reduction in the first N - 1 layers of the number-theoretic transform, and the Montgomery reduction algorithm is adopted starting from the Nth layer; where N is a positive integer, and the value of N is the layer where overflow first occurs during the transform;

[0009] Based on the processing procedures of the modular reduction operation and the polynomial product operation, the implementation of the post-quantum cryptographic algorithm on a resource-constrained processor is completed.

[0010] Further, the resource-constrained processor is a 32-bit platform of ARM Cortex-M0 / M0+, which has the following constraints: the processor does not have high-performance multiplication, and can only calculate the product of 32 bits × 32 bits to obtain a 32-bit result; the processor does not have multi-operation consecutive instructions such as multiply-add and multiply-subtract; the processor does not have a single-instruction multiple-data operation function.

[0011] Further, the ARM Cortex-M0 / M0+ processor includes 16 32-bit registers, where R0 - R12 are general-purpose registers, and R13 - R15 are special registers.

[0012] Further, for the polynomial product operation in the post-quantum cryptographic algorithm, the number-theoretic transform algorithm based on the Chinese Remainder Theorem is adopted for processing, where the forward number-theoretic transform adopts the CT-butterfly transform, and the inverse number-theoretic transform adopts the GS-butterfly transform.

[0013] Further, the input coefficients of the forward number-theoretic transform are in normal order. After CT-butterfly transform, the coefficient order is bit-reversed. The input coefficients of the inverse number-theoretic transform are bit-reversed, and after GS-butterfly transform, the coefficient order is normal.

[0014] Further, the post-quantum cryptography algorithm adopts the Saber algorithm, a lattice-based post-quantum cryptography algorithm.

[0015] Further, the post-quantum cryptography algorithm can also be Kyber, NTRU, Dilithium, and Falcon algorithms.

[0016] According to the second aspect of the embodiments of the present disclosure, a system for implementing a post-quantum cryptography algorithm for a resource-constrained processor is provided, including an ARM Cortex-M0 / M0+ processor and a post-quantum cryptography algorithm executed on the processor. Among them, the implementation of the post-quantum cryptography algorithm specifically adopts the following steps:

[0017] For the polynomial multiplication operation in the post-quantum cryptography algorithm, a number-theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing;

[0018] For the modular reduction operation in the post-quantum cryptography algorithm, a hybrid reduction algorithm based on the k-reduction algorithm and the Montgomery reduction algorithm is used for processing; among them, the hybrid reduction algorithm is specifically: for the first N - 1 layers of operations in the number-theoretic transform, the k-reduction algorithm is used for reduction, and starting from the Nth layer, the Montgomery reduction algorithm is used; where N is a positive integer, and the value of N is the layer where overflow first occurs during the transform;

[0019] Based on the processing procedures of the modular reduction operation and the polynomial multiplication operation, the implementation of the post-quantum cryptography algorithm on the resource-constrained processor is completed.

[0020] According to the third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program running on the memory. When the processor executes the program, it implements the method for implementing a post-quantum cryptography algorithm for a resource-constrained processor as described above.

[0021] According to the fourth aspect of the embodiments of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the method for implementing a post-quantum cryptography algorithm for a resource-constrained processor as described above.

[0022] Compared with the prior art, the beneficial effects of the present disclosure are:

[0023] (1) The present disclosure provides a method and system for implementing a post-quantum cryptographic algorithm for resource-constrained processors. In the resource-constrained processors, the NTT method is adopted to calculate the polynomial product in the post-quantum cryptographic algorithm. Meanwhile, a hybrid reduction algorithm combining the Montgomery reduction algorithm and the k-reduction algorithm is proposed to implement the modular reduction operation in the post-quantum cryptographic algorithm, effectively reducing the consumption of computing resources and storage resources, and realizing the effective implementation of the lattice-based post-quantum cryptographic algorithm on resource-constrained and performance-limited processors.

[0024] (2) The speeds of key-generation, encryption, and decryption of the Saber algorithm implemented based on the solution of the present disclosure are respectively 3 times, 3.2 times, and 3.4 times higher than the current optimal implementation on ARM Cortex-M0 / M0+.

[0025] Advantages of additional aspects of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. Brief Description of the Drawings

[0026] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.

[0027] Figure 1 It is a schematic diagram of the hybrid reduction algorithm described in the embodiments of the present disclosure. Detailed Embodiments

[0028] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0029] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used in the solutions described in this embodiment have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0031] In the case of no conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.

[0032] Example 1:

[0033] The purpose of this embodiment is to provide a method for implementing a post - quantum cryptographic algorithm for resource - constrained processors.

[0034] A method for implementing a post - quantum cryptographic algorithm for resource - constrained processors includes processing modular reduction operations and polynomial multiplication operations. Among them,

[0035] For the polynomial multiplication operation in the post - quantum cryptographic algorithm, a number - theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing;

[0036] For the modular reduction operation in the post - quantum cryptographic algorithm, a hybrid reduction algorithm based on the k - reduction algorithm and the Montgomery reduction algorithm is used for processing; among them, the hybrid reduction algorithm is specifically: for the first N - 1 layers of operations in the number - theoretic transform, the k - reduction algorithm is used for reduction, and starting from the Nth layer, the Montgomery reduction algorithm is used; where N is a positive integer, and the value of N is the layer where overflow first occurs during the transformation;

[0037] Based on the processing process of the modular reduction operation and the polynomial multiplication operation, the implementation of the post - quantum cryptographic algorithm on the resource - constrained processor is completed.

[0038] Furthermore, the resource - constrained processor is a 32 - bit platform of ARM Cortex - M0 / M0 +, which has the following constraints: the processor does not have high - performance multiplication and can only calculate the product of 32 bits × 32 bits to obtain a 32 - bit result; the processor does not have multiply - add and multiply - subtract multi - operation consecutive instructions; the processor does not have single - instruction multiple - data operation functions.

[0039] Furthermore, the ARM Cortex - M0 / M0 + processor includes 16 32 - bit registers, among which, R0 - R12 are general - purpose registers, and R13 - R15 are special registers.

[0040] Furthermore, for the polynomial multiplication operation in the post - quantum cryptographic algorithm, a number - theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing, among which, the forward number - theoretic transform uses the CT - butterfly transform, and the inverse number - theoretic transform uses the GS - butterfly transform.

[0041] Furthermore, the input coefficients of the forward number - theoretic transform are in normal order. After the CT - butterfly transform, the coefficient order is bit - reversed; the input coefficients of the inverse number - theoretic transform are bit - reversed, and after the GS - butterfly transform, the coefficient order is normal.

[0042] Furthermore, the post - quantum cryptographic algorithm uses the lattice - based post - quantum cryptographic algorithm Saber algorithm.

[0043] Further, the post-quantum cryptographic algorithm may also be Kyber, NTRU, Dilithium, and Falcon algorithms.

[0044] Specifically, for the sake of easy understanding, the solution described in this embodiment will be described in detail from the perspective of specific implementation as follows:

[0045] The solution described in this embodiment takes the case of using NTT to calculate the polynomial product of Saber on ARM Cortex-M0 / M0+ as an example for research. The solution can be applied to any other lattice-based cryptographic algorithm system, such as Kyber, NTRU, Dilithium, and Falcon, etc.

[0046] On the one hand, based on the instruction set on ARM Cortex-M0 / M0+, we compared different modular reduction algorithms, including Montgomery Reduction, Barrett Reduction, and k-Reduction. Based on the comparison results, we selected a hybrid reduction algorithm scheme that combines Montgomery reduction and k-reduction to handle modular reduction operations on resource-constrained processors.

[0047] On the other hand, in order to comply with the characteristics of the 32-bit registers and instruction set on ARM Cortex-M0 / M0+, we use the multi-moduli NTT method, which can ensure that the intermediate values and results of the calculation are not greater than 32 bits, and can complete all operations involved in the Saber algorithm without error in the case of the lack of long multiplication instructions on ARM Cortex-M0 / M0+. The specific process is as follows:

[0048] First of all, using NTT to calculate the product of polynomials in the Saber algorithm requires changing the modulus ring. The new modulus needs to meet the basic requirements of NTT - it is a prime number. In addition, for the accuracy of the calculation results, in the signed implementation, when calculating the product of two polynomials, the new modulus should be greater than q = 2 13 , and the value of μ changes with the changes of Saber, Lightsaber, and Firesaber.

[0049] Among them, the multi-moduli NTT can be regarded as applying the Chinese Remainder Theorem (CRT) to the NTT. The multi-moduli NTT divides a large modulus into the product of several small moduli. After calculating the polynomial product on these small rings respectively, the result on the large ring is obtained by using the CRT. For example, when calculating the product of two polynomials on Saber using the NTT, the modulus should be greater than Using the multi-moduli NTT technology, two small prime numbers q1 = 3329 and q2 = 12289 are selected, (q1, q2) = 1, and q1q2 > 2 22 , first complete the product of polynomials on the rings 3329 and 12289 respectively, then use the CRT to calculate the result of the polynomial product on the ring q1q2, and finally obtain the target result on the ring 2 13 .

[0050] Secondly, compare the efficiencies of three different reduction algorithms on the Cortex-M0 / M0+ processor. Specifically:

[0051] (1) Montgomery reduction is a fast algorithm for calculating modular reduction. To adapt to the instruction set in the ARM Cortex-M0 / M0+, we improved the signed Montgomery reduction, as shown in Algorithm 1.

[0052]

[0053] Using the Cortex-M0 / M0+ instructions to implement the modified signed Montgomery reduction, as shown in Algorithm 2, a total of 8 instructions are required.

[0054]

[0055] (2) Barrett reduction is also a fast algorithm for calculating modular reduction, as shown in Algorithm 3.

[0056]

[0057] Using the Cortex-M0 / M0+ instructions to implement the Barrett reduction, as shown in Algorithm 4, a total of 7 instructions are required.

[0058]

[0059] (3) k-reduction is also a fast algorithm for calculating modular reduction, as shown in Algorithm 5.

[0060]

[0061] The k-reduction implemented using Cortex-M0 / M0+ instructions, as shown in Algorithm 6, requires a total of 6 instructions.

[0062]

[0063] A comparison of the implementations of these three reduction algorithms on Cortex-M0 / M0+ is shown in Table 1.

[0064] Table 1: Comparison of several reduction algorithms

[0065]

[0066] a In Barrett reduction, R is not related to the word size and is not necessarily 2 16 or 2 32 . Its value is based on the limiting conditions in Barrett reduction and the size of the input C.

[0067] b The numbers in parentheses represent the number of LDR or MOVS instructions used to place the target data into the correct register. The MOVS instruction can only move an 8-bit immediate value into the register. When a large immediate value (such as the modulus or its inverse) needs to be loaded in Montgomery reduction or Barrett reduction, the LDR instruction is required.

[0068] In addition, the result in Montgomery reduction contains a factor R -1 , and we can multiply the rotation factor by R in advance. In this way, after the product of the coefficient and the processed rotation factor undergoes Montgomery reduction, it is not affected by R -1 .

[0069] Similarly, the result in k-reduction also contains a factor k, and we multiply the rotation factor by k in advance -1 , and similarly, the factor k in the product of the coefficient and the rotation factor can be eliminated when passing through k-reduction.

[0070] The limiting conditions of Barrett reduction make it not applicable to Cortex-M0 / M0+. In Barrett reduction, in order to calculate the result correctly, the limiting conditions need to be satisfied Implementing Barrett reduction in Cortex - M0 / M0+ also requires that all intermediate values and results be less than 32 bits. Taking modulus 3329 as an example, when calculating the product of two numbers modulo 3329, then Cm < 3329×3329×5039, and 3329×3329×5039 > 2 32 . Under the constraint conditions of correctness, the intermediate values may not meet the 32 - bit constraint conditions. Therefore, we do not use Barrett reduction to directly reduce the result of modular multiplication of two elements in the ring.

[0071] To obtain the fastest speed possible, we should try to choose k - reduction to calculate modular multiplication in the ring. However, if we only use k - reduction to calculate modular multiplication on Cortex - M0 / M0+, there is a risk of overflow. Because Cortex - M0 / M0+ has 32 - bit registers, and the output value of k - reduction is not within a fixed range. It increases as the input increases. When q = k·2 m +1, the input C ≤ (k·2 m ) 2 ≤ 2 32 , and the reduced result after k - reduction is in . If k - reduction is still used in the subsequent modular multiplication, the result will be reduced to and overflow may occur. To avoid data overflow, we need to add additional reduction algorithms. There are three ways to add additional reduction algorithms. One is to add k - reduction as an additional reduction algorithm (Case 1); the second is to use Montgomery reduction instead of k - reduction for the reduction algorithm in the ring (Case 2); the third is to add Barrett reduction as an additional reduction algorithm (Case 3).

[0072] We use the Cooley-Tukey butterfly to calculate the forward NTT and the Gentleman-Sande butterfly to calculate the inverse NTT. The combined use of the Cooley-Tukey butterfly and the Gentleman-Sande butterfly can avoid the "bit-reverse" operation of the coefficients. The input coefficients of the forward NTT are in normal order. After the Cooley-Tukey butterfly transformation, the coefficient order is bit-reversed. At this time, the input coefficients of the inverse NTT are bit-reversed. After the Gentleman-Sande butterfly transformation, the coefficient order is normal. Considering the forward NTT and the inverse NTT as a whole, the order of the input coefficients is the same as the order of the output coefficients.

[0073] Furthermore, the following details the computational requirements for implementing the multiplication of polynomials over the ring 3329 using the above three methods respectively on Cortex-M0 / M0+.

[0074] The k-reduction over the ring 3329 on Cortex-M0 / M0+ can be further optimized. As shown in Algorithm 7, the optimized k-reduction only requires 4+(1) instructions on Cortex-M0 / M0+.

[0075]

[0076] Among them, for the forward NTT transformation, 3329 = 13 × 2 8 +1:

[0077] The first layer: The coefficients satisfy 0 ≤ a < 2 13 , and the rotation factor ξ = 133. The Cooley-Tukey butterfly transformation is as follows:

[0078] Multiplication: 0 ≤ C = aξ < 2 13 ×133 < 2 21

[0079] k-reduction: C = C0 + C12 8 , 0 ≤ C0 < 2 8 , 0 ≤ C1 < 2 13 , -2 13 <13C0 - C1 < 2 12

[0080] Addition (subtraction): The result range is (-2 14 , 2 13 )

[0081] The second layer: The result of the first layer is the input of the second layer. The maximum rotation factor ξ = 1991. The Cooley-Tukey butterfly transformation is as follows:

[0082] Multiplication: -225 <-2 14 ×1991 < C = aξ < 2 13 ×1991 < 2 25

[0083] k - reduction: C = C0 + C12 8 , 0 ≤ C0 < 2 8 , -2 17 ≤ C1 < 2 16 , -2 16 <13C0 - C1 < 2 12 +2 17

[0084] Addition (subtraction): The result range is (-2 18 , 2 18 )

[0085] The third layer: The result of the second layer is the input of the third layer. The maximum rotation factor ξ = 2764, and the CT - butterfly transformation is as follows:

[0086] Multiplication: -2 30 < -2 18 ×2764 < C = aξ < 2 18 ×2764 < 2 30

[0087] k - reduction: C = C0 + C12 8 , 0 ≤ C0 < 2 8 , -2 22 ≤ C1 < 2 22 , -2 22 <13C0 - C1 < 2 12 +2 22

[0088] Addition (subtraction): The result range is (-2 23 , 2 23 )

[0089] In the multiplication C = aξ of the fourth - layer forward NTT transform, it will exceed 32 bits when ξ < 3329, and data overflow occurs at this time. At this time, additional reduction algorithms are needed to solve the data - overflow problem. Next, the remaining forward NTT transforms will be introduced separately according to the above three cases:

[0090] (1) Case 1:

[0091] We always use k-reduction to reduce data. As can be seen from the above, the calculation result of the third layer will affect the product of the coefficients and rotation factors in the subsequent fourth layer, and it is very likely that this product will overflow. To ensure the smooth progress of the subsequent NTT transformation, we need to perform additional reduction on the data of the third layer to keep them within a smaller range. Therefore, it is necessary to perform k-reduction on the results of the third layer again. At this time, the input of k-reduction is (-2 23 , 2 23 ), so 0 ≤ C0 < 2 8 , -2 15 ≤ C1 < 2 15 , -2 15 < 13C0 - C1 < 2 12 + 2 15 . Therefore, the range of the output of the third layer after additional k-reduction is (-2 15 , 2 16 ).

[0092] After the coefficients and rotation factors in the subsequent fourth layer are multiplied, k-reduction is applied to reduce the product, -2 27 < C = aξ < 2 28 , 0 ≤ C0 < 2 8 , -2 20 ≤ C1 < 2 12 + 2 19 , -2 20 < 13C0 - C1 < 2 12 + 2 19 . Next, addition (subtraction) is performed, and finally the result of the fourth layer is (-2 20 - 2 15 , 2 20 ).

[0093] The strategy of the fifth layer is the same as that of the third layer. After the CT-butterfly transformation, an additional k-reduction needs to be performed on the result. First, calculate the multiplication |C = aξ| < 2 32 . Next, perform k-reduction on the product, 0 ≤ C0 < 2 8 , -2 24 ≤ C1 < 2 24 , -2 24 < 13C0 - C1 < 2 12 + 2 24 . Therefore, the range of the result after addition (subtraction) is (-2 25 , 2 25 ). Finally, an additional k-reduction is performed to reduce the result again, |C| < 2 25,0≤C0<2 8 ,-2 17 ≤C1<2 17 ,-2 17 <13C0-C1<2 12 +2 17 Therefore, the result of layer 5 is (-2 17 ,2 12 +2 17 ).

[0094] In the next sixth layer, additional k-reduction is also required after the CT-butterfly transformation to ensure that the intermediate values ​​of the seventh layer do not overflow. In the sixth layer, after the coefficients and the rotation factors are multiplied, k-reduction is performed, -2 29 <C=aξ<2 29 ,0≤C0<2 8 ,-2 21 ≤C1<2 21 ,-2 21 <13C0-C1<2 12 +2 21 After addition (subtraction), add an additional k-reduction, -2 22 <C<2 22 ,0≤C0<2 8 ,-2 14 ≤C1<2 14 ,-2 14 <13C0-C1<2 12 +2 14 , so the range of the sixth layer result is (-2 14 ,2 12 +2 14 ).

[0095] For the seventh layer, the conventional CT-butterfly transform is performed. The product of the coefficients and the rotation factors is reduced using k-reduction, -2 16 <C=aξ<2 27 ,0≤C0<2 8 ,-2 18 ≤C1<2 19 ,-2 19 <13C0-C1<2 12 +2 18 , after completing the addition (subtraction), the result range is (-2 20 ,2 19 ).

[0096] (2) Scenario 2:

[0097] We use a combination of k-reduction and Montgomery reduction for modular reduction. The range of the signed Montgomery reduction result in the ring 3329 is (-3329, 3329). In the forward NTT transform, the first, second, and third layers all use k-reduction as described above. After the transformation of the first three layers, the range of the intermediate value is (-2 23 , 2 23 ).

[0098] In the subsequent fourth-layer transformation, the product |C = aξ| < 2 23+12 will exceed 32 bits. Therefore, to avoid overflow in the fourth layer, we will use Montgomery reduction in the third layer to reduce the values in the CT-butterfly transform instead of k-reduction. After Montgomery reduction and addition (subtraction) in the third layer, the result will be in the range (-2 18 - 3329, 2 18 + 3329).

[0099] In the subsequent fourth, fifth, sixth, and seventh layers, Montgomery reduction is used to reduce the product of the coefficient and the rotation factor. After completing the 7-layer forward NTT transform, the result is (-2 18 - 3329.5, 2 18 + 3329.5) ∈ (-2 19 , 2 19 ).

[0100] (3) Scenario three:

[0101] We use a combination of k-reduction and Barrett reduction for modular reduction. The range of the Barrett reduction result in the ring 3329 is (0, 6658). In the forward NTT transform, after the CT-butterfly transform in the third layer, an additional Barrett reduction is added to reduce the result, and the range of the reduced result is (0, 6658).

[0102] In the fourth layer, k-reduction is still used for reduction. The product of the coefficient and the rotation factor |C = aξ| < 6658 × 3329 < 2 25 , 0 ≤ C0 < 2 8 , |C1| < 2 17 , -2 17 <13C0 - C1 < 2 12+2 17 . After addition (subtraction), the range of the result of the fourth layer is (-2 18 , 2 18 ). In the fifth layer, k-reduction is also used for modular reduction, |C = aξ| < 2 30 , 0 ≤ C0 < 2 8 , |C1| < 2 22 , -2 22 < 13C0 - C1 < 2 12 +2 22 . Therefore, the result after addition (subtraction) is (-2 23 , 2 23 ).

[0103] To prevent overflow errors in the intermediate values for the next layer's calculation, Barrett reduction needs to be additionally applied to the result of the fifth layer. Finally, the result of the fifth layer is (0, 6658). Then, the calculation process of the sixth layer is the same as that of the fourth layer, and only k-reduction is needed to perform modular reduction on the product. The result of the sixth layer is also (-2 18 , 2 18 ). In the calculation of the next seventh layer, only k-reduction is used to perform modular reduction on the product of the coefficient and the rotation factor. The result of the seventh layer is (-2 23 , 2 23 ). Since only 7 layers of NTT transformation are performed in the ring 3329, this result of the seventh layer will not affect the calculation of the next layer, so no additional reduction algorithm needs to be added to the result of the seventh layer.

[0104] In summary, for the forward NTT transformation in Ring 3329 in Case 1, 13 × 128 × k-reduction = 13 × 128 × 5 = 8320 instructions are required. In Case 2, 2 × 128 × k-reduction + 5 × 128 × Montgomery = 2 × 128 × 5 + 5 × 128 × 8 = 6400 instructions are needed to complete the same operation. In Case 3, 7 × 128 × k-reduction + 2 × 256 × Barrett = 7 × 128 × 5 + 2 × 256 × 7 = 8064 instructions are required.

[0105] Analyze the inverse NTT by the same method. In case one, it requires 13×128×k-reduction = 13×128×5 = 8320 instructions. In case two, it requires 1×128×k-reduction + 6×128×Montgomery = 1×128×5 + 6×128×8 = 6784 instructions. In case three, it requires 7×128×k-reduction + 3×256×Barrett = 7×128×5 + 3×256×7 = 9856 instructions.

[0106] Furthermore, Table 2 and Table 3 respectively show the comparison of different instructions used in the above three cases on the ring 3329. According to the above analysis, we adopt the strategy of case two, that is, the combination of k-reduction and Montgomery reduction, to perform the forward NTT transformation and inverse NTT transformation of polynomials on the ring 3329, as Figure 1 shown.

[0107] Table 2: Comparison of Forward NTT in Ring 3329 Completed by Different Reduction Algorithms

[0108]

[0109] Table 3: Comparison of Inverse NTT in Ring 3329 Completed by Different Reduction Algorithms

[0110]

[0111] Next, the computational complexity required to implement the product of polynomials on the ring 12289 using the above three methods on Cortex-M0 / M0+ will be described in detail.

[0112] For the forward NTT transformation, 12289 = 3×2 12 +1:

[0113] The first layer: The coefficient a < 2 13 , and the rotation factor ξ = 493. The CT-butterfly transformation is as follows:

[0114] Multiplication: 0 ≤ C = aξ < 2 13 ×493 < 2 22

[0115] k-reduction: C = C0 + C1·2 12 , 0 ≤ C0 < 2 12 , 0 ≤ C1 < 2 10 , -2 10 < 3C0 - C1 < 2 15

[0116] Addition (subtraction): The result range is (-3.2 12, 3.2 12 +12289) ∈ (-2 14 , 2 15 )

[0117] Layer 2: The input value is the calculation result of Layer 1. The maximum rotation factor ξ = 5444. The CT-butterfly transformation is as follows:

[0118] Multiplication: -2 26 < -3.2 12 ×5444 < C = aξ < (3.2 12 +12289) × 5444 < 2 27

[0119] k-reduction: C = C0 + C1·2 12 , 0 ≤ C0 < 2 12 , -2 14 ≤ C1 < 2 15 , -2 15 < 3C0 - C1 < 2 15

[0120] Addition (subtraction): The result range is (-3.2 12 -2 15 , 3.2 12 +12289 + 2 15 ) ∈ (-2 16 , 2 16 )

[0121] Layer 3: The input value is the calculation result of Layer 2. The maximum rotation factor ξ = 4337. The CT-butterfly transformation is as follows:

[0122] Multiplication: -2 28 < (-3.2 12 -2 15 ) × 4337 < C = aξ < (3.2 12 +12289 + 2 15 +2 14 +2 16 ) × 11885 < 2 31

[0123] k-reduction:

[0124] C = C0 + C1·2 12 , 0 ≤ C0 < 2 12 , -2 16 ≤ C1 < 2 16 , -2 16 < 3C0 - C1 < 2 14 +2 16

[0125] Addition (subtraction): Result range

[0126] (-3.2 12 -2 15 -2 14 -2 16 ,3.2 12 +12289+2 15 +2 14 +2 16 )∈(-2 17 ,2 18 )

[0127] Fourth layer: The input value is the calculation result of the third layer, the maximum rotation factor ξ = 11885, and the CT - butterfly transformation is as follows:

[0128] Multiplication:

[0129] -2 31 <(-3.2 12 -2 15 -2 14 -2 16 )×11885<C = aξ

[0130] <(3.2 12 +12289+2 15 +2 14 +2 16 )×11885<2 31

[0131] k - reduction:

[0132] C = C0 + C1·2 12 ,0≤C0<2 12 ,-2 19 ≤C1<2 19 ,-2 19 <3C0 - C1<2 14 +2 19

[0133] Addition (subtraction): Result range (-3.2 12 -2 17 -2 19 ,3.2 12 +12289+2 17 +2 19 )∈(-2 20 ,2 20 )

[0134] Next, directly perform the calculation of the fifth-layer forward NTT. There is a risk of overflow on Cortex-m0 / m0+ for the intermediate values, and the product of the coefficient and the rotation factor may exceed 32 bits. The following explains how to perform the subsequent forward NTT transformation in three cases specifically:

[0135] (1) Case 1:

[0136] Add an additional k-reduction to the calculation result of the above fourth layer, C = C0 + C1·2 12 , 0 ≤ C0 < 2 12 , -2 8 ≤ C1 < 2 8 , -2 8 <3C0 - C1 < 2 14 +2 8 , so the result of the fourth layer is (-2 8 , 2 14 +2 8 ).

[0137] In the fifth layer, the product of the coefficient and the rotation factor -2 22 <C = aξ < 2 29 , use k-reduction for reduction, C = C0 + C1·2 12 , 0 ≤ C0 < 2 12 , -2 10 ≤ C1 < 2 17 , -2 17 <3C0 - C1 < 2 14 +2 10 , and then through addition (subtraction) operations, the result is (-2 8 -2 17 , 2 14 +2 8 +2 17 ).

[0138] In the sixth layer, -2 31 <C = aξ < 2 31 , after reduction by k-reduction, C = C0 + C1·2 12 , 0 ≤ C0 < 2 12 , -2 19 ≤ C1 < 2 19 , -2 19 <3C0 - C1 < 2 14 +2 19 , and the result of the addition (subtraction) operation is (-2 20 , 2 20 ). An additional k-reduction needs to be added to this result of the sixth layer to ensure the smooth progress of the next layer's calculation. At this time, C = C0 + C1·212 , 0 ≤ C0 < 2 12 , -2 8 ≤ C1 < 2 8 , -2 8 < 3C0 - C1 < 2 14 +2 8 . Therefore, the result of the sixth layer is (-2 8 , 2 14 +2 8 ).

[0139] The calculation of the seventh layer is the same as that of the fifth layer. After one pass through multiplication, k-reduction, and addition (subtraction), the result of the seventh layer is (-2 8 -2 17 , 2 14 +2 8 +2 17 ). Compared with the calculation of the sixth layer, the eighth layer does not need to apply k-reduction to the result after CT-butterfly transformation again. Therefore, the calculation result of the eighth layer is (-2 20 , 2 20 ).

[0140] (2) Case two:

[0141] The modular reduction algorithm in the CT-butterfly transformation of the fourth, fifth, sixth, seventh, and eighth layers all uses Montgomery reduction. Since on the ring 12289, the result of signed Montgomery reduction is (-12289, 12289), after completing the 8-layer forward NTT transformation on the ring 12289, the result is (-2 17 -5 × 12289, 2 18 +5 × 12289) ∈ (-2 18 , 2 19 ).

[0142] (3) Case three:

[0143] Add additional Barrett reduction to the results of the CT-butterfly transformation of the fourth and sixth layers. In addition, k-reduction is used for the modular reduction in the CT-butterfly transformation of each layer.

[0144] In summary, for the forward NTT transform in Ring 12289 in Case 1, it requires 8×128×k - reduction + 2×256×k - reduction, and in total it needs 8×128×6 + 2×256×6 = 9216 instructions. In Case 2, to complete the forward NTT calculation, it requires 3×128×k - reduction + 5×128×Montgomery, and in total it needs 3×128×6 + 5×128×8 = 7424 instructions. While in Case 3, it needs 8×128×k - reduction + 2×256×Barrett, and in total it needs 8×128×6 + 2×256×7 = 9728 instructions.

[0145] The same three strategies are adopted to analyze the inverse NTT on Ring 12289. In Case 1, calculating the inverse NTT requires 8×128×k - reduction + 3×256×k - reduction, that is 8×128×6 + 3×256×6 = 10752 instructions.

[0146] In Case 2, it requires 1×128×k - reduction + 7×128×Montgomery, which is 1×128×6 + 7×128×8 = 7936 instructions in total. In Case 3, to complete the calculation, it requires 8×128×k - reduction + 3×256×Barrett, and in total it needs 8×128×6 + 3×256×7 = 11520 instructions.

[0147] Furthermore, Table 4 and Table 5 respectively illustrate the comparison of different instructions used in the above three cases on Ring 12289. According to the above analysis, we adopt the strategy of Case 2, that is, the hybrid reduction method combining k - reduction and Montgomery reduction, to perform the forward NTT transform and inverse NTT transform of polynomials on Ring 12289, as Figure 1 shown.

[0148] Table 4: Comparison of Completing Forward NTT in Ring 12289 with Different Reduction Algorithms

[0149]

[0150] Table 5: Comparison of Completing Inverse NTT in Ring 12289 with Different Reduction Algorithms

[0151]

[0152] Finally, merge the NTT layer: Cortex-M0 / M0+ has 16 32-bit registers, where R0-R12 are general-purpose registers and R13-R15 are special registers, and among them, R0-R7 are called low registers. Some instructions in the Cortex-M0 / M0+ instruction set, such as MUL, can only process low registers; some instructions, such as MOV, can process almost all registers.

[0153] Due to the limitations of the number of registers and the characteristics of the instruction set in Cortex-M0 / M0+, we calculate two layers of NTT in one merge on Cortex-M0 / M0+. That is, each time 4 coefficients are loaded into the registers, and after completing the calculation of the forward NTT and inverse NTT of the two layers related to them, these 4 results are then stored in the memory, which can reduce the number of times of coefficient loading and storage. The 7-layer NTT on ring 3329 is decomposed into: 2 layers - 2 layers - 2 layers - 1 layer, and the 8-layer NTT on ring 12289 is also divided into 4 blocks: 2 layers - 2 layers - 2 layers - 2 layers.

[0154] After completing the 7-layer (8-layer) inverse NTT, it is also necessary to multiply by the coefficient 2 -7 (2 -8 ) to obtain the target result. Merging the operation of multiplying this coefficient into the point multiplication stage of the polynomial in the NTT domain can also save the operations of coefficient loading and storage.

[0155] Furthermore, the effectiveness of the solution described in this embodiment is demonstrated through experiments as follows:

[0156] We use a combination of k-reduction and Montgomery reduction to perform reduction operations on the ring, and use multi-moduli NTT to calculate the product of polynomials in Saber. Compared with the results of Cortex-M0 / M0+ listed on the Saber algorithm homepage, our results have a significant improvement. Among them, the speed in the key generation stage is 3 times that of the reference value, the speed in the encryption stage is 3.2 times that of the reference value, and the speed in the decryption stage is 3.4 times that of the reference value. The comparison results are shown in Table 6.

[0157] Table 6: The speed of Saber on ARM Cortex-M0 / M0+

[0158]

[0159] Embodiment 2:

[0160] The purpose of this embodiment is to provide a post-quantum cryptographic algorithm implementation system for resource-constrained processors.

[0161] A post-quantum cryptography algorithm implementation system for resource-constrained processors, including an ARM Cortex-M0 / M0+ processor and a post-quantum cryptography algorithm executed on the processor. Among them, the implementation of the post-quantum cryptography algorithm specifically adopts the following steps:

[0162] For the polynomial multiplication operation in the post-quantum cryptography algorithm, a number-theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing;

[0163] For the modular reduction operation in the post-quantum cryptography algorithm, a hybrid reduction algorithm based on the k-reduction algorithm and the Montgomery reduction algorithm is used for processing; among them, the hybrid reduction algorithm is specifically: for the first N-1 layers of operations in the number-theoretic transform, the k-reduction algorithm is used for reduction, and starting from the Nth layer, the Montgomery reduction algorithm is used; where N is a positive integer, and the value of N is the layer where overflow first occurs during the transformation;

[0164] Based on the processing processes of the modular reduction operation and the polynomial multiplication operation, the implementation of the post-quantum cryptography algorithm on the resource-constrained processor is completed.

[0165] Further, the system in this embodiment corresponds to the method in Embodiment 1, and its technical details have been described in detail in Embodiment 1, so they will not be repeated here.

[0166] In more embodiments, there is also provided:

[0167] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be repeated here.

[0168] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0169] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.

[0170] A computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0171] The method in the first embodiment can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0172] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0173] The method and system for implementing a post-quantum cryptographic algorithm for a resource-constrained processor provided in the above embodiment can be realized and have broad application prospects.

[0174] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for implementing a post - quantum cryptographic algorithm for resource - constrained processors, characterized in that, Including the processing of modular reduction operations and polynomial multiplication operations, where, For the polynomial multiplication operation in the post-quantum cryptographic algorithm, the number-theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing; For the modular reduction operation in the post-quantum cryptographic algorithm, a hybrid reduction algorithm based on the k-reduction algorithm and the Montgomery reduction algorithm is used for processing; where, the specific hybrid reduction algorithm is: for the first N - 2 layers of operations in the number-theoretic transform, the k-reduction algorithm is used for reduction, and starting from the N - 1th layer, the Montgomery reduction algorithm is used; where, N is a positive integer, and the value of N is the layer where overflow first occurs during the transform; Based on the processing process of the modular reduction operation and the polynomial multiplication operation, the implementation of the post-quantum cryptographic algorithm on a resource-constrained processor is completed.

2. The method for implementing a post-quantum cryptographic algorithm for a resource-constrained processor according to claim 1, wherein, The resource-constrained processor is a 32-bit platform of ARM Cortex-M0 / M0+, and it has the following constraints: the processor does not have high-performance multiplication, and can only calculate the product of 32 bits × 32 bits to obtain a 32-bit result; the processor does not have multiply-add and multiply-subtract multi-operation consecutive instructions; the processor does not have single-instruction multiple-data operation functions.

3. The method for implementing a post-quantum cryptographic algorithm for a resource-constrained processor according to claim 2, wherein, The ARM Cortex-M0 / M0+ processor includes 16 32-bit registers, where, R0 - R12 are general-purpose registers, and R13 - R15 are special registers.

4. The post-quantum cryptographic algorithm implementation method for a resource-constrained processor according to claim 1, wherein For the polynomial multiplication operation in the post-quantum cryptographic algorithm, the number-theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing, where, the forward number-theoretic transform uses the CT-butterfly transform, and the inverse number-theoretic transform uses the GS-butterfly transform.

5. The method for implementing a post-quantum cryptographic algorithm for a resource-constrained processor according to claim 4, wherein The input coefficients of the forward number-theoretic transform are in normal order. After the CT-butterfly transform, the coefficient order is bit-reversed; the input coefficients of the inverse number-theoretic transform are bit-reversed, and after the GS-butterfly transform, the coefficient order is normal.

6. The post-quantum cryptographic algorithm implementation method for a resource-constrained processor according to claim 1, wherein, The post-quantum cryptographic algorithm uses the lattice-based post-quantum cryptographic algorithm Saber algorithm.

7. The post-quantum cryptographic algorithm implementation method for a resource-constrained processor according to claim 1, wherein The post-quantum cryptographic algorithm can also be the Kyber, NTRU, Dili thium, and Falcon algorithms.

8. A post-quantum cryptographic algorithm implementation system for resource-constrained processors, characterized in that, Including an ARM Cortex-M0 / M0+ processor and a post-quantum cryptographic algorithm executed on the processor, where, the implementation of the post-quantum cryptographic algorithm specifically adopts the following steps: For the polynomial multiplication operation in the post-quantum cryptographic algorithm, the number-theoretic transform algorithm based on the Chinese Remainder Theorem is used for processing; For the modular reduction operation in the post-quantum cryptographic algorithm, a hybrid reduction algorithm based on the k-reduction algorithm and the Montgomery reduction algorithm is used for processing; where, the specific hybrid reduction algorithm is: for the first N - 2 layers of operations in the number-theoretic transform, the k-reduction algorithm is used for reduction, and starting from the N - 1th layer, the Montgomery reduction algorithm is used; where, N is a positive integer, and the value of N is the layer where overflow first occurs during the transform; Based on the processing process of the modular reduction operation and the polynomial multiplication operation, the implementation of the post-quantum cryptographic algorithm on a resource-constrained processor is completed.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and running thereon. When the processor executes the program, it implements a method for implementing a post-quantum cryptographic algorithm for a resource-constrained processor as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, it implements a method for implementing a post-quantum cryptographic algorithm for a resource-constrained processor as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Optimization method and device for lattice cipher polynomial multiplication operation based on number-theory transformation

    CN113972980A

  • Polynomial multiplication hardware implementation system suitable for lattice cryptographic algorithm

    CN114297571A