Endogenous security variable point NTT hardware implementation method based on KEY-RED algorithm

By adopting the endogenous safe variable point hardware architecture based on KEY-RED mode multiplication algorithm in the NTT algorithm hardware implementation, the problem of insufficient computing efficiency and security in the hardware implementation of NTT algorithm is solved, efficient and secure NTT calculation is realized, and it is suitable for computing scenarios with different points.

CN119945670APending Publication Date: 2025-05-06TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017451.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing NTT algorithm hardware implementations have shortcomings in terms of computing efficiency and security, especially when facing a large number of large integer polynomial multiplication calculations and physical attacks (such as SCA attacks).

Method used

The endogenous safe variable point NTT hardware implementation method based on the KEY-RED mode multiplication algorithm is adopted, and the hardware architecture including a rotation factor storage module, a delay feedback module, a random delay module, a data input module, a control module and a data allocator DEMUX is designed. Through the improved KEY-RED algorithm and a random delay unit, the efficiency of mode reduction operation and anti-time attack capability are improved.

Benefits of technology

It improves the computing efficiency and security of the NTT algorithm hardware implementation, reduces the consumption of hardware resources, enhances the resistance to SCA attacks, and realizes NTT calculation of variable points, which is suitable for computing scenarios with different points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945670A_ABST
    Figure CN119945670A_ABST
Patent Text Reader

Abstract

The invention relates to the field of post quantum cryptography and hardware acceleration, and provides an endogenous security variable point (NTT) hardware implementation method based on a KEY-RED modular multiplication algorithm, which comprises the following steps: firstly, designing a twiddle factor storage module, a delay feedback module, a random delay module, a data input module and a control module required by an NTT hardware architecture; then, on the basis of the designed module, a variable point NTT hardware architecture is built; and finally, setting a mode and inputting data, and after comprehensive design, design implementation and bit stream file generation, burning codes on the FPGA. According to the hardware implementation method, the operation efficiency of an NTT algorithm is improved through an SDF pipeline structure and a more efficient KEY-RED algorithm, the NTT algorithm of a variable point is achieved through the characteristics of a control unit and the SDF structures, the algorithm implementation flexibility is improved, random time delay is added into any SDF structure through a random time delay unit to resist side channel attacks, and the algorithm implementation efficiency is improved. Therefore, the hardware implementation security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of post-quantum cryptography and hardware acceleration, and in particular to a method for hardware implementation of an intrinsically secure variable-point NTT based on a KEY-RED modular multiplication algorithm. Background Art

[0002] With the rapid development of quantum computing, modern encryption algorithms (including the popular public key encryption algorithm PKC) face security threats. The upcoming large-scale quantum computers may crack widely adopted cryptographic systems such as RSA and elliptic curve cryptography (ECC). Therefore, it is necessary to study the next generation of PKC, namely post-quantum cryptography (PQC) to solve this problem. At this stage, many effective post-quantum algorithms have been proposed, such as lattice-based encryption algorithms (such as NTRU, Kyber, LWE, etc.) and fully homomorphic encryption algorithms.

[0003] Since most PQC algorithms involve a large number of large integer polynomial multiplication calculations, which greatly reduces the efficiency of the algorithm implementation, and the fast number theory transformation algorithm can effectively speed up the efficiency of polynomial multiplication, most of the lattice cryptographic systems with high complexity involve the application of NTT. It can be seen that the NTT algorithm plays a key role in the implementation of post-quantum cryptographic algorithms, and with the transition of existing cryptographic systems to PQC, it is very important to improve the resistance to physical attacks (SCA attacks) in the implementation of algorithm encryption algorithms. As the underlying module of the encryption algorithm, how to efficiently implement the NTT algorithm with inherent security on hardware to accelerate the calculation of polynomials and ensure the security of the algorithm implementation process, while adapting to NTT calculation scenarios with different numbers of points, has become the focus of current research.

[0004] The NTT algorithm is an extension of the Fast Fourier Transform (FFT) algorithm in the finite field. It mainly uses the characteristics of the rotation factor to accelerate the calculation of polynomial multiplication. As the number of points calculated increases, the complexity of the algorithm implementation will also increase. Taking the most basic 8-point NTT algorithm as an example, if the in-situ calculation method is used, 12 butterfly units are required, and the NTT result is obtained after 3 levels of calculation. If the number of calculation points increases, then if the in-situ calculation method is used again, the number of butterfly units required will far exceed the number of calculation points, which will consume a lot of hardware resources and reduce calculation efficiency.

[0005] The implementation of the NTT algorithm can be divided into software implementation and hardware implementation. The implementation of the software focuses on reducing the time complexity and space complexity of the algorithm. Specific measures include optimizing the loop structure, reducing memory usage, and using efficient mathematical libraries. Since the software runs on a general-purpose processor, when implementing the NTT algorithm, the calculation accuracy can be improved by optimizing the program code. Taking the modular multiplication module in the NTT operation as an example, when the software implements the NTT algorithm, the modular multiplication operation can be implemented through a loop judgment statement. The software implementation of the NTT algorithm is usually applied to the cryptography library and applied to functions such as encryption, decryption, and key operations. In general, the software implementation of the NTT algorithm is mostly used in the development of the PQC library, and it optimizes the algorithm itself. The hardware implementation focuses on the optimization of hardware resources in order to achieve more functions and higher computing speeds under limited resources. The hardware implementation of the NTT algorithm usually adopts a low-power, low-complexity hardware architecture. Common hardware structures include sequential recursive structures, pipeline structures, parallel iterations, and hybrid structures. The hardware implementation of the NTT algorithm is usually applied to scenarios that require real-time computing or efficient encryption. The present invention focuses on the hardware implementation of the security variable point of the NTT algorithm on FPGA.

[0006] SCA attack refers to side-channel attack, which is an attack method that analyzes the physical leakage of computing devices (such as power consumption, electromagnetic radiation, processing time, sound, temperature change, etc.) to infer the internal operation process or obtain sensitive information. Unlike traditional cryptanalysis, side-channel attack does not directly target the mathematical weaknesses of encryption algorithms, but infers encryption keys or other private information by capturing and analyzing the physical characteristics generated when the hardware executes the algorithm.

[0007] In the complete encryption algorithm, SCA attacks mainly involve two vulnerable nodes, namely the point-wise multiplication (PWM) stage and the modular multiplication stage.

[0008] In the hardware implementation of the NTT algorithm, the core of its processing element (PE) is the modular operation unit, which includes modular addition, modular subtraction and modular multiplication operations. At present, the research on the hardware implementation of NTT mainly focuses on optimizing the modular reduction operation in PE. Common modular reduction algorithms include Montgomery Reduction and Barrett Reduction. These two algorithms have certain advantages in software implementation, but in terms of hardware implementation, they are not friendly to hardware implementation because they both require multiple multiplication and addition operations, and the preprocessing of Barrett modular reduction also consumes a lot of resources. Summary of the invention

[0009] The technical problem to be solved by the present invention is to improve the computational efficiency and security of the hardware implementation method of the NTT algorithm.

[0010] The present invention solves the above technical problems through the following technical means:

[0011] The present invention provides a method for hardware implementation of an intrinsically secure variable point NTT based on a KEY-RED algorithm, comprising the following steps:

[0012] S1. Design the modules required for the NTT hardware architecture, including: rotation factor storage module, delay feedback module, random delay module, data input module, control module and data distributor DEMUX;

[0013] S2. Building a variable-point NTT hardware architecture based on the modules obtained in step S1;

[0014] S3, setting the mode and input data, after synthesizing the design, implementing the design, and generating the bitstream file, the code is burned into the FPGA.

[0015] Furthermore, the design of the rotation factor storage module in step S1 includes the following specific steps:

[0016] (1) Use Python language to pre-calculate the rotation factor and generate a rotation factor file;

[0017] (2) Use System Verilog to design a 10-port ROM module that can store 1024 24-bit twiddle factors;

[0018] (3) The content in the rotation factor file is stored in the ROM module.

[0019] Furthermore, the pre-calculation of the rotation factors is specifically as follows:

[0020] Select modulus q = 12289 = 3 × 2 12 +1, the calculation formula of the rotation factor of the N-point NTT algorithm is: And the calculation formula of the modified rotation factor is:

[0021] Where 0≤k≤1023

[0022] A 24-bit rotation factor is calculated; wherein k is the index of the rotation factor, indicating the position of the current rotation factor in the transformation, and is a variable used to traverse all rotation factors.

[0023] Furthermore, the delayed feedback module described in step S1 is composed of an SDF structure and an analog multiplication module, and the SDF structure is composed of an analog addition and analog subtraction module, a delay module and a data selector; the input of the delayed feedback module is connected to the analog addition module and the analog subtraction module, and the output of the analog subtraction module is connected to the delay module after passing through the data selector to achieve a feedback effect.

[0024] Furthermore, the specific operation mode of the module addition and subtraction module is as follows:

[0025] For the modular addition module, determine whether the sum of the two numbers is greater than or equal to the modulus q. If the sum of the two numbers is greater than or equal to the modulus q, subtract q from the sum of the two numbers; if the sum of the two numbers is less than q, do nothing.

[0026] For the modular subtraction module, determine whether the difference between the two numbers is less than 0. If it is less than 0, add q to the difference between the two numbers; if the difference between the two numbers is greater than 0, no operation is required.

[0027] Furthermore, the delay module is implemented by a dual-port RAM and an address generation module, the RAM is used for temporary storage of data; the address generation module implements the delay operation through a counter, and changes the delay depth by changing the counting range of the counter.

[0028] Furthermore, the random delay module is designed in step S1 in the following manner:

[0029] A hardware-based pseudo-random number generator PRNG with data storage function is added after each of the selected levels of SDF structures to generate different random delay times; after the random delay, the data is output to the next level of SDF structure.

[0030] Furthermore, the specific operation mode of the analog multiplication module in step S1 is:

[0031] Based on the KEY-RED algorithm, the product of two numbers with a bit width of less than 48 bits is split twice into high bits and low bits; the modulus q is selected as 12289 = 3×2 12 +1, and use the special properties of the selected modulus to perform fast modular reduction, and complete the modular multiplication operation by keeping the output width at 24 bits through 2 shifts and 4 additions.

[0032] Furthermore, the design data input module in step S1 is specifically designed as follows:

[0033] Design a dual-port RAM that can store 1024 24-bit input data, and the data reading and writing of the two ports do not interfere with each other; when the number of input data points is less than 1024 points, the remaining bits are filled with 0; one port of the RAM is used for data input, and the other port is connected to a distributor DEMUX, and under the control of the control module, the input NTT point number is matched with the pipeline structure;

[0034] The specific operation mode of the control module in step S1 is:

[0035] According to the input control signal MODE, the data in the twiddle factor storage module is extracted to the corresponding modular multiplication unit.

[0036] Furthermore, the step S2 is specifically as follows:

[0037] (1) Build a 10-level SDF structure in a pipeline structure;

[0038] (2) The output of the data input module is connected to the input of the data distributor DEMUX, and the output port of the data distributor DEMUX is connected to the input end of each level of SDF respectively;

[0039] (3) The delay unit depths of the 1st-level SDF to the 10th-level SDF are set to 512D, 256D, …, 1D, respectively, where D is the unit delay time.

[0040] (4) The 10 ports of the rotation factor storage module are connected to the input of the control module, and after being allocated by the control module based on the control signal MODE, they are independently connected to the modular multiplication module of each level of the SDF structure;

[0041] (5) The result of the modular operation is output in the 10th level SDF structure.

[0042] Preferably, the variable-point NTT algorithm can be implemented by changing the content of the input data in the txt file and selecting the correct MODE.

[0043] The advantages of the present invention are:

[0044] The present invention introduces a more efficient modular multiplication algorithm and improves the algorithm. It can use fewer shift and addition operations to quickly implement modular reduction, reduce the occupation of hardware resources, and improve the operation speed. At the same time, compared with the general single-path feedback results, the architecture designed by the present invention is more flexible and unit reusable. A separate SDF unit can meet the calculation of different numbers of points, and realizes the pipeline NTT configuration of variable points. The randomness of the variable points further improves the security of the circuit. At the same time, the present invention improves the ability to resist timing attacks by introducing a random delay unit. In addition, the architecture of the present invention can be further expanded to achieve higher variable point NTT calculations by increasing the number of SDF units, and has a certain scalability. It can be well applied to NTT operation scenarios with different numbers of points. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of the hardware architecture of the variable-point NTT algorithm according to an embodiment of the present invention;

[0046] Figure 2 Schematic diagram of the structure of the analog multiplication module according to an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of a classic hardware architecture of a common NTT algorithm according to an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram of the basic SDF pipeline structure of an embodiment of the present invention;

[0049] Figure 5 Schematic diagram of a single SDF structure according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0051] Example 1

[0052] The present invention adopts an improved KEY-RED algorithm, which can efficiently complete the modular reduction operation through operations such as shifting, does not require pre-calculation operations, and can be well used for hardware implementation. The modular multiplication algorithm adopted divides the product of two numbers twice into high bits and low bits, and uses the special properties of the selected modulus to perform fast modular reduction. For a 24-bit two-bit multiplier input signal, the product C after multiplication can be a maximum of 48 bits, which can be effectively reduced to 24 bits. The modular multiplication algorithm is shown in the following table Algorithm 1: Improved KEY-RED modular multiplication algorithm number theory transformation:

[0053]

[0054] In terms of the hardware architecture of the NTT algorithm, the classic architecture of the common NTT computing unit is as follows: Figure 3 As shown in FIG. 1 , it mainly includes an operation logic unit module and a control logic unit module. The disadvantage of this architecture is that it can only be used for a single modulus and point number, and cannot implement a variable-point NTT algorithm. The embodiment of the present invention uses an improved single-path delayed feedback (R2-SDF) structure to implement a variable-point NTT operation. A basic SDF pipeline structure for implementing an 8-point NTT algorithm is shown in FIG. Figure 4 As shown in FIG. 1 , the structure mainly controls the flow of data through a delay unit and two data selectors, separates the two parts of the butterfly operation data, and then performs the butterfly operation. The pipeline structure design greatly improves the operation efficiency. This embodiment improves the SDF structure, adds a random delay unit, uses the improved KEY-RED modular reduction algorithm for the modular multiplication module, and controls the selection of rotation factors and the allocation of operation points through a control unit. A variable point NTT implementation method with resistance to side channel attack SCA that can achieve 1024, 512, 256 and lower points is proposed. Its hardware architecture is as follows Figure 1 First, prepare a computer that can write System Verilog code, an FPGA development board, and a Python script file; then start designing each module and building the hardware architecture. The specific steps are as follows:

[0055] S1. Design the modules required for the NTT hardware architecture, including: rotation factor storage module, delay feedback module, random delay module, data input module, control module and data distributor DEMUX.

[0056] The specific steps of designing the rotation factor storage module in step S1 are as follows:

[0057] (1) The rotation factor is a key parameter in the NTT algorithm. Since the number of rotation factors is large, its calculation involves the calculation of the power of large integers. In hardware design, if hardware resources are used to calculate and store the rotation factors, a large amount of hardware resources will be consumed and the operation speed of the overall NTT algorithm will be slowed down. Therefore, in this embodiment, a pre-calculation method is adopted to pre-calculate the rotation factors using Python language and then store them in FPGA to reduce unnecessary hardware resource loss; wherein, the modulus selection q = 12289 = 3 × 2 12 +1, the calculation formula of the rotation factor of the N-point NTT algorithm is: Due to the symmetric properties of the rotation factors, for example: Therefore, for NTT operations with smaller points, it can share the rotation factors with higher points. In order to realize the NTT operation with variable points, the rotation factors are calculated with the highest number of points, 1024. According to the algorithm 1 in the above table, according to the principle of the improved KEY-RED algorithm, the modulus obtained after pre-calculation needs to be multiplied by 3 -2 To ensure that the correct reduction result can be obtained after modular reduction, the Python script file is used to pre-calculate the rotation factors used, and the calculation formula of the rotation factors is corrected as follows:

[0058] Where 0≤k≤1023

[0059] Finally, the required 24-bit rotation factor is calculated. Wherein, k is the index of the rotation factor, indicating the position of the current rotation factor in the transformation, and is a variable used to traverse all rotation factors.

[0060] (2) Since the input module bit width is set to 24 bits in the KER-RED algorithm, a 10-port ROM module that can store 1024 24-bit twiddle factors is designed using the System Verilog language to ensure that the twiddle factors will not overflow and that the extraction of twiddle factors in each level of the SDF structure does not interfere with each other, ensuring that the calculated twiddle factors can meet the calculation requirements of a smaller number of points.

[0061] (3) Store the contents of the rotation factor file rotation_factor.txt generated by Python into ROM, and extract the rotation factor through the address input signal.

[0062] The specific method of designing the delayed feedback module in step S1 is as follows:

[0063] The delay feedback module is composed of an SDF structure and a modular multiplication module. The SDF structure is composed of an analog addition and analog subtraction module, a delay module, and a data selector. The input of the delay feedback module is connected to the analog addition module and the analog subtraction module, and the output of the analog subtraction module is connected to the delay module after passing through the data selector to achieve the feedback effect. For a single SDF structure, such as Figure 5 As shown, since the structure of each level of SDF is basically the same, and the three modular operation modules are the same for each level of SDF, a single SDF structure can be designed and then the depth of the delay unit can be modified to ensure that the data can be processed normally in a pipeline.

[0064] The specific operation mode of the modular addition and modular subtraction module is as follows: for the modular addition module, it is determined whether the sum of the two numbers is greater than or equal to the modulus q. If the sum of the two numbers is greater than or equal to the modulus q, q is subtracted from the sum of the two numbers; if the sum of the two numbers is less than q, no operation is performed; for the modular subtraction module, it is determined whether the difference between the two numbers is less than 0. If it is less than 0, q is added to the difference between the two numbers; if the difference between the two numbers is greater than 0, no operation is required. Therefore, the modular addition and modular subtraction module pair can be implemented by a simple data selector, and the result after modular addition or modular subtraction is selected by determining the relative size of the sum of the two numbers or the difference between the two numbers and q.

[0065] The design of the delay module is implemented by using a dual-port RAM plus an address generation module. The RAM is used for temporary storage of data, and the address generation module implements the delay operation through a counter. The delay depth can be changed by simply changing the counting range of the counter. The delay depth of each level is different. To implement the 1024-point NTT algorithm, 10 levels of SDF are required. For the first-level SDF structure, the delay depth is 512 cycles, the second level is 256 cycles, and so on. The delay depth of the last-level SDF unit is 1 cycle. For the 512-point NTT algorithm, under the action of the control unit, it can be implemented by simply calling the 2nd to 10th-level SDF structures, and there is no need to redefine the delay depth.

[0066] The random delay module is designed in step S1, and the specific method is as follows:

[0067] In order to improve the anti-SCA attack capability in the implementation of the NTT algorithm, a hardware-based pseudo-random number generator PRNG with data storage function is added after each selected level of SDF structure to generate different random delay times; after the random delay, the data is output to the next level of SDF structure. Figure 1 As shown, this embodiment selects to add random delay units after the 4th, 6th, and 8th level SDF structures.

[0068] The specific operation mode of the analog multiplication module in step S1 is:

[0069] The architecture of the modular multiplication module is as follows Figure 2 As shown, based on the KEY-RED algorithm, the product of two numbers with a bit width of less than 48 bits is divided twice into high bits and low bits; the modulus q is selected as 12289 = 3×2 12 +1, and use the special properties of the selected modulus to perform fast modular reduction, and through 2 shifts and 4 additions, keep the output width at 24 bits to complete the modular multiplication operation. For example, for a 48-bit multiplication result input, first divide the input into high and low bits, then perform 1 shift and 2 additions, and repeat this twice to get the correct modular multiplication result.

[0070] The design data input module described in step S1 is specifically designed as follows:

[0071] A dual-port RAM is designed that can store 1024 24-bit input data, and the data reading and writing of the two ports do not interfere with each other; when the number of input data points is less than 1024 points, the remaining bits are filled with 0; one port of the RAM is used for data input, and the other port is connected to a distributor DEMUX, and under the control of the control module, the input NTT point number is matched with the pipeline structure.

[0072] The specific operation mode of the control module in step S1 is:

[0073] Since the variable-point NTT algorithm needs to be implemented, the rotation factors used in each level of the modular multiplication module need to be reallocated according to the number of calculated points. The control unit extracts the data in the rotation factor storage module ROM to the corresponding modular multiplication unit according to the input control signal MODE to facilitate the butterfly operation.

[0074] S2. Based on the modules obtained in step S1, a variable point NTT hardware architecture is constructed. After each module is designed, it can be called to complete the construction of the overall structure. The overall structure is as follows: Figure 1 As shown, the specific steps are:

[0075] (1) Build a 10-level SDF structure in a pipeline structure;

[0076] (2) The output of the data input module is connected to the input of the data distributor DEMUX, and the output port of the data distributor DEMUX is connected to the input end of each level of SDF respectively; according to the control signal MODE received by the control module, the data is distributed to the input end of the corresponding level of SDF structure; for example, Figure 1 As shown, when the control signal MODE is 100, 010 and 001, the data is respectively allocated to the 1st level SDF structure, the 2nd level SDF structure and the 3rd level SDF structure, corresponding to the NTT algorithms of 1024, 512 and 256 points.

[0077] (3) The delay unit depths of the 1st-level SDF structure to the 10th-level SDF structure are set to 512D, 256D, ..., 1D, respectively, where D is the unit delay time.

[0078] (4) The 10 ports of the rotation factor storage module are connected to the input of the control module, and after being allocated by the control module based on the control signal MODE, they are independently connected to the modular multiplication module of each level of the SDF structure;

[0079] (5) The result of the modular operation is output in the 10th level SDF structure.

[0080] S3, setting the mode and input data, after synthesizing the design, implementing the design, and generating the bitstream file, the code is burned into the FPGA.

[0081] After the hardware architecture of the intrinsically secure variable-point NTT based on the KEY-RED algorithm is built, the input data required for the NTT operation is stored in a txt file and called in the define.sv file. At the same time, the MODE parameters are set to perform the NTT operation at the corresponding point. At the same time, the GMP large integer library can be used to write the corresponding Python code to implement the NTT algorithm. After the NTT algorithm is performed on the input data, the result is compared with the result after FPGA calculation to verify its correctness.

[0082] In particular, the variable-point NTT algorithm can be implemented by simply changing the content of the input data in the txt file and selecting the correct MODE.

[0083] The variable point NTT algorithm provided in this embodiment has different point numbers that are manually configured, and the circuit structures used to implement different points are different, so that attackers cannot determine the circuit implementation structure, which greatly improves security. In addition, in order to improve the ability to resist SCA attacks and prevent timing attacks, this embodiment randomly selects to introduce a random delay unit after the 4th, 6th, and 8th level SDF structures, making the overall NTT algorithm implementation time unpredictable. Random delays can also make it more difficult for attackers to synchronize physical attacks such as fault injection, which greatly improves the ability to resist timing attacks.

[0084] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. The hardware implementation method of the intrinsically secure variable point NTT based on the KEY-RED algorithm is characterized by: The following steps are involved: S1. Design the modules required for the NTT hardware architecture, including: rotation factor storage module, delay feedback module, random delay module, data input module, control module and data distributor DEMUX; S2. Building a variable-point NTT hardware architecture based on the modules obtained in step S1; S3, setting the mode and input data, after synthesizing the design, implementing the design, and generating the bitstream file, burn the code into the FPGA.

2. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 1, characterized in that: The specific design steps of the rotation factor storage module described in step S1 are as follows: (1) Use Python language to pre-calculate the rotation factor and generate a rotation factor file; (2) Use System Verilog to design a 10-port ROM module that can store 1024 24-bit twiddle factors; (3) The content in the rotation factor file is stored in the ROM module.

3. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 2 is characterized in that: The pre-calculation of the rotation factor is specifically as follows: Select modulus q = 12289 = 3 × 2 12 +1, the calculation formula of the rotation factor of the N-point NTT algorithm is: And the calculation formula of the modified rotation factor is: Where 0≤k≤1023 A 24-bit rotation factor is calculated; wherein k is the index of the rotation factor, indicating the position of the current rotation factor in the transformation, and is a variable used to traverse all rotation factors.

4. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 1, characterized in that: The delay feedback module described in step S1 is composed of an SDF structure and an analog multiplication module, and the SDF structure is composed of an analog addition and analog subtraction module, a delay module and a data selector; the input of the delay feedback module is connected to the analog addition module and the analog subtraction module, and the output of the analog subtraction module is connected to the delay module after passing through the data selector to achieve a feedback effect.

5. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 4 is characterized in that: The specific operation mode of the module addition and subtraction module is as follows: For the modular addition module, determine whether the sum of the two numbers is greater than or equal to the modulus q. If the sum of the two numbers is greater than or equal to the modulus q, subtract q from the sum of the two numbers; if the sum of the two numbers is less than q, do nothing. For the modular subtraction module, determine whether the difference between the two numbers is less than 0. If it is less than 0, add q to the difference between the two numbers. If the difference between the two numbers is greater than 0, no action is required.

6. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 4 is characterized in that: The delay module is implemented by a dual-port RAM and an address generation module, wherein the RAM is used for temporary storage of data; the address generation module implements delay operation through a counter, and changes the delay depth by changing the counting range of the counter.

7. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 4 is characterized in that: The random delay module described in step S1 is specifically designed as follows: A hardware-based pseudo-random number generator PRNG with data storage function is added after each of the selected levels of SDF structure to generate different random delay times; After a random delay, the data is output to the next level SDF structure.

8. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 1, characterized in that: The specific operation mode of the analog multiplication module in step S1 is: Based on the KEY-RED algorithm, the product of two numbers with a bit width of less than 48 bits is split twice into high bits and low bits; the modulus q is selected as 12289 = 3×2 12 +1, and use the special properties of the selected modulus to perform fast modular reduction, and complete the modular multiplication operation by keeping the output width at 24 bits through 2 shifts and 4 additions.

9. The method for implementing the intrinsically secure variable-point NTT hardware based on the KEY-RED algorithm according to claim 1, characterized in that: The data input module in step S1 is specifically designed as follows: Design a dual-port RAM that can store 1024 24-bit input data, and the data reading and writing of the two ports do not interfere with each other; when the number of input data points is less than 1024 points, the remaining bits are filled with 0; one port of the RAM is used for data input, and the other port is connected to a distributor DEMUX, and under the control of the control module, the input NTT point number is matched with the pipeline structure; The specific operation mode of the control module in step S1 is: According to the input control signal MODE, the data in the twiddle factor storage module is extracted to the corresponding modular multiplication unit.

10. The method for implementing the intrinsically secure variable point NTT hardware based on the KEY-RED algorithm according to claim 2, characterized in that: The step S2 is specifically as follows: (1) Build a 10-level SDF structure in a pipeline structure; (2) The output of the data input module is connected to the input of the data distributor DEMUX, and the output port of the data distributor DEMUX is connected to the input end of each level of SDF respectively; (3) Setting the delay unit depths of the 1st to 10th level SDF to 512D, 256D, ..., 1D, respectively, where D is the unit delay time; (4) The 10 ports of the rotation factor storage module are connected to the input of the control module, and after being allocated by the control module based on the control signal MODE, they are independently connected to the modular multiplication module of each level of the SDF structure; (5) The result of the modular operation is output in the 10th level SDF structure.