NTT hardware implementation method

By coupling the NTT hardware accelerator to the LSU pipeline of the processor, using the address generation module and the butterfly operation module for direct reading and writing of data, the problem of large data transmission overhead in the NTT hardware implementation is solved and the processing performance is improved.

CN120144090APending Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204528.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the existing NTT hardware implementation, the data transmission overhead is large and the transmission time is long, which affects the processing performance.

Method used

The NTT hardware accelerator is coupled to the LSU pipeline of the processor, and the necessary addresses are generated through the address generation module, and the butterfly operation module is used to read and write data from the main memory at one time through the LSU, reducing the transmission and communication overhead of data in the processor pipeline.

Benefits of technology

It reduces the overhead and time of data transmission and improves the overall processing performance of the NTT processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144090A_ABST
    Figure CN120144090A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of information communication, and provides an NTT hardware implementation method which is applied to an NTT hardware accelerator, the NTT hardware accelerator is coupled to an LSU of a processor, and the NTT hardware accelerator comprises an address generation module and a butterfly operation module. The NTT hardware implementation method comprises the steps that an address generation module generates addresses of two input coefficients, addresses of two output results and an address of a twiddle factor; and the butterfly operation module reads the two input coefficients and the twiddle factor from the main memory at one time through the LSU based on the addresses of the two input coefficients and the address of the twiddle factor, performs butterfly operation on the two input coefficients and the twiddle factor, and writes a butterfly operation result into the main memory through the LSU based on the addresses of the two output results. According to the method and the device, the data transmission overhead during NTT hardware implementation can be reduced, and the data transmission time is shortened at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information and communication technologies, and particularly relates to a method for hardware implementation of NTT. Background Art

[0002] With the rapid development of quantum algorithms and quantum computing technologies, traditional public-key encryption algorithms are facing unprecedented challenges. To address these challenges, post-quantum cryptographic algorithms have emerged, along with hardware integration processor implementation solutions for the core operations of these algorithms - the Number Theoretic Transform (NTT).

[0003] Currently, existing NTT coupling schemes are tightly coupled to the decoding stage and the execution stage. Tightly coupling to the decoding stage easily causes the critical path to appear in the decoding stage, thereby affecting the main frequency of the processor. At the same time, the coupled NTT structure still requires register storage and pipelining after data processing is completed, that is, the additional transmission overhead is still unavoidable. Similarly, coupling to the execution stage also requires data pipelining and storage, and its transmission overhead is also unavoidable. Moreover, both of the above have a common drawback, that is, the input data to be processed needs to be stored in a register before data processing can be performed, and after the data is processed, the result is pipelined and transmitted to the memory module to be written back to the main memory, which requires an additional two cycles and will affect the overall processing performance of the NTT processor in some cases. Summary of the Invention

[0004] The embodiments of this application provide a method for hardware implementation of NTT, which can solve the problems of large data transmission overhead and long transmission time during the hardware implementation of NTT.

[0005] The embodiments of this application provide a method for hardware implementation of NTT, which is applied to an NTT hardware accelerator. The NTT hardware accelerator is coupled to the LSU of the processor. The NTT hardware accelerator includes an address generation module and a butterfly operation module. The method for hardware implementation of NTT includes:

[0006] The address generation module generates the addresses of two input coefficients, the addresses of two output results, and the address of a rotation factor;

[0007] The butterfly operation module, based on the addresses of two input coefficients and the address of a rotation factor, reads two input coefficients and a rotation factor from the main memory at one time through the LSU, performs a butterfly operation on the two input coefficients and the rotation factor, and writes the butterfly operation result to the main memory through the LSU based on the addresses of the two output results.

[0008] Optionally, before the step of the address generation module generating the addresses of two input coefficients, the addresses of two output results, and the address of a rotation factor, the method for hardware implementation of NTT further includes:

[0009] Configure the order \(n\) and the modulus \(q\) of the NTT hardware accelerator through custom instructions.

[0010] Optionally, before the step of the address generation module generating the addresses of two input coefficients, the addresses of two output results, and the address of a twiddle factor, the NTT hardware implementation method further includes:

[0011] Transfer the parameters of the source register of the processor to the NTT hardware accelerator for configuration; wherein, if the parameter of the source register is 1, the NTT operation is enabled.

[0012] Optionally, the NTT hardware implementation method further includes:

[0013] If the parameter of the source register is 0, perform the INTT operation.

[0014] Optionally, the lengths of the addresses of two input coefficients and the length of the address of a twiddle factor are both 32 bits.

[0015] Optionally, after the step of the address generation module generating the addresses of two input coefficients, the addresses of two output results, and the address of a twiddle factor, the NTT hardware implementation method further includes:

[0016] The address generation module concatenates the addresses of two 32-bit input coefficients and the address of a 32-bit twiddle factor into a 96-bit address and sends it to the read address port of the main memory.

[0017] Optionally, the NTT hardware implementation method further includes:

[0018] When performing the NTT operation, the ports of the main memory and the processor are connected to the interaction interface of the NTT hardware accelerator, and a 96-bit data path is enabled.

[0019] The above solution of the present application has the following beneficial effects:

[0020] In the embodiment of the present application, by coupling the NTT hardware accelerator to the LSU pipeline of the processor, the NTT hardware accelerator is coupled to the memory access stage, and its data interface is directly connected to the interaction interface between the processor and the main memory, thereby reducing the transmission and communication overhead of data in the processor pipeline and shortening the data transmission time.

[0021] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a schematic structural diagram of an NTT hardware accelerator provided by an embodiment of the present application;

[0024] Figure 2 It is a flowchart of an NTT hardware implementation method provided by an embodiment of the present application. Detailed implementation manners

[0025] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0026] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0027] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0028] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0029] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0031] In view of the problems of large data transmission overhead and long transmission time in the current NTT hardware implementation, the embodiments of this application provide a method for implementing NTT in hardware. This method couples the NTT hardware accelerator to the LSU pipeline of the processor, realizes coupling the NTT hardware accelerator to the memory access stage, and its data interface is directly connected to the interaction interface between the processor and the main memory, thereby reducing the transmission communication overhead of data in the processor pipeline and shortening the data transmission time at the same time.

[0032] The method for implementing NTT in hardware provided by this application will be exemplarily described below in conjunction with specific embodiments.

[0033] The method for implementing NTT in hardware provided by this application is applied to an NTT hardware accelerator, as Figure 1 shown. This NTT hardware accelerator is coupled to the load store unit (LSU) of the processor. The NTT hardware accelerator includes an address generation module and a butterfly operation module. As Figure 2 shown, the method for implementing NTT in hardware provided by this application includes the following steps:

[0034] Step 21, the address generation module generates the addresses of two input coefficients, the addresses of two output results, and the address of a rotation factor.

[0035] Step 22, based on the addresses of two input coefficients and the address of a rotation factor, the butterfly operation module reads two input coefficients and a rotation factor from the main memory at one time through the LSU, performs a butterfly operation on the two input coefficients and the rotation factor, and writes the butterfly operation result into the main memory through the LSU based on the addresses of the two output results.

[0036] Among them, the address generation module is mainly used to generate the addresses for accessing the main memory (RAM). It has five output ports, namely the addresses of two input coefficients, the addresses of two output results, and the address of a rotation factor. As an optional example, this address generation module can be implemented by a controller, and the processor can be the processor RI5CY suitable for low-power application scenarios.

[0037] The butterfly operation module includes three input ports and two output ports. The input ports receive two input coefficients and a rotation factor read from the RAM. The two output data are the two results of the butterfly operation. This butterfly operation module is mainly used to perform the butterfly operation on the two input coefficients and the rotation factor, and write the two results of the butterfly operation into the main memory through the LSU according to the addresses of the two output results.

[0038] In some embodiments of the present application, before starting the NTT operation, the order n and the modulus q of the NTT hardware accelerator can be configured through instructions. Specifically, the order n and the modulus q of the NTT hardware accelerator can be configured through custom instructions, thus greatly enhancing the flexibility of the design. This flexibility enables the design solution of the present application to support a variety of different post-quantum cryptography algorithms (PQC) and polynomial operations of different security levels, and adapt to diverse encryption requirements.

[0039] In practical applications, the NTT hardware accelerator can be configured accordingly according to different algorithms. Instructions with different configuration parameters contain different parameter information. For example, when configuring the modulus q, the parameter q of the source register rs1 of the processor is transmitted to the NTT hardware accelerator for configuration; when configuring the order n, the order n of the source register rs1 of the processor is transmitted to the NTT hardware accelerator for configuration.

[0040] In some embodiments of the present application, the NTT operation can be started by transmitting the parameters of the source register rs1 of the processor to the NTT hardware accelerator for configuration. Specifically, in the same way as configuring the modulus q, the parameters of the source register are transmitted to the NTT hardware accelerator for configuration. If the parameter of the source register is 1, the NTT operation is started; if the parameter of the source register is 0, the INTT operation is performed. That is, if NTT is performed, the value of rs1 is 1, otherwise it is 0. It should be noted that after starting the NTT operation, the instruction operations in the front of the pipeline will be paused until the NTT operation ends.

[0041] It is worth mentioning that the input data interface of the NTT hardware accelerator is directly connected to the main memory interface. Through the designed control logic module (i.e., the address generation module), the input of data, the output results of the butterfly operation, and the addresses for memory access are controlled to ensure the correctness of the data stream. When performing NTT / INTT (INTT is the inverse operation of NTT) operations, the lengths of the addresses of the above two input coefficients and the address of a rotation factor can all be 32 bits. After the address generation module generates the addresses of the two input coefficients and the address of a rotation factor, the address generation module will concatenate the addresses of the two 32-bit input coefficients and the address of a 32-bit rotation factor into a 96-bit address and send it to the read address port of the main memory (RAM). In the next cycle, the RAM directly sends the three read data to the butterfly operation module for processing through the storage interface. Since the butterfly operation module of this application is pure combinational logic, the butterfly operation is immediately performed after the data is input to obtain two output results, and then the results are written back to the RAM according to the write address in the next cycle.

[0042] It should be noted that the read and write data ports of this application are dynamically configurable. When performing NTT operations, the ports of the main memory and the processor are connected to the interaction interface of the NTT hardware accelerator, and a 96-bit data path is enabled. In practical applications, when performing NTT / INTT operations, the data path can adopt sizes of 96 bits and 64 bits, and can access up to three 32-bit data in the RAM at one time, significantly improving the throughput of data processing. This efficient data access mechanism makes the NTT hardware accelerator more rapid and efficient when processing a large amount of data; while in most other instruction operation cases, the original 32-bit data bit width is maintained. This dynamic configuration ability enables the NTT hardware accelerator to flexibly adapt to different data processing requirements while maintaining high efficiency and low power consumption.

[0043] In practical applications, after the NTT operation is enabled, the NTT hardware accelerator starts to access the main memory. The access address has a width of 96 bits (bit) and can address three pieces of data, including two input coefficients and a rotation factor; in the next cycle, the three pieces of data are transmitted to the NTT hardware accelerator and immediately execute the calculation to output the results; in the next cycle, the two output results and the addresses of the two results are interacted with the main memory, and the results are written back to the corresponding positions. The subsequent data to be processed (i.e., two input coefficients and a rotation factor) can repeat the above process.

[0044] It can be seen that the NTT hardware accelerator of the present application realizes the triple advantages of high performance, high throughput, and high flexibility through tight integration into the LSU, dynamic data path configuration, and flexible configuration of parameters, providing an efficient and scalable platform for the hardware implementation of post-quantum cryptographic algorithms. At the same time, compared with other design schemes, it significantly reduces the overhead of data transfer to registers and avoids the reduction of the main frequency caused by integration into the decoding stage. This integration method not only optimizes the data flow but also maintains the performance of the processor.

[0045] Specifically, the present application cleverly integrates the extension of the RISC-V instruction set, the adjustment of the processor architecture, and the implementation of the NTT hardware accelerator to form a unified whole.

[0046] First, according to the data access timing of the processor memory access module, an NTT hardware accelerator with an NTT / INTT enable signal is designed and implemented. The NTT hardware accelerator, according to the aforementioned structure, includes three input ports and seven output ports for efficient data access and interaction with the RAM.

[0047] Next, a new instruction set is extended to activate the NTT hardware accelerator and configure its parameters. When designing the instruction format, the occupied instruction formats in the original instruction set are cleverly avoided. For example, in this case, the R-type instruction is adopted, and 0000011 is selected as the value of the funct7 field (the funct7 field is part of the RISC-V instruction set) because this value has not been used in the previous instruction set. At the same time, the funct3 field (the funct3 field is part of the RISC-V instruction set) is used to distinguish the NTT enable instruction and the parameter configuration instruction. After designing the instruction format, it is necessary to recompile the RISC-V toolchain or use the method of inline assembly combined with the.word format to ensure that the toolchain can recognize and compile the extended instructions. As long as the toolchain can recognize and compile these extended instructions, the expected functions can be achieved.

[0048] Finally, and most importantly, the NTT hardware accelerator is integrated into the memory access stage of the processor pipeline. By using a multiplexer, the interaction signals (such as memory access address, write-back data, and data input, etc.) in the NTT execution stage are selectively connected to the signals of the original memory access instructions. When performing the NTT operation, the corresponding selection signal will connect the ports of the RAM and the processor to the NTT interaction interface and enable the 96-bit data path; for other instructions, it will be connected back to the original LSU interface and the unused 64-bit data path will be masked to save power consumption.

[0049] Through this series of carefully designed steps, the NTT hardware implementation method of this application not only improves the efficiency of the NTT operation, but also maintains the flexibility and energy efficiency of the processor, providing an efficient and scalable solution for the hardware implementation of post-quantum cryptographic algorithms.

[0050] The above are the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle described in this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A NTT hardware implementation method, characterized in that: Applied to NTT hardware accelerator, the NTT hardware accelerator is coupled to the LSU of the processor, the NTT hardware accelerator includes an address generation module and a butterfly operation module, and the NTT hardware implementation method includes: The address generation module generates addresses of two input coefficients, addresses of two output results, and an address of a rotation factor; The butterfly operation module reads the two input coefficients and the rotation factor from the main memory at one time through the LSU based on the addresses of the two input coefficients and the address of the one rotation factor, performs a butterfly operation on the two input coefficients and the rotation factor, and writes the butterfly operation results into the main memory through the LSU based on the addresses of the two output results.

2. The NTT hardware implementation method according to claim 1, characterized in that: Before the step of generating, by the address generation module, addresses of two input coefficients, addresses of two output results, and an address of a rotation factor, the NTT hardware implementation method further includes: The order n and modulus q of the NTT hardware accelerator are configured through custom instructions.

3. The NTT hardware implementation method according to claim 1, characterized in that: Before the step of generating, by the address generation module, addresses of two input coefficients, addresses of two output results, and an address of a rotation factor, the NTT hardware implementation method further includes: The parameters of the source register of the processor are transmitted to the NTT hardware accelerator for configuration; wherein, if the parameter of the source register is 1, the NTT operation is started.

4. The NTT hardware implementation method according to claim 3, characterized in that: The NTT hardware implementation method further includes: If the parameter of the source register is 0, an INTT operation is performed.

5. The NTT hardware implementation method according to claim 1, characterized in that: The lengths of the addresses of the two input coefficients and the length of the address of the one rotation factor are both 32 bits.

6. The NTT hardware implementation method according to claim 5, characterized in that: After the address generation module generates the addresses of two input coefficients, the addresses of two output results, and the address of a rotation factor, the NTT hardware implementation method further includes: The address generation module concatenates two 32-bit input coefficient addresses and one 32-bit rotation factor address into a 96-bit address and sends it to the read address port of the main memory.

7. The NTT hardware implementation method according to claim 1, characterized in that: The NTT hardware implementation method further includes: When performing NTT operations, the ports of the main memory and the processor are connected to the interactive interface of the NTT hardware accelerator, and a 96-bit data path is enabled.