SRAM computing integrated NTT accelerator, chip and electronic equipment

By combining SRAM storage and computing technology with the NTT hardware accelerator, the problem of time-consuming lattice cryptographic polynomial multiplication operations is solved, and highly parallel processing and low-energy consumption NTT operations are achieved, which is suitable for large-scale data processing scenarios.

CN119988803BActive Publication Date: 2025-10-10SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510059996.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-10-10
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

In the existing technology, the polynomial multiplication operation of lattice cipher takes a long time, especially in the case of high bit width and high number of operation points, resulting in high power consumption and excessive consumption of computing resources, and cannot effectively process large-scale data or high-precision calculations. In particular, when processing high-resolution images, the frequent interaction between the storage unit and the computing unit leads to excessive power consumption.

Method used

The SRAM storage and computing technology is combined with the NTT hardware accelerator, including the storage and computing butterfly unit, array decoder, controller, modular reduction simple element and inter-column data router to realize butterfly operation and store intermediate data. Fast modular operation is performed through modular reduction simple element. Combined with constant time reduction method and high parallel data arrangement, the modular reduction operation of the butterfly structure is optimized.

Benefits of technology

It achieves high parallel processing of NTT operations, reduces the delay of polynomial multiplication, adapts to large-scale data processing, reduces data transmission time and overhead, improves computing efficiency and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988803B_ABST
    Figure CN119988803B_ABST
Patent Text Reader

Abstract

The application discloses an SRAM computing integrated NTT accelerator, a chip and electronic equipment, wherein the NTT accelerator comprises a computing integrated butterfly unit, an array decoder, a controller, a module approximation simple unit and an inter-column data router; the computing integrated butterfly unit is used for accelerating the NTT algorithm, realizing two kinds of butterfly operation operations, and storing input data of one-time butterfly operation and intermediate data generated in the calculation process; the array decoder and the controller are used for controlling the working state of each computing integrated butterfly unit when the butterfly operation is performed; the module approximation simple unit is used for processing the modulo operation of the data output after the multiplication of the multiplier is completed; the inter-column data router is connected with the input and output of each butterfly unit, is used for writing initial data into the butterfly unit, and is used for transmitting the calculation result of each stage to the corresponding unit, and preparing for the operation of the next stage. The application supports bit parallel operation, guarantees that the butterfly operation of one stage of NTT can be completed at the same time, and maximizes the operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security, and in particular to an SRAM storage and computing integrated NTT accelerator, chip, and electronic equipment. Background Art

[0002] Public-key cryptography provides a reliable way to ensure the confidentiality, integrity, and security of data, playing a critical role in global digital communications systems. However, the security of current public-key cryptography systems is threatened by the increasing computing power of adversaries and the continuous advancement of cryptanalysis techniques. Increasing key sizes also increases the computational cost required for honest parties to communicate securely: larger key sizes result in increased consumption of computing and storage resources, which in turn impacts the efficiency and real-time nature of communications. In particular, the immense computing power of quantum computers poses a threat to the security and stability of modern encryption technologies. To address this challenge and protect information security, research into post-quantum cryptographic algorithms is imperative.

[0003] Among the currently mainstream post-quantum cryptography, lattice cryptography offers significant advantages over other post-quantum cryptography methods in terms of security, computational efficiency, and versatility in designing public-key encryption, digital signatures, and cryptographic agreement protocols. Polynomial multiplication in lattice cryptography is mostly performed on polynomial rings. In particular, the multiplication of two high-order polynomials consumes significant computational cost, making improving computational speed a pressing issue. Number theoretic transforms (NTTs) can reduce the time complexity of polynomial computations and are applicable to polynomial multiplication in lattice cryptography based on error learning problems over rings. They also avoid the precision loss associated with complex number multiplication in fast Fourier transforms, leading to their widespread use in many lattice cryptography systems. However, NTT operations require the cryptographic algorithm to use a specific modulus, and different lattice cryptography algorithms have specific parameters. This modulus affects the computational speed and security of the overall algorithm. As the bit width of encryption operands or the number of encryption points increases, most designs exhibit high power consumption and are unable to properly address the performance bottlenecks that may be encountered in application scenarios involving processing large-scale data or high-precision calculations. A large amount of computing resources are required in the encryption and decryption process, especially when processing high-resolution images. Frequent data interaction between storage units and computing units will lead to excessive power consumption.

[0004] Storage-computing integration is an emerging computing model that aims to integrate computing resources directly within the storage medium, thus overcoming the limitations of excessive bandwidth usage. This reduces the energy consumption and latency of memory access, achieving a close integration of data and computing. Originally proposed to address the memory wall problem in large-scale data processing and highly parallel computing scenarios, it has been widely used in fields such as neural networks, machine learning, and artificial intelligence. Similarly, lattice-based post-quantum cryptography schemes also encounter the memory wall problem as the number of operation points and bit width increase. Incorporating storage-computing integration technology into the hardware architecture of lattice-based post-quantum cryptography schemes can effectively reduce data transmission time and overhead, thereby improving computing performance and efficiency. Recent research has also combined storage-computing integration with lattice-based cryptography. However, individual butterfly structures can only handle operations such as multiplication and modular reduction through bit-serial operations, which is time-consuming. Summary of the Invention

[0005] In order to at least partially solve one of the technical problems existing in the prior art, the purpose of the present invention is to provide an NTT hardware accelerator, chip and electronic device combined with SRAM storage and computing technology, which is suitable for accelerating the NTT part of the CRYSTALS-KYBER algorithm.

[0006] The first technical solution adopted by the present invention is:

[0007] An SRAM storage and computing integrated NTT accelerator, including: 128 storage and computing integrated butterfly units, 1 array decoder, 1 controller, 1 modular simple unit, and 1 inter-column data router;

[0008] The storage-computation integrated butterfly unit is used to accelerate the NTT algorithm, implement two butterfly operations, and store the input data of a butterfly operation and the intermediate data generated during the calculation process;

[0009] The array decoder and controller are used to control the working state of each storage-computation-in-one butterfly unit when performing butterfly operations;

[0010] The modular reduction simple element is used to process the modular operation of the 24-bit width data output after the multiplier completes the 12×12-bit multiplication;

[0011] The inter-column data router is connected to the input and output of each butterfly unit, and is used to write initial data into the butterfly unit and transmit the calculation results of each stage to the corresponding unit to prepare for the next stage of calculation.

[0012] Furthermore, the memory-computation integrated butterfly unit includes an SRAM memory array circuit and a near-memory logic operation circuit;

[0013] The SRAM memory array includes 12×12 6T-SRAM memory cells, a total of 144 cells, 12 pre-charge circuits, 12 write circuits, and 12 read circuits;

[0014] 144 6T-SRAM memory cells are used to store 12 12-bit data in rows;

[0015] One precharge circuit is shared by a column of 1×12 6T-SRAM and is used to charge the bit line to a high level;

[0016] One write circuit is shared by a column of 1×12 6T-SRAM and is used to write data into the SRAM memory cell;

[0017] One read circuit is shared by a column of 1×12 6T-SRAM and is used to read data from the SRAM memory cells;

[0018] The near-memory logic operation circuit includes 12 sense amplifiers, a data driver, a two-to-one multiplexer, a unified adder, and a range selector;

[0019] The sense amplifier is used to collect the 12-bit data read out of the SRAM storage array every cycle;

[0020] The data driver is used to drive the intermediate results generated during the calculation process to the write control port of the SRAM storage array;

[0021] The range selector is placed after the unified adder and is used to select the result after the unified adder operation as the final result output;

[0022] The two-to-one multiplexer is used to select the working mode of the circuit. The circuit is controlled to work in CT mode or GS mode through a control signal. When the control signal is 0, the circuit works in CT mode, and the entire architecture is used to perform NTT operations. When the control signal is 1, the circuit works in GS mode, and the architecture is used to perform INTT operations.

[0023] Furthermore, the unified adder integrates the hardware resources of the adder that need to be reused in the butterfly operation process, specifically implementing 12-bit multiplication, fast modulo operation and modular addition in the butterfly operation to perform addition operations, which are selected by the control signal sel;

[0024] 1) When sel = 0, the unified adder acts as an accumulator for 12-bit × 12-bit multiplication. In this mode, the unified adder accumulates the shifted 12-bit data in each clock cycle to perform a 12-bit by 12-bit multiplication.

[0025] 2) When sel = 1, the unified adder also works as an accumulator, but its purpose is to perform Mod3329 reduction on the input data. Specifically, the unified adder accumulates the intermediate results of the Mod3329 reduction of the simple elements in each cycle, and then passes through the range selector to obtain the final modular reduction result.

[0026] 3) When sel=2, the unified adder and the range selector module together form a modular adder; in this mode, the unified adder and the range selector work together to perform modular addition operations.

[0027] Furthermore, based on the constant time reduction method, the modular simple element is based on the constant time reduction method, using the recursive algorithm 2 12 ≡2 9 -2 8 +1(mod3329) converts the number larger than 12 bits in the NTT operation into several numbers smaller than 12 bits on the integer ring.

[0028] Furthermore, a calculation method of XORing the weight and the corresponding bit is introduced to extract the weight of the recursive result. The weight extraction method is: in units of rows, the weight of the inverted bit is recorded as 1, and the weight of the inverted bit is recorded as 0.

[0029] Furthermore, the specific steps of NTT calculation are as follows:

[0030] 1) Preparation phase: 256 polynomial coefficients and 128 twiddle factors are pre-stored in the memory structure of the corresponding butterfly unit. A single butterfly unit needs to pre-store 8 twiddle factors, corresponding to 8 stages of calculation.

[0031] 2) First-stage operation: Each butterfly unit calls the data at a specified location in the unit's storage array to start the butterfly operation. When the first-stage butterfly operation is completed, the calculation result is written back to the previous address.

[0032] 3) Data migration: The data router writes the initial data into the butterfly unit and transmits the calculation results of this stage to the corresponding unit, adjusting the data position between units to prepare for subsequent operations;

[0033] 4) Similarly, after completing 8 stages of butterfly operations, the final result can be output.

[0034] Furthermore, the inter-column data router adopts an efficient routing method: as much data as possible is left in the original location, and only one data in each butterfly structure needs to be transferred.

[0035] Furthermore, the routing mapping mode of the inter-column data router is as follows: the address of the router internal register transposition operation is the current data routing stage router i And the function of original address index k:

[0036] 1) The first data routing stage:

[0037] addr = {~k[msb-router i +1],k[msb-router i :0]}

[0038] 2) Intermediate data routing stage:

[0039] addr = {k[msb:msb-router i +2],~k[msb-router i +1]}

[0040] 3) The last data routing stage:

[0041] addr = {k[msb:msb-router i +2],~k[msb-router i +1],k[msb-router i ∶0]}

[0042] Among them, MSB represents the most significant bit.

[0043] The second technical solution adopted by the present invention is:

[0044] A chip comprising the SRAM storage-computing integrated NTT accelerator as described above.

[0045] The third technical solution adopted by the present invention is:

[0046] An electronic device comprises the chip described above.

[0047] The NTT hardware architecture proposed by the present invention supports bit-parallel operations, ensuring that butterfly operations in one stage of NTT can be completed simultaneously, maximizing operational efficiency. It is suitable for large-scale data processing applications and reduces the latency of polynomial multiplication during encryption and decryption calculations. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.

[0049] Figure 1 This is a schematic diagram of the NTT hardware architecture combined with SRAM storage and computing technology in an embodiment of the present invention.

[0050] Figure 2 It is a schematic diagram of the storage and computing integrated butterfly unit structure in an embodiment of the present invention.

[0051] Figure 3 Schematic diagram of the circuit structure of a 6T-SRAM memory cell in an embodiment of the present invention.

[0052] Figure 4 2 is a schematic diagram of the hardware for modular reduction operation in an embodiment of the present invention.

[0053] Figure 5 This is a schematic diagram of the efficient routing method stages in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention. The step numbers in the following embodiments are provided for ease of explanation only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0055] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention.

[0056] In the description of the present invention, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0057] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0058] In response to existing technical problems, the present invention proposes an NTT hardware accelerator combined with SRAM storage and computing technology, which is suitable for accelerating the NTT part of the CRYSTALS-KYBER algorithm. This solution is proposed for application scenarios with high computing speed requirements and relatively loose resource usage. The modular addition, subtraction and multiplication required by the butterfly structure in the NTT operation are completed through the SRAM array and near-memory computing circuits, realizing a composite butterfly unit that supports two butterfly operations, and combining the constant-time modular reduction method with the storage and computing technology to pre-load the weights that need to take the modular modulus to quickly realize the modular reduction operation link in the butterfly structure. The proposed high-parallel data arrangement method is used to accelerate the implementation of the NTT and INTT operation links in the encryption algorithm, effectively reducing the area and energy consumption. By executing multiple butterfly structures in parallel and an efficient inter-unit data routing method, data can be efficiently circulated between each unit, and the NTT calculation can be completed quickly and efficiently.

[0059] like Figure 1 As shown, this embodiment provides an NTT hardware architecture combined with SRAM storage and computing technology, which consists of 128 storage and computing integrated butterfly units, ensuring that the butterfly operations of one stage of NTT can be completed simultaneously, maximizing computing efficiency. The array decoder and controller are used to control the working state of each storage and computing integrated butterfly unit when performing butterfly operations. The modular simple element and the butterfly unit jointly implement 12-bit modular operations. Since the SRAM storage array can only read and write one row of data at the same time, the input and output of the butterfly unit in this article are both 12-bit numbers. The input and output of each butterfly unit are connected to the inter-column data router, which uniformly controls the reading and writing. The inter-unit data router is not only used to write the initial data into the butterfly unit, but also responsible for transmitting the calculation results of each stage to the corresponding unit, preparing for the next stage of calculation.

[0060] In preparation for the NTT operation, 256 polynomial coefficients and 128 twiddle factors are pre-stored in the memory structure of the corresponding butterfly unit. Because each butterfly unit is only used once in each stage of the NTT operation, a single butterfly unit needs to store 8 twiddle factors, corresponding to 8 stages of calculation.

[0061] At the start of an NTT operation, each butterfly unit begins performing a butterfly operation by accessing data from a specified location in the unit's internal memory array. After a certain computational time, the first stage of butterfly operation is complete, and the calculation results E and O are written back to the addresses previously occupied by A and B. This is when the data movement phase begins, and the data router begins to work, adjusting the data locations between units in preparation for subsequent operations. The second stage of butterfly operation then begins, and so on. After completing eight stages of butterfly operations, the final result is output.

[0062] like Figure 2 As shown, Figure 2 This is a schematic diagram of the integrated storage and computation butterfly unit structure described in an embodiment, comprising an SRAM storage array circuit and a near-memory logic operation circuit. The SRAM array circuit is used to store the input data of a butterfly operation and the intermediate data generated during the calculation process. The near-memory logic operation circuit is used to implement two butterfly operations.

[0063] As an optional embodiment, the SRAM memory array includes 144 6T-SRAM memory cells in a 12×12 array, 12 precharge circuits, 12 write circuits, and 12 read circuits. The 144 6T-SRAM memory cells are used to store 12 12-bit data in rows; one precharge circuit is shared by a column of 1×12 6T-SRAM cells and is used to charge the bit line to a high level; one write circuit is shared by a column of 1×12 6T-SRAM cells and is used to write data to the SRAM memory cells; and one read circuit is shared by a column of 1×12 6T-SRAM cells and is used to read data from the SRAM memory cells.

[0064] Specifically, if Figure 3 As shown, Figure 3The circuit diagram of the 6T-SRAM memory cell described in this embodiment is shown, where M1, M3, M5, and M6 are NMOS transistors, conducting at a high level; M2 and M4 are PMOS transistors, conducting at a low level. M1, M2, and M3, M4 each form two inverters, which are connected in the first position here. This end-to-end inverter structure is a key structure for storing data in memory. The BL (Bit Line) and BLB connected to the source / gate of M5 and M6 are bit lines, used for reading and writing data. The WL (Word Line) connected to the drain of M5 and M6 is the word line, used to control read and write operations. Each bit of data in the SRAM is stored in the two cross-connected inverters composed of M1, M2, M3, and M4 (i.e., the Q terminal and / Q terminal in the figure). The two NMOS transistors M5 and M6 are control switches, used to control the transfer of data from the memory cell to the bit lines.

[0065] As an optional embodiment, the near-memory logic circuit includes 12 sense amplifiers, a data driver, a two-to-one multiplexer, a unified adder, and a range selector. The sense amplifiers are used to collect the 12 bits of data read out of the SRAM storage array each cycle; the data driver is used to drive the intermediate results generated during the calculation process to the write control port of the SRAM storage array.

[0066] The unified adder integrates the hardware resources of the adder that need to be reused during the butterfly operation. Specifically, it implements 12-bit multiplication, fast modulo operation, and modulo addition in the butterfly operation to perform addition operations. The selection is made by the control signal sel:

[0067] 1) When sel = 0, the unified adder acts as an accumulator for 12-bit × 12-bit calculations. In this mode, the unified adder accumulates the shifted 12-bit data in each clock cycle to implement a 12-bit by 12-bit multiplication operation.

[0068] 2) When sel=1, the unified adder also works as an accumulator, but its purpose is to perform Mod3329 reduction on the input data;

[0069] 3) When sel=2, the unified adder and the range selector module together form a modular adder; in this mode, the unified adder and the range selector work together to perform modular addition operations.

[0070] The range selector, placed after the unified adder, selects the unified adder's result as the final output. A two-to-one multiplexer selects the circuit's operating mode. A control signal controls the circuit's operation between CT mode and GS mode. When the control signal is 0, the circuit operates in CT mode, and the entire architecture is used to perform NTT operations. When the control signal is 1, the circuit operates in GS mode, and the architecture is used to perform INTT operations.

[0071] like Figure 4 As shown, Figure 4 This is the overall diagram of the modular simple unit and the butterfly unit working together. Its storage structure is based on rows, starting from the first row, and stores 9 12-bit weights mw0 to mw8 in sequence. It stores the weight mw in each cycle. x Read out row by row, perform XOR operation with the corresponding operand bit, and then calculate the result s obtained in each cycle x The sum is accumulated to get S. Because the weight array is set to 9 rows, it takes 9 clock cycles to perform a modular reduction operation on a 24-bit operand in this architecture.

[0072] like Figure 5 As shown, Figure 5 This is a schematic diagram of the stages of this efficient routing method. Figure 5 The two merged blocks in the figure form a butterfly unit storage structure. The gray addresses indicate that the data in them needs to be moved through the data router. Figure 5 As can be seen in the figure, only one data item needs to be moved in each butterfly structure. Using this method, the 8-point NTT operation only needs to move four data items each time it switches phases, leaving the remaining four data items unchanged. This way, we keep as much data as possible in its original location.

[0073] With this allocation method, in each routing stage, the data router only needs to receive one data from each butterfly unit. After receiving the output data from the 128 butterfly units, the data router temporarily stores it in order in the 12×128 registers and then performs the transposition operation. After the transposition is completed, the new data in the register is output to the butterfly unit in order to ensure the calculation of the next stage. Since 256-point NTT requires 8 stages of butterfly operations, there are a total of 7 data routing stages between butterfly operations. We use the router i =1~7 to represent these stages. Here, the address of the router internal register transposition operation is the current data routing stage router i And the function of the original address index k, the mapping method is as follows: the transposition operation and address mapping relationship of the device:

[0074] If it is the first data routing stage, then:

[0075] addr = {~k[msb-router i +1],k[msb-router i :0]}

[0076] If it is the last data routing stage, then:

[0077] addr = {k[msb:msb-router i +2],~k[msb-router i +1]}

[0078] If it is any other stage, then:

[0079] addr = {k[msb:msb-router i +2],~k[msb-router i +1],k[msb-router i ∶0]}

[0080] INTT follows a similar mapping strategy and shares the same data router. Using this mapping strategy to implement register transposition requires two clock cycles for the data router to complete. Including the read and write operations of the SRAM storage structure, a single data routing stage requires a total of four clock cycles.

[0081] In summary, the NTT hardware accelerator of the present invention has at least the following advantages and beneficial effects compared to the prior art:

[0082] (1) The NTT hardware architecture proposed in this paper supports bit-parallel operations, ensuring that butterfly operations in one stage of NTT can be completed simultaneously, maximizing operational efficiency. It is suitable for large-scale data processing applications and reduces the latency of polynomial multiplication during encryption and decryption calculations.

[0083] (2) The modular reduction simple element proposed in this invention is based on a constant-time reduction method. It uses a recursive algorithm to convert numbers larger than 12 bits in the NTT operation into equivalent numbers smaller than 12 bits on the integer ring. A calculation method that uses an exclusive-OR operation between weights and corresponding digits is introduced, and weights are extracted from the recursive results to enable fast operations such as negation and summation of corresponding digits in the storage and calculation array, thus quickly implementing the modular reduction function.

[0084] (3) The present invention proposes an efficient routing method that optimizes the arrangement of data between different butterfly units, so that only half of the data needs to be moved in the data movement phase after each butterfly operation phase to ensure the normal implementation of the overall function, thereby greatly reducing the burden on the data router.

[0085] Based on the above-mentioned NTT accelerator, this embodiment further provides a chip, which includes the above-mentioned NTT accelerator and thus has corresponding functions and beneficial effects.

[0086] This embodiment also provides an electronic device in which the above-mentioned chip is installed. The electronic device includes electronic products such as computers, smart terminals and servers.

[0087] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0088] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.

Claims

1. An SRAM storage and computing integrated NTT accelerator, characterized in that: include: 128 memory-computing butterfly units, 1 array decoder, 1 controller, 1 modular simple unit, and 1 inter-column data router; The storage-computation integrated butterfly unit is used to accelerate the NTT algorithm, implement two butterfly operations, and store the input data of a butterfly operation and the intermediate data generated during the calculation process; The array decoder and controller are used to control the working state of each storage-computation-in-one butterfly unit when performing butterfly operations; The modular reduction simple element is used to process the modular operation of the 24-bit wide data output by the multiplier after completing the 12×12-bit multiplication. The inter-column data router is connected to the input and output of each butterfly unit, and is used to write the initial data into the butterfly unit and transmit the calculation results of each stage to the corresponding unit in preparation for the next stage of operation.

2. The SRAM storage and computing integrated NTT accelerator according to claim 1, characterized in that: The memory-computation integrated butterfly unit includes an SRAM memory array circuit and a near-memory logic operation circuit; The SRAM memory array includes 12×12 6T-SRAM memory cells, a total of 144 cells, 12 pre-charge circuits, 12 write circuits, and 12 read circuits; 144 6T-SRAM memory cells are used to store 12 12-bit data in rows; One precharge circuit is shared by a column of 1×12 6T-SRAM and is used to charge the bit line to a high level; One write circuit is shared by a column of 1×12 6T-SRAM and is used to write data into the SRAM memory cell; One read circuit is shared by a column of 1×12 6T-SRAM and is used to read data from the SRAM memory cells; The near memory logic operation circuit includes 12 sense amplifiers, 1 data driver, and 1 two-to-one multiplexer. 1 unified adder, 1 range selector; The sense amplifier is used to collect the 12-bit data read out of the SRAM storage array every cycle; The data driver is used to drive the intermediate results generated during the calculation process to the write control port of the SRAM storage array; the range selector is placed after the unified adder and is used to select the result after the unified adder operation as the final result output; The two-to-one multiplexer is used to select the working mode of the circuit. The circuit is controlled to work in CT mode or GS mode through a control signal. When the control signal is 0, the circuit works in CT mode, and the entire architecture is used to perform NTT operations. When the control signal is 1, the circuit works in GS mode, and the architecture is used to perform INTT operations.

3. The SRAM storage and computing integrated NTT accelerator according to claim 2, characterized in that: The unified adder integrates the hardware resources of the adder that need to be reused in the butterfly operation process, specifically implementing 12-bit multiplication, fast modulo operation and modulo addition in the butterfly operation to perform addition operations, which are selected by the control signal sel; When sel = 0, the unified adder acts as an accumulator for 12-bit × 12-bit multiplication. In this mode, the unified adder accumulates the shifted 12-bit data in each clock cycle to implement a 12-bit by 12-bit multiplication operation. When sel=1, the unified adder also works as an accumulator, performing Mod 3329 reduction on the input data; When sel=2, the unified adder and the range selector module together form a modular adder; in this mode, the unified adder and the range selector work together to perform modular addition operations.

4. The SRAM storage and computing integrated NTT accelerator according to claim 1, characterized in that: The modular reduction is based on a constant-time reduction method using recursive algorithm 2. 12 ≡2 9 -2 8 +1(mod3329) converts the number larger than 12 bits in the NTT operation into several numbers smaller than 12 bits on the integer ring.

5. The SRAM storage and computing integrated NTT accelerator according to claim 4, characterized in that: The weight is extracted from the recursive result by introducing the calculation method of XOR between the weight and the corresponding bit. The weight extraction method is: in units of rows, the weight of the inverted bit is recorded as 1, and the weight of the inverted bit is recorded as 0.

6. The SRAM storage and computing integrated NTT accelerator according to claim 1, characterized in that: The specific steps of NTT calculation are as follows: 1) Preparation phase: 256 polynomial coefficients and 128 twiddle factors are pre-stored in the memory structure of the corresponding butterfly unit. A single butterfly unit needs to pre-store 8 twiddle factors, corresponding to 8 stages of calculation. 2) First-stage operation: Each butterfly unit calls the data at a specified location in the unit's storage array to start the butterfly operation. When the first-stage butterfly operation is completed, the calculation result is written back to the previous address. 3) Data migration: The data router writes the initial data into the butterfly unit and transmits the calculation results of this stage to the corresponding unit, adjusting the data position between units to prepare for subsequent operations; 4) Similarly, after completing 8 stages of butterfly operations, the final result can be output.

7. The SRAM storage and computing integrated NTT accelerator according to claim 1, characterized in that: The inter-column data router adopts an efficient routing method: as much data as possible is left in the original location, and only one data in each butterfly structure needs to be transferred.

8. A chip, characterized in that: It includes an SRAM storage and computing integrated NTT accelerator as described in any one of claims 1-7.

9. An electronic device, characterized in that: Comprising the chip as claimed in claim 8.

Citation Information

Patent Citations

  • Efficient lightweight NTT multiplier circuit based on lattice cipher

    CN115756386A

  • High-performance polynomial multiplication hardware acceleration architecture for lattice cryptographic chip

    CN118963703A