Dspless polynomial multiplier suitable for falcon signature verification

CN122672747APending Publication Date: 2026-09-01HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610823763.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

已有的部分优化工作通过利用模数特性节省一个DSP,但底层的14位×14位乘法运算仍依赖另一个DSP;另有工作通过设计专用位宽部分积乘法器替代DSP,将14位乘法拆分为若干低位宽乘法,但带来了关键路径过长、最大工作频率下降等问题,难以实现面积与速度的良好折中

Benefits of technology

1、本发明提出多项式乘法器顶层四模块协同架构,包括控制模块、NTT/INTT模块、PPM模块和FIFO缓存模块四个相互独立的顶层子模块。其中,FIFO缓存模块以四个512深度的先进先出(First-In First-Out,FIFO)存储器在NTT/INTT模块与PPM模块之间建立流水线交叠的数据通路,使Falcon签名验证所需的两轮正向NTT变换、逐点模乘(Point-wisePolynomial Multiplication,PPM)与逆向INTT变换沿时间轴交叠执行,大幅压缩了多项式乘法器的整体计算延时。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122672747A_ABST
    Figure CN122672747A_ABST
Patent Text Reader

Abstract

This invention discloses a DSP-free polynomial multiplier suitable for Falcon signature verification, comprising: a control module, an NTT / INTT module, a PPM module, and a FIFO buffer module. The NTT / INTT module adopts a fully pipelined, non-storage-based R2-MDC architecture, containing ten cascaded configurable BFUs, and using LUT RAM memory instead of BRAM to store rotation factors. The PPM module contains two parallel DSP-free simplified Barrett modular multiplication units, performing point-by-point modular multiplication of frequency domain coefficients. The FIFO buffer module buffers data between modules, allowing the two rounds of NTT, point-by-point modular multiplication, and INTT to overlap and pipeline along the time axis. The DSP-free simplified Barrett modular multiplication unit calls the Karatsuba-Vedic DSP-free 14-bit multiplication unit to complete the underlying multiplication and coordinate error correction. This invention achieves zero DSP and zero BRAM resource consumption for the entire polynomial multiplier, significantly reducing latency and area-time product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of post-quantum cryptography hardware acceleration technology, specifically relating to a hardware accelerator suitable for polynomial multiplication operations in the signature verification process of the Falcon digital signature algorithm. It covers a fully pipelined non-storage number theoretic transform (NTT) module, an inverse number theoretic transform (INTT) module, a Barrett modular multiplication unit with zero digital signal processor (DSP) consumption, and a Karatsuba-Vedic DSP-free multiplier design. Background Technology

[0002] With the rapid development of quantum computing technology, the widely deployed public key infrastructure faces severe threats. The RSA algorithm, based on the integer factorization problem, and elliptic curve cryptography, based on the discrete logarithm problem, can be efficiently cracked by Shor's quantum algorithm. To address this security crisis, the National Institute of Standards and Technology (NIST) launched the Post-Quantum Cryptography (PQC) standardization project in 2016. Falcon, one of the selected lattice-based digital signature schemes, has played a core role in the post-quantum era. In the Falcon signature verification process, polynomial multiplication is the core computational operation, typically implemented through NTT, INTT, and pointwise modular multiplication. NTT / INTT and modular multiplication units together constitute the computational core of the entire algorithm.

[0003] Existing NTT hardware architectures are mainly divided into two categories: memory-based in-situ iterative architectures and pipelined multipath delay commutator (MDC) architectures. In-situ iterative architectures consume fewer resources but have complex addressing logic. When processing two consecutive NTT rounds in the Falcon algorithm, the second NTT round must wait for the first round to finish completely, resulting in low pipeline efficiency. MDC pipelined architectures have low latency, but traditional implementations typically store the rotation factor in Block Random Access Memory (BRAM). However, the BRAM capacity of Xilinx 7 series Field-Programmable Gate Arrays (FPGAs) is fixed at 18 Kb or 36 Kb, leading to significant resource waste for storing rotation factors with small capacity requirements and introducing a higher equivalent silicon area.

[0004] In terms of modular multiplication implementation, Falcon uses a prime modulus q=12289, while the traditional Barrett reduction pre-computation constant is 21845. This constant cannot be implemented by simple shifting and addition / subtraction, and usually requires the FPGA's internal DSP to perform the multiplication. Some existing optimizations save one DSP by utilizing the modulus characteristics, but the underlying 14-bit × 14-bit multiplication operation still relies on another DSP. Other work replaces the DSP by designing a dedicated bit-width partial product multiplier, splitting the 14-bit multiplication into several low-bit-width multiplications, but this brings problems such as excessively long critical paths and a decrease in maximum operating frequency, making it difficult to achieve a good trade-off between area and speed. Summary of the Invention

[0005] To address the shortcomings of the existing technologies, this invention proposes a DSP-free polynomial multiplier suitable for Falcon signature verification. The aim is to reduce computational latency, optimize the area-time product (ATP) index, and ensure the correctness of modular multiplication operations while achieving a completely zero-DSP and zero-BRAM resource consumption for the entire polynomial multiplier.

[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The present invention provides a DSP-free polynomial multiplier suitable for Falcon signature verification, characterized in that it includes: a control module, an NTT / INTT module, a PPM module, and a FIFO cache module; The control module generates a start signal start_ntt for NTT mode and a start signal start_intt for INTT mode, as well as a main counter signal cnt, according to the current working mode. These signals are used to control the NTT / INTT module, PPM module, and FIFO buffer module to work together. The current working mode is either NTT mode, PPM mode, or INTT mode. The NTT / INTT module adopts a fully pipelined non-storage R2-MDC architecture, containing ten cascaded configurable butterfly operation units BFU0 to BFU9, and storing all rotation factors in a LUT RAM memory. It is used to receive 1024-point polynomial coefficient pairs from external input. After performing a forward NTT transformation in NTT mode, the tenth configurable butterfly operation unit BFU9 outputs the frequency domain coefficients. After performing an inverse INTT transformation in INTT mode, the first configurable butterfly operation unit BFU0 outputs the time domain coefficients. The PPM module includes two independent and parallel DSP-free simplified Barrett modular multiplication units. Each DSP-free simplified Barrett modular multiplication unit receives two sets of frequency domain coefficients from the FIFO buffer module in each clock cycle and returns a modular multiplication result. The two DSP-free simplified Barrett modular multiplication units perform point-by-point modular multiplication operations in parallel and write the modular multiplication result back to the FIFO buffer module. The FIFO cache module includes four FIFO memories, FIFO0 to FIFO3. The first FIFO memory, FIFO0, and the second FIFO memory, FIFO1, are used to cache two sets of frequency domain coefficients of the NTT / INTT module in NTT mode and serve as inputs to the PPM module. The third FIFO memory, FIFO2, and the fourth FIFO memory, FIFO3, are used to cache the modular multiplication result of the PPM module and serve as inputs to the NTT / INTT module in INTT mode.

[0007] The DSP-free polynomial multiplier for Falcon signature verification described in this invention is also characterized by the following operation of the NTT / INTT module: The main counting signal cnt generated by the control module synchronously drives the local counters of the ten configurable butterfly operation units BFU0 to BFU9. Each local counter starts counting according to the main counting signal cnt. The highest bit c[MSB] of the local counter is used as a flag signal to drive the rotation factor address generation unit and selection switching unit inside the control module. The rotation factor address generation unit provides the corresponding rotation factor for each configurable butterfly operation unit according to the address index, and all rotation factors are stored in the LUT RAM memory; The selection switching unit generates a selection signal sel, which, in conjunction with the two sets of registers inside the control module, routes the polynomial coefficients to the input of the corresponding configurable butterfly operation unit. In NTT mode, sel selects the first register group, and each configurable butterfly operation unit switches to CT butterfly operation function through the mode selection signal mode. The polynomial coefficients flow along the forward path from BFU0 to BFU9. The output of each configurable butterfly operation unit is sent to the next configurable butterfly operation unit after being delayed and buffered, until BFU9 outputs the frequency domain coefficients, thereby completing the forward NTT conversion. In INTT mode, sel switches to the second register group, and each configurable butterfly operation unit switches to the GS butterfly operation function through the mode selection signal mode. The polynomial coefficients flow in reverse path from BFU9 to BFU0. The output of each configurable butterfly operation unit is sent to the next configurable butterfly operation unit after being delayed and buffered, until BFU0 outputs the time domain coefficients, thus completing the reverse INTT transformation.

[0008] Furthermore, each configurable butterfly operation unit in the NTT / INTT module includes a DSP-free simplified Barrett modular multiplication unit, a modular addition unit, a modular subtraction unit, and a modular division by 2 module, which can be switched to a CT butterfly structure or a GS butterfly structure via a mode selection signal: When mode=0, a single configurable butterfly operation unit switches to NTT mode and is configured for CT butterfly operation function: the input operand B is first multiplied with the rotation factor W by the DSP-free simplified Barrett modular multiplication unit, and the resulting multiplication is sent in parallel to the modular addition unit and the modular subtraction unit through a multiplexer, so as to perform operations with the input operand A respectively, to obtain the first output operand U=(A+B×W) mod q and the second output operand V=(AB×W) mod q; where q is the Falcon modulus and mod represents the modulo operation; When mode=1, a single configurable butterfly operation unit switches to INTT mode and is configured as a GS butterfly operation function: input operands A and B are first fed into the modulus addition unit and the modulus subtraction unit in parallel. The modulus subtraction result output by the modulus subtraction unit is input into the DSP-free simplified Barrett modulus multiplication unit and multiplied with the inverse rotation factor W_inv. The multiplication result and the output of the modulus addition unit are both scaled by the modulus divide by 2 module. Finally, the first output operand U=(A+B) / 2 mod q and the second output operand V=(AB)×W_inv / 2 mod q, where W_inv is the modulus inverse of W. The modulus-by-2 module utilizes the property that the Falcon modulus q is odd and is calculated as follows: When the least significant bit D[0] of the input data D is 0, D is directly right-shifted by 1 bit and then output as D / 2; When D[0]=1, first calculate D+q, then shift right by 1 bit and output (D+q) / 2.

[0009] Furthermore, the DSP-free simplified Barrett modular multiplication unit in a single configurable butterfly arithmetic unit employs a 5-stage pipeline and processes the two 14-bit operands M and N as follows: The first-stage pipeline calls the Karatsuba-Vedic DSP-free 14-bit multiplication unit to calculate M and N, obtains the 28-bit product P_K, and latches it in a register; The second-stage pipeline extracts the high 18 bits of P_K from the register, denoted as the high-order word c1, and the low 10 bits, denoted as the low-order word c0. It also takes the low 6 bits of c1 as the low-order truncation value c1p. It performs two shifts and additions on c1 using the first hardware factor f1 to obtain the first intermediate accumulation variable d1, d1=(c1>>2)+(c1>>4); where >> indicates a right shift operation. The third-level pipeline performs a shift and addition operation on d1 using the second hardware factor f2 to obtain the second intermediate accumulated variable d2 = (d1 >> 4) + (d1 >> 8); then adds d1 and d2 to obtain the third intermediate accumulated variable d3; shifts d3 right by 2 bits to obtain the estimated quotient Q'; thus obtaining the remainder estimate r' = c1p - (Q'[2:0] << 3) - (Q'[3:0] << 2) in 6-bit representation; where Q'[2:0] represents the lower 3 bits of the estimated quotient Q', Q'[3:0] represents the lower 4 bits of the estimated quotient Q', and the first hardware factor f1 and the second hardware factor f2 are both represented as the sum of several powers of 2; The fourth-stage pipeline calculates the pre-correction result R'={r', c0}-Q', and simultaneously pre-calculates three candidate correction values ​​R'-q, R'-2q, and R'-3q. A 2-bit selection signal r_sel is generated based on the comparison results of R' with q, 2q, and 3q, respectively. The four candidate values ​​are then latched together with r_sel. Here, { ,} denotes concatenation. The fifth-level pipeline selects the final modular multiplication result R from four candidate values ​​R', R'-q, R'-2q, and R'-3q based on r_sel using a pure 4-to-1 multiplexer, to ensure that R falls within the interval [0, q).

[0010] Furthermore, the Karatsuba-Vedic DSP-free 14-bit multiplication unit in the DSP-free simplified Barrett modular multiplication unit processes two 14-bit operands M and N and outputs a 28-bit product P_K as follows: The Karatsuba decomposition strategy is used to split the 14-bit operand M into the high 7 bits M_H=M[13:7] and the low 7 bits M_L=M[6:0], and the 14-bit operand N is split into the high 7 bits N_H=N[13:7] and the low 7 bits N_L=N[6:0]; The Vedic 8-bit multiplication unit is called to expand the three sub-multiplications in parallel: the low-order product P_LL=M_L×N_L, the high-order product P_HH=M_H×N_H, and the mixed product P_MID=S_M×S_N, where S_M=M_H+M_L is the sum of two segments of M, and S_N=N_H+N_L is the sum of two segments of N. The maximum value of S_M and S_N is 254, and they are both represented by 8-bit unsigned numbers. Using the Karatsuba mathematical property, calculate the cross term C_r = P_MID - P_LL - P_HH; The lower 14 bits of P_HH are placed in the [27:14] bit field of P_K, the lower 15 bits of C_r are placed in the [21:7] bit field of P_K, and the lower 16 bits of P_LL are placed in the [15:0] bit field of P_K. The three bit fields are then superimposed at their respective bit weight positions to output a 28-bit P_K.

[0011] Furthermore, the Vedic 8-bit multiplication unit in the Karatsuba-Vedic DSP-free 14-bit multiplication unit receives two 8-bit operands E and F, and processes them as follows: Divide the 8-bit operand E into four groups of two bits each: E0, E1, E2, and E3, corresponding to the [1:0], [3:2], [5:4], and [7:6] bit fields of E, respectively. Similarly, divide the 8-bit operand F into four groups of two bits each: F0, F1, F2, and F3, corresponding to the [1:0], [3:2], [5:4], and [7:6] bit fields of F, respectively. By using AND gates and adders, each group of E is partially multiplied by the four groups of F. After performing 16 2-bit × 2-bit partial product operations, the partial product {T_ij|i,j=0,1,2,3} is obtained in 4-bit representation, where T_ij represents the partial product of the i-th group of E and the j-th group of F. The 16 partial products {T_ij|i,j=0,1,2,3} are aggregated into 7 anti-diagonal accumulation terms G0 to G6 according to the anti-diagonal rule. Among them, the (k+1)th anti-diagonal accumulation term Gk aggregates all partial products that satisfy the subscript condition i+j=k, where k takes values ​​from 0 to 6. G0 to G6 are mapped to the 16-bit product space in sequence according to their starting bits 0, 2, 4, 6, 8, 10, and 12. Then, they are vertically accumulated using a concatenation-then-addition strategy as follows: In the hardware circuit, the fourth anti-diagonal accumulator G3 and the first anti-diagonal accumulator G0 are concatenated into a 16-bit concatenated row Row1 by purely physical connections according to their bit weights; the fifth anti-diagonal accumulator G4 and the second anti-diagonal accumulator G1 are concatenated into a 16-bit concatenated row Row2 by their bit weights; the sixth anti-diagonal accumulator G5 and the third anti-diagonal accumulator G2 are concatenated into a 16-bit concatenated row Row3 by their bit weights; and the seventh anti-diagonal accumulator G6 is zero-padded into a 16-bit concatenated row Row4. After performing the first 16-bit addition on Row1 and Row2, the first-stage sum S12 is obtained; after performing the second 16-bit addition on Row3 and Row4, the second-stage sum S34 is obtained; and after performing the third 16-bit addition on S12 and S34, the product P_V is obtained.

[0012] Furthermore, the workflow of the PPM module is as follows: The two DSP-free simplified Barrett modular multiplication units of the PPM module read two frequency domain coefficients from FIFO0 and FIFO1 in each clock cycle, perform two sets of point-by-point modular multiplications in parallel, obtain two product results, and write them into FIFO2 and FIFO3 in sequence. After the PPM module completes all 1024-point modular multiplication, it sends a completion signal to the control module. The control module then pulls the start_intt signal high, triggering the NTT / INTT module to read data from FIFO2 and FIFO3 and perform the inverse INTT transformation.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes a top-level four-module collaborative architecture for a polynomial multiplier, comprising four independent top-level sub-modules: a control module, an NTT / INTT module, a PPM module, and a FIFO cache module. The FIFO cache module uses four 512-depth First-In-First-Out (FIFO) memories to establish overlapping pipeline data paths between the NTT / INTT and PPM modules. This allows the two rounds of forward NTT transformation, point-wise polynomial multiplication (PPM), and inverse INTT transformation required for Falcon signature verification to be executed overlapping along the time axis, significantly reducing the overall computational latency of the polynomial multiplier.

[0014] 2. This invention designs a Karatsuba-Vedic DSP-free 14-bit multiplication unit as the underlying computational unit for a simplified Barrett modular multiplication unit without DSP. This unit uses Karatsuba decomposition to break down the 14-bit × 14-bit multiplication into three parallel Vedic 8-bit multiplications. The Vedic 8-bit multiplication unit compresses the vertical accumulation into only three 16-bit additions through a 2-bit grouping and concatenation-then-add strategy, eliminating the carry-lookahead adder chain. Furthermore, an approximate Barrett reduction factor is introduced, transforming high-order multiplication into pure shift addition and subtraction operations, ultimately achieving a completely zero-DSP consumption for the entire modular multiplication unit.

[0015] 3. This invention employs a fully pipelined, non-storage-based radix-2 multipath delay commutator (R2-MDC) architecture, replacing BRAM memory with look-up table random access memory (LUT RAM) to store rotation factors, reducing BRAM consumption to zero. Combined with a completely DSP-free modular multiplication design, the entire polynomial multiplier achieves zero DSP and zero BRAM consumption, achieving optimal ATP among similar designs. Attached Figure Description

[0016] Figure 1 This is a diagram of the overall hardware architecture of the polynomial multiplier of the present invention. Figure 2 This is a structural diagram of the NTT / INTT module based on the R2-MDC pipeline architecture of this invention; Figure 3 This is a comparison chart of cycle consumption between the MDC architecture and the in-situ architecture for polynomial multiplier indexing and two-round NTT computation of the present invention; Figure 4 A butterfly data flow diagram for NTT-CT transform and INTT-GS transform; Figure 5 This is a hardware circuit diagram of the configurable butterfly arithmetic unit (BFU) of the present invention; Figure 6 This invention provides a simplified Barrett modular multiplication unit circuit diagram without DSP. Figure 7 This is a block diagram of the Karatsuba-Vedic DSP-free 14-bit multiplication unit structure of the present invention; Figure 8 This is a circuit diagram of the Vedic 8-bit multiplier of the present invention; Figure 9 This diagram shows a comparison between the Vedic 8-bit multiplier circuit of this invention and the strategy of concatenating before adding, and the traditional carry-lookahead addition scheme. Figure 10This is a hardware structure diagram of the NTT / INTT based on the in-situ iterative architecture of this invention. Detailed Implementation

[0017] In this example, as Figure 1 As shown, a DSP-free polynomial multiplier suitable for Falcon signature verification includes four independent top-level sub-modules: a control module, an NTT / INTT module, a PPM module, and a FIFO cache module.

[0018] The control module dynamically generates start signals start_ntt and start_intt, as well as the main counter signal cnt, based on the current operating mode (NTT mode, PPM mode, or INTT mode), driving the other three sub-modules to work collaboratively. The NTT / INTT module adopts a fully pipelined, non-storage-based R2-MDC architecture, containing ten cascaded BFUs (BFU0 to BFU9) to perform forward NTT or inverse INTT transformations on polynomials of length N=1024 in Falcon. The PPM module contains two independent, parallel, DSP-free simplified Barrett modular multiplication units, performing point-by-point modular multiplication on the frequency domain coefficients output by the NTT / INTT module. The FIFO buffer module contains four 512-depth FIFO memories (FIFO0 to FIFO3) to buffer intermediate frequency domain data between the NTT / INTT module and the PPM module. Through this top-level four-module collaborative architecture, two rounds of forward NTT transformation, point-by-point modular multiplication, and inverse INTT transformation are executed in an overlapping pipeline along the time axis, significantly reducing the overall computational latency of the polynomial multiplier.

[0019] In this example, such as Figure 2 As shown, the NTT / INTT module adopts a fully pipelined, non-memory-based R2-MDC architecture, containing ten cascaded BFUs (BFU0 to BFU9). In the Falcon algorithm, the polynomial length N=1024, requiring 10 stages of iterative computation to complete the NTT and INTT transformation. The control module generates start signals start_ntt and start_intt. The main counter counts according to the start signals and outputs a count signal cnt. Each BFU is equipped with a local counter, which controls the address generation and selection switching logic of the rotation factor based on the cnt signal. The highest bit c[MSB] of the local counter serves as a flag signal to drive the relevant control timing. The rotation factor table provides rotation factors for each level of BFU according to the address index. All rotation factors are stored in LUT RAM memory, eliminating the use of BRAM memory and thus eliminating the occupation of on-chip BRAM. The gating control unit generates a selection signal sel, which, in conjunction with two sets of registers, routes data to the input of the corresponding BFU.

[0020] like Figure 3As shown, in the forward NTT transformation, data flows from BFU0 to BFU9. Taking the NTT transformation of the public key pk as an example, the first stage processes data pairs with an index difference of 512. Given that the system inputs two data points per cycle, processing the input of a 1024-point polynomial takes 512 clock cycles, and processing two polynomials requires a total of 1024 clock cycles. The second stage reduces the index difference to 256, requiring an alignment delay of 256 cycles. This continues, with the buffer delay decreasing logarithmically from 512 to 1 clock cycle at the tenth stage. Each BFU requires 8 clock cycles to complete one data calculation. Adding the 2 clock cycles required for the multiplexer and control circuitry, the total pipeline delay from BFU0 to BFU9 is 601 clock cycles. The total execution time for two forward NTT transformations is 1024 + 601 + 8 = 1633 clock cycles.

[0021] like Figure 4 As shown, the NTT stage uses Cooley-Tukey (CT) butterfly operation, with the formulas: U = (A + B × W) mod q, V = (AB × W) mod q; the INTT stage uses Gentleman-Sande (GS) butterfly operation, with the formulas: U = (A + B) / 2 mod q, V = (AB) × W_inv / 2 mod q, where W_inv is the modulo inverse of W. The two are dual in their operational structures and can be implemented in the same BFU hardware through mode switching. The INTT direction is the reverse flow of data from BFU9 to BFU0. The inverse transformation requires 1023 cycles to fill the pipeline, 90 cycles for inter-stage control overhead, and 8 cycles to clear the pipeline, totaling 1121 clock cycles.

[0022] In this example, such as Figure 5 As shown, each BFU in the NTT / INTT module is a configurable butterfly operation unit. The BFU contains a DSP-free simplified Barrett modular multiplication unit, a modular addition unit, a modular subtraction unit, and a modular division by 2 module. The switching between NTT mode (CT butterfly) and INTT mode (GS butterfly) is realized through the mode selection signal mode.

[0023] In NTT mode (mode=0): Input operand B first enters the DSP-free simplified Barrett modular multiplication unit and is multiplied by the rotation factor W. The product is sent in parallel to the modular addition and modular subtraction units through a multiplexer, and is operated with the input operand A respectively. The direct outputs are U=(A+B×W) mod q and V=(AB×W) mod q.

[0024] In INTT mode (mode=1): Input operands A and B are first fed into the modulo addition and modulo subtraction units in parallel. The output of the modulo subtraction unit then enters the DSP-free simplified Barrett modulo multiplication unit and is multiplied by the modulo inverse of the inverse twitch factor W_inv. The outputs of the addition and multiplication paths are scaled by the modulo-2 module, and the final outputs are U=(A+B) / 2 mod q and V=(AB)×W_inv / 2 mod q. The modulo-2 module takes advantage of the fact that q=12289 is an odd number: when the least significant bit D[0] of the input data D is 0, the output D>>1=D / 2; when D[0]=1, D+q is calculated first (the result is even), and the output (D+q)>>1=(D+q) / 2.

[0025] In this example, such as Figure 6 As shown, the DSP-free simplified Barrett modular multiplication unit performs modular multiplication on a modulus q = 12289, taking two 14-bit operands M and N as inputs and outputting a 14-bit modular multiplication result R. In the Falcon algorithm, the operand width Wq = 14 bits, and the underlying product P_K has a width of 28 bits. The traditional Barrett reduction factor is denoted as ζ, where ζ = ceil(2 18 / q)=21845. Here, ceil represents rounding down. This constant cannot be implemented using bit shifting or addition / subtraction; traditional designs require a DSP to perform the multiplication.

[0026] This design replaces the original Barrett reduction factor ζ with a hardware-friendly constant. This hardware-friendly constant can be decomposed into a product of several powers of 2, thus transforming high-bit-width multiplication into shift-based addition and subtraction operations. In this embodiment, an approximate value of 21840 is used as the hardware-friendly constant, and its decomposition is 21840 = (2 2 +1)×(2 8 +2 4 +1)×2 4 The three factors are denoted as h1, h2, and h3, respectively. Further derivation yields two hardware factors, f1 and f2, allowing the calculation of c1 × 21840 to be decomposed into: d1 = c1 × f1, d2 = d1 × f2, where f1 and f2 can both be expressed as the sum of several powers of 2. The entire process is implemented using bit shifting and addition / subtraction. Based on this, a 5-stage pipeline structure, error boundary derivation, and a 4-to-1 multiplexer error correction circuit are employed to realize a simplified Barrett modular multiplication unit without a DSP.

[0027] The entire module unit adopts a 5-stage pipeline structure: The first-level latch calls the 28-bit product P_K calculated by the Karatsuba-Vedic's DSP-less 14-bit multiplication unit; Level 2 extracts the high 18 bits of the high-order word c1 and the low 10 bits of the low-order word c0 from P_K, and takes the low 6 bits of c1 as the low-order truncation value c1p. Then, it performs two shifts and additions on c1 with f1 to obtain the first intermediate accumulation variable d1, d1=(c1>>2)+(c1>>4). The third level uses f2 to shift and add d1 to obtain the second intermediate accumulation variable d2, d2=(d1>>4)+(d1>>8). Then, add d1 and d2 to obtain the third intermediate accumulation variable d3, d3=d1+d2. Shift d3 right by 2 bits to obtain the estimated quotient Q', i.e. Q'=d3>>2. Then, calculate the remainder estimate r' according to r'=c1p-(Q'[2:0]<<3)-(Q'[3:0]<<2), where r' is represented by 6 bits. The fourth-level parallel pre-computation pre-correction result R'={r', c0}-Q' and three candidate correction values ​​R'-q, R'-2q, and R'-3q are generated and latched simultaneously; The fifth stage outputs the final result R based on r_sel using a pure 4-to-1 multiplexer. This stage does not contain any subtraction or comparison logic, and the critical path is shortened to the delay of a single-stage multiplexer.

[0028] In this example, such as Figure 7 As shown, the Karatsuba-Vedic DSP-free 14-bit multiplication unit is responsible for calculating the product of two 14-bit unsigned integers M and N, outputting a 28-bit result P_K, without using any DSP resources. Its structure consists of two layers: an upper Karatsuba decomposition layer and a lower Vedic 8-bit multiplication layer.

[0029] In the Karatsuba decomposition layer, the 14-bit operands are split into high 7 bits and low 7 bits respectively: M_H=M[13:7] (high 7 bits of M), M_L=M[6:0] (low 7 bits of M); N_H=N[13:7] (high 7 bits of N), N_L=N[6:0] (low 7 bits of N).

[0030] The three-way multiplication is expanded in parallel: the low-order product P_LL = M_L × N_L, the high-order product P_HH = M_H × N_H, and the mixed product P_MID = S_M × S_N, where S_M = M_H + M_L is the sum of the two segments of M, and S_N = N_H + N_L is the sum of the two segments of N. The maximum value of S_M and S_N is 254, represented by 8-bit unsigned numbers, which meets the input bit width requirement of the Vedic 8-bit multiplication unit. The cross term C_r = P_MID - P_LL - P_HH, and the Karatsuba mathematical property guarantees that C_r is always non-negative, so no signed operation is required.

[0031] The final 28-bit output is achieved through bit concatenation: the lower 14 bits of P_HH are placed in the [27:14] bit segment of P_K, the lower 15 bits of C_r are placed in the [21:7] bit segment of P_K, and the lower 16 bits of P_LL are placed in the [15:0] bit segment of P_K. These three segments are superimposed at their corresponding bit weights, achieved through purely physical interconnects, requiring no additional shifters or adders. During synthesis, the constraint attribute `use_dsp` is set to `no` to ensure that the synthesis tool does not infer the DSP.

[0032] In this example, such as Figure 8 As shown, the Vedic 8-bit multiplier divides each of the two 8-bit operands into four groups of two bits each: groups E0 to E3 correspond to the [1:0], [3:2], [5:4], and [7:6] bit fields of operand E, respectively; groups F0 to F3 correspond to the [1:0], [3:2], [5:4], and [7:6] bit fields of operand F, respectively. The underlying reconstruction consists of 16 2-bit × 2-bit partial product operations, denoted as T_{ij}. Each 2-bit × 2-bit operation is implemented using an AND gate and an adder. The result is represented by 4 bits, with a maximum value of 9, eliminating the need for a dedicated multiplier.

[0033] The products of each part are aggregated into seven anti-diagonal cumulative terms G0 to G6 according to the weights i+j=k: G0=T_{00}, G1=T_{01}+T_{10}, G2=T_{02}+T_{11}+T_{20}, G3=T_{03}+T_{12}+T_{21}+T_{30}, G4=T_{13}+T_{22}+T_{31}, G5=T_{23}+T_{32}, G6=T_{33}. These seven cumulative terms are naturally non-overlapping in the 16-bit space, with their starting positions being bits 0, 2, 4, 6, 8, 10, and 12 respectively (calculated with a 2-bit grouping granularity).

[0034] like Figure 9 As shown, this invention employs a concatenation-then-add strategy to replace the traditional carry-lookahead adder tree scheme. Utilizing the non-overlapping bit weights of each Gk in the 16-bit space, G3 and G0 are concatenated into a 16-bit concatenation row Row1 by purely physical connections, G4 and G1 into a 16-bit concatenation row Row2, G5 and G2 into a 16-bit concatenation row Row3, and G6 is zero-padded into a 16-bit concatenation row Row4. No addition operations are required within each row. Adding Row1 to Row2 yields the first-stage addition sum S12, adding Row3 to Row4 yields the second-stage addition sum S34, and adding S12 to S34 outputs the product P_V. Only three 16-bit additions are required throughout the entire process. Compared to the 7-level addition latency of the traditional carry-lookahead adder tree, this scheme reduces the critical path depth to 3 levels, effectively shortening the combinational logic latency. The entire module ensures that the synthesis tool does not infer the DSP by setting the constraint attribute use_dsp ​​to no.

[0035] The operation error introduced by the above-mentioned Barrett reduction factor replacement is verified by theoretical derivation and full-coverage software simulation, and the error coefficient l traverses integer values in [0,3]. The theoretical derivation process is as follows. The deviation between the estimated quotient Q' and the real quotient Q is defined as the quotient deviation k, where k=Q-Q'. The maximum value of the high-order word c1 is analyzed: c1_max=ceil(12288 2 / 1024)=147456. On this basis, through continuous scaling derivation of the floor function inequality, we obtain k<c1_max / 49152+1+1 / 2+1911 / 4096, which simplifies to k<5, that is, k∈[0,4], and the maximum value of the quotient deviation k is 4. It can be known from the range of k that the maximum value of the estimated remainder r' does not exceed 60, which is sufficiently represented by 6 bits. The calculation is reduced to the operation between the lower 6-bit low-order truncated value c1p of the high-order word c1 and the lower 4 bits of the estimated quotient Q'. Further, an intermediate variable s is defined, which satisfies s×q=R'-R (where R is the real modular multiplication result), and s is the algebraic multiple factor of the modulus q that differs between R' and R. By imposing a floor function inequality constraint on the value range of the product P_K, it is proved that s∈{-1, 0}. Combining the two types of errors, the total error coefficient l=k+s, and l∈[0,3]. Based on this, a 4-to-1 multiplexer is designed, which selects the correct result from R', R'-q, R'-2q, and R'-3q according to the comparison results of R' with q, 2q, and 3q, ensuring that the final output R∈[0, q).

[0036] The error boundary derivation is based on the property of the floor function, and the key steps are as follows: the estimated quotient Q' and the real quotient Q satisfy the relation Q'≈c1×21840 / 2 28 , Q=ceil(c1×ζ / 2 18 ). The deviation of the two is k=Q-Q'. By applying a scaling inequality to the floor function, we get k<c1_max / 49152+1 / 2+1911 / 4096. Substituting c1_max=147456 for simplification gives k<5, that is, k∈[0,4].

[0037] Software simulation adopts Python enumeration verification: traversing all combinations of M, N∈[0, q) for a total of q²≈150 million test cases, and counting the actual value of l for each test case. The results all fall into the integer interval [0,3], which verifies the completeness of the theoretical derivation and ensures that the hardware 4-to-1 multiplexer circuit can correct errors correctly for all inputs.

[0038] Through accurate error bound derivation and correction mechanism, the present invention proves that the value range of the error coefficient l introduced by the approximate Barrett reduction factor is [0,3], designs a 4-to-1 multiplexer to eliminate all errors, and shortens the critical path to the delay of a single-stage multiplexer through a 5-stage pipeline, achieving zero DSP resource consumption while ensuring the correctness of modular multiplication.

[0039] The 4-to-1 multiplexer generates a 2-bit selection signal r_sel based on three groups of parallel comparison results of R' with q, 2q, and 3q. The encoding rule is: when r_sel=00, R' is output (when R'<q); when r_sel=01, R'-q is output (when q≤R'<2q); when r_sel=10, R'-2q is output (when 2q≤R'<3q); when r_sel=11, R'-3q is output (when R'≥3q). Since R'-q, R'-2q, and R'-3q have been precomputed in parallel in the 4th pipeline stage, the 5th stage only contains a pure combinational logic 4-to-1 multiplexer, and the critical path depth is equivalent to the delay of a single-stage LUT, making it the pipeline stage with the shortest delay in the entire modular multiplication unit.

[0040] In this example, the PPM module instantiates two independent and parallel DSP-free simplified Barrett modular multiplication units, and each modular multiplication unit reuses the aforementioned DSP-free simplified Barrett modular multiplication unit structure, which consumes no DSP resources throughout the whole process.

[0041] After the NTT / INTT module completes two rounds of forward NTT transformation on two polynomials (public key pk and signature s2) in NTT mode, the two groups of frequency-domain coefficients are stored in FIFO0 and FIFO1 of the FIFO buffer module respectively. The two DSP-free simplified Barrett modular multiplication units of the PPM module read 2 frequency-domain coefficients from FIFO0 and FIFO1 respectively in each clock cycle, execute 2 groups of point-wise modular multiplications in parallel, return 2 product results per cycle, and write the point-wise modular multiplication results into FIFO2 and FIFO3 sequentially. After the PPM module completes the point-wise modular multiplication of all 1024 points, it sends a completion signal to the control module, the control module pulls up the start_intt signal, and triggers the NTT / INTT module to read data from FIFO2 and FIFO3 and perform inverse INTT transformation. The two-way parallel processing capability of the PPM module is matched with the dual coefficient per cycle throughput of the NTT / INTT module, and with the help of inter-stage buffering of the FIFO buffer module, NTT transformation, point-wise modular multiplication and INTT transformation are executed in overlapping pipeline on the time axis, which greatly reduces the total delay of the polynomial multiplier.

[0042] As Figure 10As shown, to further verify the applicability of the DSP-free simplified Barrett modular multiplication unit under different architectures, this invention also implements an in-situ iterative NTT / INTT architecture. This architecture uses a dual-port RAM memory to store polynomial coefficients. Read and write addresses are generated by an address generation unit and a counter generation unit. The BFU reads the corresponding coefficient, performs a radix-2 butterfly operation, and writes the result back to the RAM memory. The bit-reversal unit rearranges the bit order of the output coefficients after iteration. This invention implements three configurations: one parallel channel (containing two BFUs), two parallel channels (containing four BFUs), and four parallel channels (containing eight BFUs). By adopting a general addressing scheme, memory access correctness is guaranteed, and the lowest latency achievable with this architecture is achieved. All BFUs call the DSP-free simplified Barrett modular multiplication unit, thus enabling a fair comparison between the in-situ iterative NTT and the MDC architecture with zero DSP consumption. Comparative experimental results show that the MDC architecture has a significant advantage over the in-situ iterative architecture in the area-time product of two rounds of NTT.

[0043] The comprehensive experimental results show that this design was successfully synthesized, placed, routed, and functionally simulated on Artix-7 xc7a200tfbg484-3 and Virtex-7xc7vx485tffg1157-2 FPGA platforms. The DSP-free simplified Barrett modular multiplication unit consumes 88 slices and 0 DSPs on the Artix-7 platform, with a maximum operating frequency of 270 MHz and an ATP of 1304; on the Virtex-7 platform, it consumes 43 slices and 0 DSPs, with a maximum operating frequency of 408 MHz and an ATP of 422. Compared to existing Barrett modular multiplication designs using DSPs, this design demonstrates significant advantages in both area and frequency. In a two-round NTT operation scenario, the MDC architecture (containing 10 BFUs) consumes 0 DSPs and 0 BRAMs on the Artix-7 platform, operates at a frequency of 255 MHz, and has a total latency of approximately 6.4 μs for two rounds of NTT, with an ATP of 13964. On the Virtex-7 platform, it operates at a frequency of 298 MHz, with a latency of 5.48 μs and an ATP of 11636, achieving the best ATP level among similar designs.

Claims

1. A DSP-free polynomial multiplier suitable for Falcon signature verification, characterized in that, include: Control module, NTT / INTT module, PPM module and FIFO buffer module; The control module generates a start signal start_ntt for NTT mode and a start signal start_intt for INTT mode, as well as a main counter signal cnt, according to the current working mode. These signals are used to control the NTT / INTT module, PPM module, and FIFO buffer module to work together. The current working mode is either NTT mode, PPM mode, or INTT mode. The NTT / INTT module adopts a fully pipelined non-storage R2-MDC architecture, containing ten cascaded configurable butterfly operation units BFU0 to BFU9, and storing all rotation factors in a LUT RAM memory. It is used to receive 1024-point polynomial coefficient pairs from external input. After performing a forward NTT transformation in NTT mode, the tenth configurable butterfly operation unit BFU9 outputs the frequency domain coefficients. After performing an inverse INTT transformation in INTT mode, the first configurable butterfly operation unit BFU0 outputs the time domain coefficients. The PPM module includes two independent and parallel DSP-free simplified Barrett modular multiplication units. Each DSP-free simplified Barrett modular multiplication unit receives two sets of frequency domain coefficients from the FIFO buffer module in each clock cycle and returns a modular multiplication result. The two DSP-free simplified Barrett modular multiplication units perform point-by-point modular multiplication operations in parallel and write the modular multiplication result back to the FIFO buffer module. The FIFO cache module includes four FIFO memories, FIFO0 to FIFO3. The first FIFO memory, FIFO0, and the second FIFO memory, FIFO1, are used to cache two sets of frequency domain coefficients of the NTT / INTT module in NTT mode and serve as inputs to the PPM module. The third FIFO memory, FIFO2, and the fourth FIFO memory, FIFO3, are used to cache the modular multiplication result of the PPM module and serve as inputs to the NTT / INTT module in INTT mode.

2. The DSP-free polynomial multiplier for Falcon signature verification according to claim 1, characterized in that, The NTT / INTT module operates as follows: The main counting signal cnt generated by the control module synchronously drives the local counters of the ten configurable butterfly operation units BFU0 to BFU9. Each local counter starts counting according to the main counting signal cnt. The highest bit c[MSB] of the local counter is used as a flag signal to drive the rotation factor address generation unit and selection switching unit inside the control module. The rotation factor address generation unit provides the corresponding rotation factor for each configurable butterfly operation unit according to the address index, and all rotation factors are stored in the LUT RAM memory; The selection switching unit generates a selection signal sel, which, in conjunction with the two sets of registers inside the control module, routes the polynomial coefficients to the input of the corresponding configurable butterfly operation unit. In NTT mode, sel selects the first register group, and each configurable butterfly operation unit switches to CT butterfly operation function through the mode selection signal mode. The polynomial coefficients flow along the forward path from BFU0 to BFU9. The output of each configurable butterfly operation unit is sent to the next configurable butterfly operation unit after being delayed and buffered, until BFU9 outputs the frequency domain coefficients, thereby completing the forward NTT conversion. In INTT mode, sel switches to the second register group, and each configurable butterfly operation unit switches to the GS butterfly operation function through the mode selection signal mode. The polynomial coefficients flow in reverse path from BFU9 to BFU0. The output of each configurable butterfly operation unit is sent to the next configurable butterfly operation unit after being delayed and buffered, until BFU0 outputs the time domain coefficients, thus completing the reverse INTT transformation.

3. The DSP-free polynomial multiplier for Falcon signature verification according to claim 1, characterized in that, Each configurable butterfly operation unit in the NTT / INTT module includes a DSP-free simplified Barrett modular multiplication unit, a modular addition unit, a modular subtraction unit, and a modular division by 2 module, which can be switched to either a CT butterfly structure or a GS butterfly structure via a mode selection signal: When mode=0, a single configurable butterfly operation unit switches to NTT mode and is configured for CT butterfly operation function: the input operand B is first multiplied with the rotation factor W by the DSP-free simplified Barrett modular multiplication unit, and the resulting multiplication is sent in parallel to the modular addition unit and the modular subtraction unit through a multiplexer, so as to perform operations with the input operand A respectively, to obtain the first output operand U=(A+B×W) mod q and the second output operand V=(AB×W) mod q; where q is the Falcon modulus and mod represents the modulo operation; When mode=1, a single configurable butterfly operation unit switches to INTT mode and is configured as a GS butterfly operation function: input operands A and B are first fed into the modulus addition unit and the modulus subtraction unit in parallel. The modulus subtraction result output by the modulus subtraction unit is input into the DSP-free simplified Barrett modulus multiplication unit and multiplied with the inverse rotation factor W_inv. The multiplication result and the output of the modulus addition unit are both scaled by the modulus divide by 2 module. Finally, the first output operand U=(A+B) / 2 mod q and the second output operand V=(AB)×W_inv / 2 mod q, where W_inv is the modulus inverse of W. The modulus-by-2 module utilizes the property that the Falcon modulus q is odd and is calculated as follows: When the least significant bit D[0] of the input data D is 0, D is directly right-shifted by 1 bit and then output as D / 2; When D[0]=1, first calculate D+q, then shift right by 1 bit and output (D+q) / 2.

4. The DSP-free polynomial multiplier for Falcon signature verification according to claim 3, characterized in that, The DSP-free simplified Barrett modular multiplication unit in a single configurable butterfly arithmetic unit employs a 5-stage pipeline and processes two 14-bit operands M and N as follows: The first-stage pipeline calls the Karatsuba-Vedic DSP-free 14-bit multiplication unit to calculate M and N, obtains the 28-bit product P_K, and latches it in a register; The second-stage pipeline extracts the high 18 bits of P_K from the register, denoted as the high-order word c1, and the low 10 bits, denoted as the low-order word c0. It also takes the low 6 bits of c1 as the low-order truncation value c1p. It performs two shifts and additions on c1 using the first hardware factor f1 to obtain the first intermediate accumulation variable d1, d1=(c1>>2)+(c1>>4); where >> indicates a right shift operation. The third-level pipeline performs a shift and addition operation on d1 using the second hardware factor f2 to obtain the second intermediate accumulated variable d2 = (d1 >> 4) + (d1 >> 8); then adds d1 and d2 to obtain the third intermediate accumulated variable d3; shifts d3 right by 2 bits to obtain the estimated quotient Q'; thus obtaining the remainder estimate r' = c1p - (Q'[2:0] << 3) - (Q'[3:0] << 2) in 6-bit representation; where Q'[2:0] represents the lower 3 bits of the estimated quotient Q', Q'[3:0] represents the lower 4 bits of the estimated quotient Q', and the first hardware factor f1 and the second hardware factor f2 are both represented as the sum of several powers of 2; The fourth-stage pipeline calculates the pre-correction result R'={r', c0}-Q', and simultaneously pre-calculates three candidate correction values ​​R'-q, R'-2q, and R'-3q. It then generates a 2-bit selection signal r_sel based on the comparison results of R' with q, 2q, and 3q, respectively, and latches the four candidate values ​​together with r_sel. Here, { ,} indicates concatenation. The fifth-level pipeline selects the final modular multiplication result R from four candidate values ​​R', R'-q, R'-2q, and R'-3q based on r_sel using a pure 4-to-1 multiplexer, to ensure that R falls within the interval [0, q).

5. The DSP-free polynomial multiplier for Falcon signature verification according to claim 4, characterized in that, The Karatsuba-Vedic DSP-free 14-bit multiplication unit in the DSP-free simplified Barrett modular multiplication unit processes two 14-bit operands M and N and outputs a 28-bit product P_K as follows: The Karatsuba decomposition strategy is used to split the 14-bit operand M into the high 7 bits M_H=M[13:7] and the low 7 bits M_L=M[6:0], and the 14-bit operand N is split into the high 7 bits N_H=N[13:7] and the low 7 bits N_L=N[6:0]; The Vedic 8-bit multiplication unit is called to expand the three sub-multiplications in parallel: the low-order product P_LL=M_L×N_L, the high-order product P_HH=M_H×N_H, and the mixed product P_MID=S_M×S_N, where S_M=M_H+M_L is the sum of two segments of M, and S_N=N_H+N_L is the sum of two segments of N. The maximum value of S_M and S_N is 254, and they are both represented by 8-bit unsigned numbers. Using the Karatsuba mathematical property, calculate the cross term C_r = P_MID - P_LL - P_HH; The lower 14 bits of P_HH are placed in the [27:14] bit field of P_K, the lower 15 bits of C_r are placed in the [21:7] bit field of P_K, and the lower 16 bits of P_LL are placed in the [15:0] bit field of P_K. The three bit fields are then superimposed at their respective bit weight positions to output a 28-bit P_K.

6. The DSP-free polynomial multiplier for Falcon signature verification according to claim 5, characterized in that, The Karatsuba-Vedic DSP-free 14-bit multiplication unit's Vedic 8-bit multiplication unit receives two 8-bit operands E and F, and processes them as follows: Divide the 8-bit operand E into four groups of two bits each: E0, E1, E2, and E3, corresponding to the [1:0], [3:2], [5:4], and [7:6] bit fields of E, respectively. Similarly, divide the 8-bit operand F into four groups of two bits each: F0, F1, F2, and F3, corresponding to the [1:0], [3:2], [5:4], and [7:6] bit fields of F, respectively. By using AND gates and adders, each group of E is partially multiplied by the four groups of F. After performing 16 2-bit × 2-bit partial product operations, the partial product {T_ij|i,j=0,1,2,3} is obtained in 4-bit representation, where T_ij represents the partial product of the i-th group of E and the j-th group of F. The 16 partial products {T_ij|i,j=0,1,2,3} are aggregated into 7 anti-diagonal accumulation terms G0 to G6 according to the anti-diagonal rule. Among them, the (k+1)th anti-diagonal accumulation term Gk aggregates all partial products that satisfy the subscript condition i+j=k, where k takes values ​​from 0 to 6. G0 to G6 are mapped to the 16-bit product space in sequence according to their starting bits 0, 2, 4, 6, 8, 10, and 12. Then, they are vertically accumulated using a concatenation-then-addition strategy as follows: In the hardware circuit, the fourth anti-diagonal accumulator G3 and the first anti-diagonal accumulator G0 are concatenated into a 16-bit concatenated row Row1 by purely physical connections according to their bit weights; the fifth anti-diagonal accumulator G4 and the second anti-diagonal accumulator G1 are concatenated into a 16-bit concatenated row Row2 by their bit weights; the sixth anti-diagonal accumulator G5 and the third anti-diagonal accumulator G2 are concatenated into a 16-bit concatenated row Row3 by their bit weights; and the seventh anti-diagonal accumulator G6 is zero-padded into a 16-bit concatenated row Row4. After performing the first 16-bit addition on Row1 and Row2, the first-stage sum S12 is obtained; after performing the second 16-bit addition on Row3 and Row4, the second-stage sum S34 is obtained; and after performing the third 16-bit addition on S12 and S34, the product P_V is obtained.

7. The DSP-free polynomial multiplier for Falcon signature verification according to claim 1, characterized in that, The workflow of the PPM module is as follows: The two DSP-free simplified Barrett modular multiplication units of the PPM module read two frequency domain coefficients from FIFO0 and FIFO1 in each clock cycle, perform two sets of point-by-point modular multiplications in parallel, obtain two product results, and write them into FIFO2 and FIFO3 in sequence. After the PPM module completes all 1024-point modular multiplication, it sends a completion signal to the control module. The control module then pulls the start_intt signal high, triggering the NTT / INTT module to read data from FIFO2 and FIFO3 and perform the inverse INTT transformation.