A dilithium algorithm-based polynomial multiplier circuit
Patent Information
- Application Number
- CN202311145527.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-06
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-09-06
AI Technical Summary
[0003]Beckwith等人在Dilithium算法中提出了2×2的多项式乘法硬件电路,设计采用一次计算两轮蝶形运算的结构,这样8轮的蝶形运算被压缩到4轮,虽然减少了计算周期,但是会拉长关键路径,影响电路的最大主频频率;由于没有考虑复用的方式,设计中调用了两个多项式乘法模块,极大的增加了硬件面积开销
[0041] 1. Traditional NTT transforms employ single-round or two-round calculation methods in butterfly operations. Single-round calculation, with four inputs, requires a long computation cycle, which can be significant in large-scale polynomial multiplication operations. Two-round calculation, while effectively reducing the computation cycle by calculating data for two rounds within one clock cycle, lengthens the critical path and thus affects the circuit's maximum frequency. This invention uses a hybrid single/dual-round approach. The first six rounds of butterfly operations use an eight-input single-round calculation, while the last two rounds use a dual-round calculation. Internal registers are inserted for timing, effectively compressing the computation cycle while also considering the circuit's maximum clock frequency, thereby better reducing computation time and optimizing ATP parameters.
Smart Images

Figure CN117111883B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of polynomial multipliers, specifically a polynomial multiplier hardware circuit that is applied to improve speed, ensure data correctness and integrity. Background Technology
[0002] With the rapid development of the internet, frequent information exchange requires massive data processing. Public-key encryption and digital signatures necessitate numerous polynomial multiplications and addition / subtraction operations. As information processing becomes increasingly prominent, the design of efficient polynomial multipliers has gained growing attention from researchers. Directly multiplying polynomials in the constant domain results in excessive time complexity and enormous hardware implementation costs. Therefore, various processing methods are employed for the polynomial coefficients in the constant domain: DFT, FFT, and NTT transforms. These methods convert the polynomial coefficients in the constant domain to their corresponding processing domains for multiplication. Finally, the resulting polynomials are transformed back to the constant domain using IDFT, IFFT, and INTT transforms to obtain the final polynomial multiplication result.
[0003] Beckwith et al. proposed a 2×2 polynomial multiplication hardware circuit in the Dilithium algorithm. The design employed a two-round butterfly operation structure, compressing the eight-round butterfly operation into four rounds. While this reduced the computation cycle, it lengthened the critical path and affected the circuit's maximum clock frequency. Because it didn't consider reuse, it called two polynomial multiplication modules, significantly increasing hardware area overhead. To reduce area overhead, Zhao et al. proposed a reusable polynomial multiplication circuit compatible with both NTT and INTT conversions. It used registers to store intermediate data from the butterfly operation, effectively reducing area overhead, but still increasing the computation cycle. Summary of the Invention
[0004] To address the shortcomings of the existing technology, this invention proposes a polynomial multiplier circuit based on the Dilithium algorithm, which aims to improve the efficiency and reliability of polynomial multiplication operations while saving hardware resources, reducing the computation cycle, and increasing the operating frequency and throughput.
[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0006] The present invention provides a polynomial multiplier circuit based on the Dilithium algorithm, which includes: an NTT / INTT module, a point-by-point multiplication module, and a polynomial RAM group.
[0007] The NTT / INTT module includes: an NTT_INTT selection module, a vector address control module, a rotation factor address control module, a vector RAM group, a rotation factor ROM, and a BF module;
[0008] The point-by-point multiplication module includes: a matrix address generation module, a matrix RAM group, and a point-by-point multiplication module;
[0009] The polynomial RAM group contains t polynomial RAM blocks, wherein the depth of each polynomial RAM block is V;
[0010] The vector RAM group contains t vector RAM blocks, where the depth of each vector RAM block is V;
[0011] The matrix RAM group contains t matrix RAM blocks, where the depth of each matrix RAM block is M;
[0012] The rotation factor ROM contains a rotation factor data;
[0013] The NTT_INTT selection module acquires the mode input signal and selects the NTT operation mode or the INTT operation mode through the selector, and then sends the NTT enable signal / INTT enable signal to the vector address control module, the rotation factor address control module and the BF module respectively.
[0014] The vector address control module generates the corresponding read / write address of the vector polynomial based on the received NTT enable signal / INTT enable signal and sends it to the vector RAM group;
[0015] The rotation factor control module generates a corresponding rotation factor read address based on the received NTT enable signal / INTT enable signal and sends it to the rotation factor ROM;
[0016] The vector RAM group reads the corresponding x vector polynomial data according to the read / write address of the vector polynomial and sends them to the dot product module and the BF module respectively.
[0017] The rotation factor ROM reads the corresponding y rotation factor data according to the rotation factor read address and then sends it to the BF module;
[0018] The BF module selects the NTT operation mode or the INTT operation mode according to the received NTT enable signal / INTT enable signal, and performs Montgomery reduction, bit truncation, multiplication and addition / subtraction operations on the input x vector polynomial data and y twitch factor data in an orderly manner. After completing n rounds of butterfly operation, the NTT operation result is stored in the vector RAM group or the INTT operation result is stored in the polynomial RAM group.
[0019] The matrix address generation module generates the read address of the matrix RAM group based on the polynomial multiplication enable signal sent externally and sends it to the matrix RAM group;
[0020] The matrix RAM group reads the corresponding x matrix polynomial data according to the received read address and transmits them sequentially to the dot product module;
[0021] The dot product module performs multiplication and reduction operations on x matrix polynomial data and x vector polynomial data in an orderly manner, thereby storing the final calculation result in the polynomial RAM group.
[0022] The hardware circuit of the polynomial multiplier based on Dilithium described in this invention is characterized in that the NTT / INTT selection module first starts the NTT operation mode and sends the data of the NTT operation mode to the vector address control module and the rotation factor address control module respectively. After waiting for the NTT operation result obtained after the NTT transformation to complete the reduction and multiplication operation with the matrix polynomial data, the INTT operation mode is started, and the data of the INTT operation mode is resent to the vector address control module and the rotation factor address control module.
[0023] The vector address control module includes: a multi-address counting module, a butterfly operation round counting module, a vector address generation module, a vector base address generation module, and a vector address synthesis module;
[0024] The multi-address counting module receives the NTT enable signal / INTT enable signal and sends an enable input signal to the butterfly operation wheel counting module;
[0025] After receiving the enable input signal, the butterfly operation round number module starts recording the current butterfly operation round number and sends the round number signal to the multi-address counting module and the vector base address generation module respectively;
[0026] The multi-address counting module performs vector address counting on the address values of each round of butterfly operation based on the NTT enable signal / INTT enable signal and the round number signal, and then sends the address count data to the vector address generation module.
[0027] The vector address generation module calculates the initial address of the vector polynomial based on the address count data received for each round of butterfly operation and sends it to the vector address synthesis module.
[0028] The vector base address generation module calculates the number of p-term polynomials in the vector polynomial based on the received round number signal, thereby generating the base address of the vector polynomial and sending it to the vector address synthesis module;
[0029] The vector address synthesis module calculates the read / write address of the vector polynomial based on the initial address and the base address of the vector polynomial. The read address includes the effective read address and the effective read signal of the vector polynomial, and sends them to the vector RAM group and a register. The register outputs the effective write address and the effective write signal of the vector polynomial based on the effective read address and the effective read signal, and sends them to the vector RAM group.
[0030] The rotation factor address control module includes: an address increment / decrement counter module, a butterfly wheel increment / decrement counter module, and a rotation factor address generation module;
[0031] If the input signal received by the address increment / decrement counting module is an NTT enable signal, the twitch factor address starts to increment; if the input signal received is an INTT enable signal, the twitch factor address starts to decrement; thereby calculating the number of polynomial coefficients of the p terms in each round of butterfly operation and sending the twitch factor address count data to the twitch factor address generation module.
[0032] If the received enable signal is an NTT enable signal, the butterfly operation round count module starts counting from 1 to 8; if the received enable signal is an INTT enable signal, the butterfly operation round count starts counting from 8 to 1. This determines the current butterfly operation round number and the number of p-term polynomials processed, and sends the butterfly operation round number to the twitch factor address generation module.
[0033] The rotation factor address generation module calculates the effective read address and effective read signal of the rotation factor based on the received rotation factor address count data and the number of butterfly operation rounds, and sends them to the rotation factor ROM.
[0034] If the received mode signal is an NTT enable signal, the BF module delays x vector data through a register for one clock cycle to obtain x delayed vector data. The x delayed vector data and y twitch factor data are multiplied to obtain x multiplication results. These x multiplication results are then added / subtracted pairwise with the x vector data to obtain x added data and x subtracted data, which are sent to the first selector and the second selector, respectively. The first and second selectors, based on the mode input signal, output the x added data and x subtracted data as the NTT operation results over two clock cycles.
[0035] If the received mode signal is an INTT enable signal, the BF module delays x vector data through a register for one clock cycle to obtain x delayed vector data. Then, it performs pairwise addition / subtraction operations on the x vector data and x delayed vector data to obtain x added data and x subtracted data respectively. The x subtracted data are multiplied by y inverted twitch factor data to obtain x multiplied results. These x multiplied results are sent to the first selector, and the x added data are sent to the second selector. The first and second selectors, based on the mode input signal, output the x added data and x subtracted data as the INTT operation results over two clock cycles.
[0036] The matrix address generation module performs calculations sequentially starting from 0 based on the received polynomial multiplication enable signal, and sends the valid read signal and valid read address of the matrix to the matrix RAM group. It reads x matrix polynomial data in each clock cycle. When the coefficients of the last p-term matrix polynomial are read out, the matrix address generation module pulls the valid read signal low and sends it to the matrix RAM group, stopping the reading of data from the matrix RAM group.
[0037] The dot product module includes: a Montgomery reduction module and a reduction module;
[0038] The Montgomery reduction module multiplies the received vector polynomial data and matrix polynomial data. The multiplication result is then shifted left by 13 bits and left by 26 bits respectively to obtain two shifted results. The two shifted results are added to the multiplication result to obtain the first addition result. The result of left-shifting the multiplication result by 23 bits is subtracted from the first addition result to obtain the first subtraction result. The first subtraction result is left-shifted by 23 bits to obtain the first shifted result, and then added to the first subtraction result to obtain the second addition result. The first subtraction result is left-shifted by 13 bits and then subtracted from the second addition result and the multiplication result to obtain the output result m, which is then sent to the reduction module.
[0039] After receiving the output result m and performing a left shift operation of 22 bits, the reduction module adds the shifted result to the output result to obtain a third addition result. The third addition result is then left-shifted by 23 bits to obtain a third shift result. This third shift result is then left-shifted by 23 bits and added to the third shift result to obtain a fourth addition result. The third shift result is then left-shifted by 13 bits and subtracted from the fourth addition result and the output result m to obtain the polynomial multiplication output data mult.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] 1. Traditional NTT transforms employ single-round or two-round calculation methods in butterfly operations. Single-round calculation, with four inputs, requires a long computation cycle, which can be significant in large-scale polynomial multiplication operations. Two-round calculation, while effectively reducing the computation cycle by calculating data for two rounds within one clock cycle, lengthens the critical path and thus affects the circuit's maximum frequency. This invention uses a hybrid single / dual-round approach. The first six rounds of butterfly operations use an eight-input single-round calculation, while the last two rounds use a dual-round calculation. Internal registers are inserted for timing, effectively compressing the computation cycle while also considering the circuit's maximum clock frequency, thereby better reducing computation time and optimizing ATP parameters.
[0042] 2. Since NTT and INTT circuits are similar, this invention adopts a time-division multiplexing approach to design the NTT / INTT hardware circuit. The hardware circuit can be configured into NTT mode and INTT mode through the mode selection signal. This approach avoids calling two modules to implement NTT and INTT conversion, thereby effectively reducing area overhead.
[0043] 3. The vector RAM group and matrix RAM group designed in this invention each contain 8 RAM blocks. The data after NTT transformation of the vector RAM is written back to the vector RAM group, without calling additional RAM to store the NTT-transformed data. The method of using 8 RAM blocks to form a RAM group can ensure that 8 data read / write operations are performed in one clock cycle, achieving higher throughput. Attached Figure Description
[0044] Figure 1 This is a block diagram of the hardware circuit structure of the polynomial multiplier of the present invention;
[0045] Figure 2 This is a circuit structure diagram of the vector address control module of the present invention;
[0046] Figure 3 This is a circuit structure diagram of the rotation factor address control module of the present invention;
[0047] Figure 4 This is a circuit structure diagram of the BF module of the present invention;
[0048] Figure 5 This is a diagram illustrating the butterfly operation method and calculation cycle of the NTT transformation in this invention.
[0049] Figure 6 This is the overall operation timing diagram of the polynomial multiplier of the present invention;
[0050] Figure 7 This is the circuit diagram of the dot product module of the present invention. Detailed Implementation
[0051] In this example, as Figure 1 As shown, a polynomial multiplier circuit based on the Dilithium algorithm includes: an NTT / INTT module, a point-by-point multiplication module, and a polynomial RAM group;
[0052] The NTT / INTT module includes: NTT_INTT selection module, vector address control module, rotation factor address control module, vector RAM group, rotation factor ROM, and BF module;
[0053] The point-by-point multiplication module includes: a matrix address generation module, a matrix RAM group, and a point-by-point multiplication module;
[0054] The polynomial RAM group contains t polynomial RAM blocks, where the depth of each polynomial RAM block is V; in this example, when the security level is 2, t = 8, V = 128, and the data bit width is 24 bits.
[0055] The vector RAM group contains t vector RAM blocks, where the depth of each vector RAM block is V; in this example, when the security level is 2, t = 8, V = 128, and the data bit width is 24 bits.
[0056] The matrix RAM group contains t matrix RAM blocks, where the depth of each matrix RAM block is M; in this example, when the security level is 2, t = 8, M = 512, and the data bit width is 26 bits.
[0057] The rotation factor ROM contains a rotation factor data; in this example, the rotation factor ROM stores 255 rotation factors with a data bit width of 24 bits.
[0058] After the NTT_INTT selection module obtains the mode input signal and selects the NTT operation mode or the INTT operation mode through the selector, it sends the NTT enable signal / INTT enable signal (NTT_en / INTT_en) to the vector address control module, the twitch factor address control module and the BF module respectively.
[0059] The vector address control module generates the corresponding read / write address of the vector polynomial based on the received NTT enable signal / INTT enable signal (NTT_en / INTT_en) and sends it to the vector RAM group;
[0060] The rotation factor control module generates the corresponding rotation factor read address based on the received NTT enable signal / INTT enable signal (NTT_en / INTT_en) and sends it to the rotation factor ROM;
[0061] The vector RAM group reads the corresponding x vector polynomial data according to the read / write address of the vector polynomial and sends them to the dot product module and the BF module respectively; in this example, x = 8;
[0062] The rotation factor ROM reads the corresponding y rotation factor data according to the rotation factor read address and sends it to the BF module; in this example, y = 1;
[0063] The BF module selects either the NTT or INTT operation mode based on the received NTT enable signal / INTT enable signal (NTT_en / INTT_en). It then sequentially performs Montgomery reduction, bit truncation, multiplication, and addition / subtraction operations on the input x vector polynomial data and y twitch factor data, completing n rounds of butterfly operations. Finally, it stores the NTT operation result in a vector RAM group or the INTT operation result in a polynomial RAM group. In this example, x = 8, y = 1, and n = 8.
[0064] The matrix address generation module generates the read address of the matrix RAM group based on the polynomial multiplication enable signal (poly_mult_en) sent externally and sends it to the matrix RAM group;
[0065] The matrix RAM group reads the corresponding x matrix polynomial data according to the received read address and transmits them sequentially to the dot product module; in this example, x = 8;
[0066] The NTT / INTT selection module first starts the NTT operation mode and sends the data of the NTT operation mode to the vector address control module and the rotation factor address control module respectively. After the NTT operation result obtained after the NTT transformation is completed and the matrix polynomial data is reduced and multiplied, the INTT operation mode is started, and the data of the INTT operation mode is resent to the vector address control module and the rotation factor address control module.
[0067] In this example, such as Figure 2 As shown, the vector address control module includes: a multi-address counting module, a butterfly operation rounds module, a vector address generation module, a vector base address generation module, and a vector address synthesis module;
[0068] The multi-address counting module receives the NTT enable signal / INTT enable signal (NTT_en / INTT_en) and sends the enable input signal (en) to the butterfly operation wheel counting module;
[0069] After receiving the enable input signal (en), the butterfly operation round number module starts recording the current round number of the butterfly operation and sends the round number signal (round) to the multi-address counting module and the vector base address generation module respectively;
[0070] The multi-address counting module performs vector address counting on the address values of each round of butterfly operation based on the NTT enable signal / INTT enable signal (NTT_en / INTT_en) and the round number signal (round), and then obtains the address count data (addr_cnt) and sends it to the vector address generation module.
[0071] The vector address generation module calculates the initial address (poly_addr) of the vector polynomial based on the address count data (addr_cnt) received from each round of butterfly operation and sends it to the vector address synthesis module.
[0072] The vector base address generation module calculates the number of p terms in the vector polynomial based on the received round signal, thereby generating the base address (basic_addr) of the vector polynomial and sending it to the vector address synthesis module; in this example, p = 256;
[0073] The vector address synthesis module calculates the read / write address of the vector polynomial based on its initial address (poly_addr) and base address (basic_addr). The read address includes the effective read address (r_addr) and effective read signal (r_en) of the vector polynomial, and sends them to the vector RAM group and a register. The register outputs the effective write address (w_addr) and effective write signal (w_en) of the vector polynomial based on the effective read address (r_addr) and effective read signal (r_en), and sends them to the vector RAM group.
[0074] In this example, such as Figure 3 As shown, the rotation factor address control module includes: an address increment / decrement counter module, a butterfly wheel increment / decrement counter module, and a rotation factor address generation module;
[0075] If the received input signal is the NTT enable signal (NTT_en), the twitch factor address starts incrementing; if the received input signal is the INTT enable signal (INTT_en), the twitch factor address starts decrementing. This allows the module to calculate the number of polynomial coefficients in each round of butterfly operation and send the twitch factor address count data (addr_cnt) to the twitch factor address generation module.
[0076] If the received enable signal is the NTT enable signal (NTT_en), the butterfly operation round count module starts counting from 1 to 8; if the received enable signal is the INTT enable signal (INTT_en), the butterfly operation round count starts counting from 8 to 1. This determines the current round number of the butterfly operation and the number of p-term polynomials processed, and sends the round number to the twitch factor address generation module. In this example, p = 256.
[0077] The rotation factor address generation module calculates the effective read address (r_addr) and effective read signal (r_en) of the rotation factor based on the received rotation factor address count data (addr_cnt) and the number of butterfly operation rounds (round_cnt), and sends them to the rotation factor ROM.
[0078] In this example, such as Figure 4 As shown, if the mode signal received by the BF module is the NTT enable signal (NTT_en), x vector data (din0~din7) are delayed by one clock cycle through the register to obtain x vector delayed data (din0_mid~din7_mid); the x vector delayed data (din0_mid~din7_mid) and y twitch factor data (zeta) are multiplied to obtain x multiplication results; the x multiplication results are then added / subtracted pairwise with the x vector data to obtain x added data and x subtracted data, which are sent to the first selector and the second selector, respectively. The first selector and the second selector, based on the mode input signal, use two clock cycles to output the x added data and x subtracted data as the NTT operation results (NTT_dout0~NTT_dout7, NTT_dout0_mid~NTT_dout7_mid); in this example, x=8, y=1;
[0079] If the received mode signal from the BF module is the INTT enable signal (INTT_en), x vector data (din0~din7) are delayed by one clock cycle through the register to obtain x vector delayed data (din0_mid~din7_mid). The x vector data (din0~din7) and the x vector delayed data (din0_mid~din7_mid) are then added / subtracted pairwise to obtain x added data and x subtracted data respectively. The x subtracted data and y twitch factor data (zeta) after inversion are multiplied to obtain x multiplied results. The x multiplied results and x added data are sent to the first selector and the second selector respectively. The first selector and the second selector, based on the mode input signal, use two clock cycles to output the x added data and x subtracted data as the INTT operation results (INTT_dout0~INTT_dout7, INTT_dout0_mid~INTT_dout7_mid). In this example, x=8, y=1.
[0080] In practice, a hybrid single / dual-wheel approach is used for NTT calculations, and its butterfly operation method and calculation cycle diagram are shown below. Figure 5 As shown. The first six rounds are single-round calculations, processing 8 input data points. Using a pipelined structure, processing 256 data points requires 33 × 6 clock cycles. The last two rounds are dual-round calculations, also requiring 33 clock cycles to process 256 data points. Each round switch in the butterfly operation requires one clock cycle of waiting time; therefore, the NTT transformation of a 256-term polynomial requires a total of 34 × 7 = 238 clock cycles.
[0081] In its implementation, the Dilithium algorithm includes multiple security levels. Taking security level 2 as an example, its overall timing diagram is as follows: Figure 6 As shown. The vector consists of 4×1 coefficients of a 256-term polynomial, and the matrix consists of 4×4 coefficients of a 256-term polynomial. Both the NTT and INTT transformations of the 256-term polynomial require 238 clock cycles, and the matrix is multiplied row-by-row by the vector, requiring 128 clock cycles per row. Therefore, at security level 2, the total number of clock cycles required for polynomial multiplication is 238×4 + 128×4 + 238×4 = 2416.
[0082] The matrix address generation module performs calculations sequentially starting from 0 based on the received polynomial multiplication enable signal, and sends the valid read signal and valid read address of the matrix to the matrix RAM group. It reads x matrix polynomial data every clock cycle. When the last group of p matrix polynomial coefficients is read, the matrix address generation module pulls the valid read signal low and sends it to the matrix RAM group, stopping the reading of data from the matrix RAM group. In this example, x = 8, p = 256.
[0083] In this example, such as Figure 7 As shown, the dot product module includes: the Montgomery reduction module and the reduction module;
[0084] The Montgomery reduction module multiplies the received vector polynomial data (vector_din) and matrix polynomial data (matrix_din). The multiplication result is then shifted left by 13 bits and left by 26 bits respectively, resulting in two shifted results. The two shifted results are added to the multiplication result to obtain the first addition result. The result of left-shifting the multiplication result by 23 bits is subtracted from the first addition result to obtain the first subtraction result. The first subtraction result is then left-shifted by 23 bits to obtain the first shifted result, which is then added to the first subtraction result to obtain the second addition result. The first subtraction result is then left-shifted by 13 bits and subtracted from the second addition result and the multiplication result to obtain the output result m (mont_dout), which is then sent to the reduction module.
[0085] After receiving the output result m(mont_dout) and performing a left shift operation of 22 bits, the reduction module adds the shifted result to the output result to obtain the third addition result. The third addition result is then left shifted by 23 bits to obtain the third shift result. This third shift result is then left shifted by 23 bits and added to the third shift result to obtain the fourth addition result. Finally, the third shift result is left shifted by 13 bits and subtracted from the fourth addition result and the output result m(mont_dout) to obtain the polynomial multiplication output data mult(poly_mult_dout).
Claims
1. A polynomial multiplier circuit based on the Dilithium algorithm, characterized in that, include: NTT / INTT module, point-by-point multiplication module, polynomial RAM group; The NTT / INTT module includes: an NTT_INTT selection module, a vector address control module, a rotation factor address control module, a vector RAM group, a rotation factor ROM, and a BF module; The point-by-point multiplication module includes: a matrix address generation module, a matrix RAM group, and a point-by-point multiplication module; The polynomial RAM group contains t polynomial RAM blocks, wherein the depth of each polynomial RAM block is V; The vector RAM group contains t vector RAM blocks, where the depth of each vector RAM block is V; The matrix RAM group contains t matrix RAM blocks, where the depth of each matrix RAM block is M; The rotation factor ROM contains a rotation factor data; The NTT_INTT selection module acquires the mode input signal and selects the NTT operation mode or the INTT operation mode through the selector, and then sends the NTT enable signal / INTT enable signal to the vector address control module, the rotation factor address control module and the BF module respectively. The vector address control module generates the corresponding read / write address of the vector polynomial based on the received NTT enable signal / INTT enable signal and sends it to the vector RAM group; The rotation factor control module generates a corresponding rotation factor read address based on the received NTT enable signal / INTT enable signal and sends it to the rotation factor ROM; The vector RAM group reads the corresponding x vector polynomial data according to the read / write address of the vector polynomial and sends them to the dot product module and the BF module respectively. The rotation factor ROM reads the corresponding y rotation factor data according to the rotation factor read address and then sends it to the BF module; The BF module selects the NTT operation mode or the INTT operation mode according to the received NTT enable signal / INTT enable signal, and performs Montgomery reduction, bit truncation, multiplication and addition / subtraction operations on the input x vector polynomial data and y twitch factor data in an orderly manner. After completing n rounds of butterfly operation, the NTT operation result is stored in the vector RAM group or the INTT operation result is stored in the polynomial RAM group. The matrix address generation module generates the read address of the matrix RAM group based on the polynomial multiplication enable signal sent externally and sends it to the matrix RAM group; The matrix RAM group reads the corresponding x matrix polynomial data according to the received read address and transmits them sequentially to the dot product module; The dot product module performs multiplication and reduction operations on x matrix polynomial data and x vector polynomial data in an orderly manner, thereby storing the final calculation result in the polynomial RAM group.
2. The polynomial multiplier circuit based on the Dilithium algorithm according to claim 1, characterized in that, The NTT / INTT selection module first starts the NTT operation mode and sends the data of the NTT operation mode to the vector address control module and the rotation factor address control module respectively. After the NTT operation result obtained after the NTT transformation is completed and the matrix polynomial data is reduced and multiplied, the INTT operation mode is started, and the data of the INTT operation mode is resent to the vector address control module and the rotation factor address control module.
3. The polynomial multiplier circuit based on the Dilithium algorithm according to claim 1, characterized in that, The vector address control module includes: a multi-address counting module, a butterfly operation round counting module, a vector address generation module, a vector base address generation module, and a vector address synthesis module; The multi-address counting module receives the NTT enable signal / INTT enable signal and sends an enable input signal to the butterfly operation wheel counting module; After receiving the enable input signal, the butterfly operation round number module starts recording the current round number of the butterfly operation and sends the round number signal to the multi-address counting module and the vector base address generation module respectively; The multi-address counting module performs vector address counting on the address values of each round of butterfly operation based on the NTT enable signal / INTT enable signal and the round number signal, and then sends the address count data to the vector address generation module. The vector address generation module calculates the initial address of the vector polynomial based on the address count data received for each round of butterfly operation and sends it to the vector address synthesis module. The vector base address generation module calculates the number of p-term polynomials in the vector polynomial based on the received round number signal, thereby generating the base address of the vector polynomial and sending it to the vector address synthesis module; The vector address synthesis module calculates the read / write address of the vector polynomial based on the initial address and the base address of the vector polynomial. The read address includes the effective read address and the effective read signal of the vector polynomial, and sends them to the vector RAM group and a register. The register outputs the effective write address and the effective write signal of the vector polynomial based on the effective read address and the effective read signal, and sends them to the vector RAM group.
4. The polynomial multiplier circuit based on the Dilithium algorithm according to claim 1, characterized in that, The rotation factor address control module includes: an address increment / decrement counter module, a butterfly wheel increment / decrement counter module, and a rotation factor address generation module; If the input signal received by the address increment / decrement counting module is an NTT enable signal, the twitch factor address starts to increment; if the input signal received is an INTT enable signal, the twitch factor address starts to decrement; thereby calculating the number of polynomial coefficients of the p terms in each round of butterfly operation and sending the twitch factor address count data to the twitch factor address generation module. If the received enable signal is an NTT enable signal, the butterfly operation round count module starts counting from 1 to 8; if the received enable signal is an INTT enable signal, the butterfly operation round count starts counting from 8 to 1. This determines the current number of butterfly operation rounds and the number of p-term polynomials processed, and sends the butterfly operation round count to the twitch factor address generation module. The rotation factor address generation module calculates the effective read address and effective read signal of the rotation factor based on the received rotation factor address count data and the number of butterfly operation rounds, and sends them to the rotation factor ROM.
5. The polynomial multiplier circuit based on the Dilithium algorithm according to claim 1, characterized in that, If the received mode signal is an NTT enable signal, the BF module delays x vector data through a register for one clock cycle to obtain x delayed vector data. The x delayed vector data and y twitch factor data are multiplied to obtain x multiplication results. These x multiplication results are then added / subtracted pairwise with the x vector data to obtain x added data and x subtracted data, which are sent to the first selector and the second selector, respectively. The first and second selectors, based on the mode input signal, output the x added data and x subtracted data as the NTT operation results over two clock cycles. If the received mode signal is an INTT enable signal, the BF module delays x vector data through a register for one clock cycle to obtain x delayed vector data. Then, it performs pairwise addition / subtraction operations on the x vector data and x delayed vector data to obtain x added data and x subtracted data respectively. The x subtracted data are multiplied by y inverted twitch factor data to obtain x multiplied results. These x multiplied results are sent to the first selector, and the x added data are sent to the second selector. The first and second selectors, based on the mode input signal, output the x added data and x subtracted data as the INTT operation results over two clock cycles.
6. The polynomial multiplier circuit based on the Dilithium algorithm according to claim 1, characterized in that, The matrix address generation module performs calculations sequentially starting from 0 based on the received polynomial multiplication enable signal, and sends the valid read signal and valid read address of the matrix to the matrix RAM group. It reads x matrix polynomial data in each clock cycle. When the coefficients of the last p-term matrix polynomial are read out, the matrix address generation module pulls the valid read signal low and sends it to the matrix RAM group, stopping the reading of data from the matrix RAM group.
7. The polynomial multiplier circuit based on the Dilithium algorithm according to claim 1, characterized in that, The dot product module includes: a Montgomery reduction module and a reduction module; The Montgomery reduction module multiplies the received vector polynomial data and matrix polynomial data. The multiplication result is then shifted left by 13 bits and left by 26 bits respectively to obtain two shifted results. The two shifted results are added to the multiplication result to obtain the first addition result. The result of left-shifting the multiplication result by 23 bits is subtracted from the first addition result to obtain the first subtraction result. The first subtraction result is left-shifted by 23 bits to obtain the first shifted result, and then added to the first subtraction result to obtain the second addition result. The first subtraction result is left-shifted by 13 bits and then subtracted from the second addition result and the multiplication result to obtain the output result m, which is then sent to the reduction module. After receiving the output result m and performing a left shift operation of 22 bits, the reduction module adds the shifted result to the output result to obtain a third addition result. The third addition result is then left-shifted by 23 bits to obtain a third shift result. This third shift result is then left-shifted by 23 bits and added to the third shift result to obtain a fourth addition result. The third shift result is then left-shifted by 13 bits and subtracted from the fourth addition result and the output result m to obtain the polynomial multiplication output data mult.