Sparse polynomial multiplication accelerator applied to HQC algorithm

By designing the sparse polynomial multiplication accelerator of the HQC algorithm, the multiplication process is optimized by using the sparse polynomial characteristics, the security problems and low computing efficiency of traditional algorithms in the quantum computing environment are solved, and low resource consumption and efficient computing are achieved.

CN120353431APending Publication Date: 2025-07-22ZHEJIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510399201.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing public key cryptographic algorithms such as RSA and ECC are no longer safe in the face of quantum computer attacks, and traditional polynomial multiplication algorithms are inefficient in sparse polynomials. The existing accelerators fail to effectively utilize sparse characteristics, resulting in insufficient resource overhead and computational efficiency.

Method used

A sparse polynomial multiplication accelerator applied to the HQC algorithm is designed, using sparse polynomial non-zero coefficient index register group, status register group, multiplication unit, control state machine and address generation module. By optimizing the multiplication process and hardware architecture, resource consumption is reduced and computing efficiency is improved.

Benefits of technology

It realizes the sparse polynomial multiplication operation in bit-sized RAM storage space, reduces the number of multiplications and hardware resource consumption, has low power consumption and high computing efficiency, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353431A_ABST
    Figure CN120353431A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse polynomial multiplication accelerator circuit applied to an HQC algorithm, and belongs to the field of post quantum cryptography algorithm hardware acceleration. Comprising a sparse polynomial non-zero coefficient index register set, a state register set, a multiplication and addition unit, a control state machine and an address generation module. The sparse polynomial non-zero coefficient index register set is used for storing indexes of sparse polynomial non-zero coefficients and providing the indexes for the control state machine, and the control state machine further inputs the obtained indexes to the address generation module; the state register group is used for configuring and representing the working state of the accelerator, storing information of registers and providing information for the address generation module and the control state machine; the address generation module is used for calculating address information and feeding back a result to the control state machine; and the control state machine reads the coefficient of the dense polynomial from the external RAM and inputs the coefficient into the multiplication and addition unit for calculation, and after the multiplication and addition unit completes calculation, the control state machine writes a calculation result into the external RAM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hardware acceleration of post - quantum cryptographic algorithms, and particularly relates to a sparse polynomial multiplication accelerator applied to the HQC algorithm. Background Art

[0002] The security of existing public - key cryptographic algorithms, such as RSA and ECC, depends on mathematical difficult problems such as integer factorization and elliptic - curve discrete logarithm. However, with the rapid development of quantum computers and the maturity of quantum - computing theories and methods, the mathematical problems relied on by these public - key cryptographic algorithms will be cracked by sufficiently mature quantum computers in polynomial time, and the existing public - key cryptosystems will no longer be secure. Therefore, it is urgent to develop new public - key cryptographic algorithms that can resist quantum - computer attacks to ensure information security in the post - quantum era. The HQC algorithm is a post - quantum cryptographic algorithm based on coding. The underlying mathematical difficult problems it relies on have deep research accumulations and quantum hardness recognized by the academic mainstream. Therefore, the security of this algorithm is sufficiently guaranteed. At the same time, it also has good performance in performance tests compared with other algorithms and is expected to be standardized in the fourth - round screening of NIST.

[0003] Polynomial multiplication is the most time - consuming operation in the HQC algorithm. However, traditional polynomial - multiplication algorithms such as the Schoolbook algorithm and the Karatsuba algorithm cannot achieve ideal computational efficiency on sparse polynomials, and existing polynomial - multiplication accelerator designs suitable for low - power devices rarely focus on utilizing the sparse characteristics of polynomials. Therefore, in the research on sparse polynomial multiplication accelerators for the HQC algorithm, how to balance low resource overhead and high computational efficiency has become an urgent problem for researchers in this field. Summary of the Invention

[0004] To solve the problems in the prior art, the present invention provides a sparse polynomial multiplication accelerator applied to the HQC algorithm.

[0005] The technical solution adopted by the present invention is as follows:

[0006] In a first aspect, the present invention discloses a sparse polynomial multiplication accelerator applied to the HQC algorithm, including a sparse polynomial non-zero coefficient index register bank, a status register bank, a multiply-accumulate unit, a control state machine, and an address generation module; the sparse polynomial non-zero coefficient index register bank is used to store the indexes of the non-zero coefficients of the externally input sparse polynomial and provide the indexes to the control state machine, and the control state machine will further input the obtained indexes to the address generation module; the status register bank is used to configure and represent the working state of the accelerator according to the externally or control state machine's status input, and store the Hamming weight of the externally input sparse polynomial, the starting address of the dense polynomial in the external RAM, the length of the dense polynomial, and the starting address of the result polynomial in the external RAM, and provide the information stored in the status register bank to the address generation module and the control state machine; the address generation module is used to calculate address information based on the indexes input by the control state machine and the input of the status register bank and feedback the calculated address information to the control state machine; the control state machine reads the coefficients of the dense polynomial from the external RAM based on the address information and the starting address of the dense polynomial in the external RAM and inputs them into the multiply-accumulate unit for calculation. After the multiply-accumulate unit completes the multiplication calculation, the control state machine writes the calculation result of the multiply-accumulate unit to the external RAM based on the starting address of the result polynomial in the external RAM.

[0007] In a second aspect, the present invention discloses a sparse polynomial multiplication calculation method applied to the HQC algorithm for the accelerator, including:

[0008] 1) The control state machine initializes the calculation group number g and sets the calculation group number g to 0;

[0009] 2) The sparse polynomial non-zero coefficient index register bank receives and stores the indexes of the non-zero coefficients of the externally input sparse polynomial; the Hamming weight register receives the Hamming weight of the externally input sparse polynomial, the dense polynomial starting address register receives the starting address of the externally input dense polynomial in the RAM; the polynomial length register receives the length of the externally input polynomial; the result polynomial starting address register receives the starting address of the externally input result polynomial in the RAM; the working state register receives the externally written 0x3 to enable the control state machine to start polynomial multiplication calculation;

[0010] 3) The control state machine requests and reads the dense polynomial coefficient u0 and the dense polynomial coefficient u n-1 from the external RAM where they are located. The size of one memory row is Nbit. Concatenate the N-(nmod N)bit from u0 to the dense polynomial coefficient u (N-(n mod N)-1) to the back of u n-1 to obtain and write it to u n-1The storage row where it is located;

[0011] 4) Clear the values of the counter i of the 0th path and the counter j of the 1st path of the multiply-accumulate unit, and at the same time clear the values of the register of the 0th path and the register of the 1st path of the multiply-accumulate unit; the control state machine then reads the index P of the corresponding non-zero coefficient of the sparse polynomial from the sparse polynomial non-zero coefficient index register group based on i and j i and P j and index P i and P j and the calculated group number g are output to the address generation module, and the address offset D i 、address offset DN i 、address offset D j and address offset DN j are received; then the control state machine requests the storage row where it is located;

[0012] 5) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg00 of the multiply-accumulate unit, and requests the storage row where it is located;

[0013] 6) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg01 of the multiply-accumulate unit, the control state machine controls the multiply-accumulate unit to perform the 0th path calculation and updates the value of the 0th path register way0; the control state machine then requests the storage row where it is located, and increments the value of the counter i of the 0th path by 1;

[0014] 7) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg01 of the multiply-accumulate unit, the control state machine then requests the storage row where it is located, if i at this time is equal to ω, perform step 9), otherwise continue to perform step 8);

[0015] 8) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg11 of the multiply-accumulate unit, the control state machine controls the multiply-accumulate unit to perform the 1st path calculation and updates the value of the 1st path register way1; the control state machine then requests the storage row where it is located, and increments the value of the counter j of the 1st path by 1, and then perform step 5);

[0016] 9) The control state machine reads The entire row of data in the storage row where it is located is written into the register reg11 of the multiply-accumulate unit, and the control state machine controls the multiply-accumulate unit to perform the first-way calculation and update the value of the first-way register way1; then the value of the zero-way register way0 is written out to the external RAM at w 2gN on the storage row where it is located; if at this time and then return to step 1), otherwise proceed to step 10);

[0017] 10) The control state machine writes the value of the first-way register way1 to the external RAM at w (2g+1)N on the storage row where it is located, and increments the value of the calculation group number g by 1; if then it indicates that the multiplication calculation of the entire sparse polynomial is completed, and the control state machine writes 0x1 to the working state register; otherwise return to step 4) to continue the multiplication calculation of the sparse polynomial.

[0018] Thirdly, the present invention also discloses an SoC integrating the sparse polynomial multiplication accelerator applied to the HQC algorithm.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0020] (1) The present invention only needs RAM storage space of bit size to complete the sparse polynomial multiplication operation in the HQC algorithm. Among them, the first bits store the dense polynomial, the latter bits store the multiplication result polynomial, and the sparse polynomial is converted into the form of its non-zero coefficient indexes and stored in the sparse polynomial non-zero coefficient index register group. If the polynomial is specified to be stored in an N-bit alignment during compilation, then the accelerator can directly read and write data on the RAM space specified by the compiler without the need to allocate additional RAM storage space.

[0021] (2) The present invention adopts a configurable hardware architecture design, makes full use of the existing hardware resources during the operation process, reduces the resource consumption while efficiently completing various operations required during the sparse polynomial multiplication operation, thereby improving the calculation efficiency.

[0022] (3) The present invention has strong scalability and a wide application range, and has advantages such as low resource consumption and low power consumption. The sparse polynomial multiplication operation of all parameters in the HQC algorithm can be realized by configuring the register group, and the configuration process is simple, meeting the requirements for flexibility in many application scenarios. Description of the Drawings

[0023] Figure 1 It is a schematic diagram of a simplified representation method of a sparse polynomial;

[0024] Figure 2 Schematic diagram of the row and column index rules for the elements of the circulant matrix;

[0025] Figure 3 Schematic diagram of the simplification process of matrix - form multiplication;

[0026] Figure 4 Overall hardware architecture diagram of the sparse polynomial multiplication accelerator;

[0027] Figure 5 Schematic diagram of the storage methods of the dense polynomial and the result polynomial in the RAM;

[0028] Figure 6 Circuit structure diagram of the address generation module;

[0029] Figure 7 Schematic diagram of the extraction methods of the 0 - th extraction module and the 1 - st extraction module;

[0030] Figure 8 Flowchart of the control state machine algorithm;

[0031] Figure 9 For splicing u0 and u n-1 Schematic diagram of the method for the data in the storage row where it is located;

[0032] Figure 10 Schematic diagram of the implementation method of sparse polynomial multiplication on HQC128 in a specific embodiment;

[0033] Figure 11 Partial simulation waveform diagram of sparse polynomial multiplication on HQC128 in a specific embodiment;

[0034] Figure 12 An SoC integration scheme diagram of the sparse polynomial multiplication accelerator. Detailed implementation manners

[0035] The present invention will be further described and explained below in conjunction with the detailed implementation manners. The described embodiments are only demonstrations of the present disclosure content and do not delimit the scope of limitation. Without conflict, the technical features of each implementation manner in the present invention can be combined accordingly.

[0036] In order to solve the problems in the prior art and fill the technical gap in this field, the present invention proposes a sparse polynomial multiplication accelerator circuit applied to the HQC algorithm, including a sparse polynomial non-zero coefficient index register group, a status register group, a multiplication-addition unit, a control state machine and an address generation module. The multiplication-addition unit implements two-way multiplication-addition operations, and reduces the hardware circuit overhead by module multiplexing; the control state machine implements an efficient sparse polynomial multiplication algorithm, reducing the number of multiplications required for the operation; the address generation module implements parallel calculation of four addresses; and the register group can also be configured through the AXI bus interface, so that the accelerator can support different parameters of the HQC algorithm. The present invention aims to speed up the calculation speed of sparse polynomial multiplication in the HQC algorithm with lower power consumption and hardware resource overhead while taking into account the computational efficiency. In addition, the present invention can also be applied to the calculation of other post-quantum cryptographic algorithms in the finite field GF(2 n ), such as the sparse polynomial multiplication of the BIKE algorithm.

[0037] The purpose of the present invention is to provide a sparse polynomial multiplication accelerator hardware architecture for the HQC algorithm, by taking the sparse polynomial non-zero coefficient index P in the circulant matrix rot(U) as q The corresponding column vector By XORing the corresponding elements in the accelerator, the time complexity of multiplication is reduced to O(ωn), which effectively reduces the number of multiplications in the operation process, thereby improving the calculation speed of sparse polynomial multiplication. A configurable design is also made for the hardware architecture of the accelerator circuit, reducing the hardware resource consumption of the accelerator circuit.

[0038] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0039] The multiplication in the HQC algorithm is defined as the finite field GF(2 n )(It can also be expressed as ), the coefficient operation follows the finite field GF(2) (which can also be expressed as ), that is, addition is equivalent to XOR operation, and multiplication is equivalent to logical AND operation. express A vector space with dimension n on , where n is a positive integer related to the security strength of the algorithm and its value is equal to the length of the dense polynomial, The elements in can be considered as a ring A polynomial or row vector on x n -1 is The irreducible polynomials in .

[0040] For polynomial U, , define the ring polynomial multiplication on , where U is a dense polynomial, whose vector form is V is a sparse polynomial, whose vector form is The Hamming weight of V is ω, that is, V has only ω non-zero coefficients; W is the multiplication result, whose vector form is: where w k = ∑ e+f≡k mod n u e v f , k ∈ {0, 1, …, n - 1}. The polynomial multiplication can be expressed in the form of a polynomial group as:

[0041]

[0042] Define the circulant matrix Then the polynomial multiplication can be further expressed as:

[0043]

[0044] Existing implementations of polynomial multiplication mostly use divide-and-conquer methods such as the Schoolbook algorithm or the Karatsuba algorithm. For example, the code submitted by the HQC algorithm team to the fourth-round post-quantum cryptography algorithm solicitation of NIST uses the Karatsuba algorithm to implement polynomial multiplication operations. However, the time complexity of the Schoolbook algorithm multiplication is O(n 2 ), and the time complexity of the Karatsuba algorithm multiplication with base-2 decomposition is And neither of them can utilize the sparse characteristics of the polynomials in the HQC algorithm and cannot achieve the optimal efficiency. Therefore, it is necessary to further observe the characteristics of polynomial multiplication in HQC and design more efficient algorithms and hardware architectures.

[0045] v is defined as a sparse polynomial containing ω non-zero coefficients. The following gives the relationship between the values and indices of the coefficients of polynomial V: Let P = {P0, P1, …, P ω-1} ∈ {0, 1, …, n - 1} be the set of indices of non-zero coefficients, then Figure 1 Illustrates this relationship more vividly. The following transforms the polynomial multiplication in matrix form shown in formula 2: First Figure 2 gives the elements on each row and column of the circulant matrix rot(U). It is not difficult to observe that the element in the a-th row and b-th column of the circulant matrix rot(U) is u (n-b+a)mod n, that is, the index of the element in the \(a\)-th row and \(b\)-th column is \((n - b + a)\bmod n\); the process of simplifying the polynomial multiplication in the matrix form shown in Formula 2 is as Figure 3 shown, Figure 3 marks some non-zero coefficients of \(V\) of the sparse polynomial, such as Since a number multiplied by 0 gives 0 and multiplied by 1 gives itself, that is so the terms with a result of 0 can be ignored in the addition, and only specific columns of the circulant matrix need to be selected for addition, such as Figure 3 the \(P0\) column, \(P1\) column, \(P\) ω-1 column in, and finally the simplified matrix form is obtained. At this time, Formula 1 becomes:

[0046]

[0047] It is not difficult to find that only \(\omega\cdot n\) multiplications are needed to complete the polynomial multiplication operation. Let the column vector be the column with column index \(P\) b of the circulant matrix \(rot(U)\), and the multiplication result can be further obtained. In order to make full use of the sparse characteristics of the polynomial and improve the algorithm efficiency, the hardware architecture of the present invention is implemented based on the foregoing proposed calculation method.

[0048] Such as Figure 4 shown, the present invention is a sparse polynomial multiplication accelerator applied to the HQC algorithm, including the following sub-modules: AXI bus interface, sparse polynomial non-zero coefficient index register group, status register group, multiply-accumulate unit, address generation module, and control state machine. These modules are introduced separately below.

[0049] The AXI bus interface enables the accelerator to act as an AXI bus slave, receive the register configuration signal sent by the AXI bus master, and the AXI bus master reads the working state of the accelerator.

[0050] The sparse polynomial non-zero coefficient index register group includes 149 16-bit registers. Because among all the parameters of the HQC algorithm, the maximum value of \(\omega\) is 149, the maximum value of \(n\) is 57637, and \(2\) 15 <57637<2 16 , so by setting the number of registers to 149 and the register bit width to 16 bits, all parameters can be supported. These registers all have a unique address offset, ranging from 0x000 to 0x128, so that the AXI bus master can write the sparse polynomial non-zero coefficient index into the register. Among them, the set \(P\) of the indexes of the non-zero coefficients in the sparse polynomial non-zero coefficient index register group is \(\{P0, P1, \ldots, P\) ω-1} ∈ {0, 1, …, n - 1}, where ω is the number of non - zero coefficients in the sparse polynomial; P0 is the index of the first non - zero coefficient, P1 is the index of the second non - zero coefficient; n is the length of the dense polynomial.

[0051] The status register set includes the working status register status, Hamming weight ω register omega, starting address register addrU of the dense polynomial U, polynomial length n register nlen, and starting address register addrW of the result polynomial W.

[0052] Working status register status: The address offset is 0x12C, which is used to configure and represent the working status of the accelerator. The AXI host writes 0x3 to make the control state machine start calculating the polynomial multiplication. After the calculation is completed, the control state machine writes 0x1 to indicate that the multiplication calculation is completed. Hamming weight ω register omega: The address offset is 0x130, which is used to store the value of the Hamming weight ω of the sparse polynomial V. Starting address register addrU of the dense polynomial U: The address offset is 0x134, which is used to store the starting address of this polynomial in the external RAM. Polynomial length n register nlen: The address offset is 0x138, which is used to store the length of the polynomial. Starting address register addrW of the result polynomial W: The address offset is 0x13C, which is used to store the starting address of this polynomial in the external RAM.

[0053] The following introduces the storage methods of the dense polynomial U and the result polynomial W in the RAM. As Figure 5 illustrates the storage addresses of each coefficient of U and W in the RAM. Here, one RAM storage behavior is N bits. For example, the address of u0 is addrU, the address of u N is addrU + N, the address of w0 is addrW, and so on; Figure 5 also illustrates the start - end range of the addresses of all bits within each storage row. For example, u0 to u N-1 In the same storage row, the address range is addrU to addrU + N - 1. Due to the characteristics of the RAM, each read operation can only read N bits of the same storage row, and each write operation can only write N bits into the same storage row.

[0054] The address generation module is used to generate four address information according to the current calculation group number g, polynomial length n, and the indexes P i and P j , that is, the address generation module generates an address offset D i calculated based on the index P i , address offset DN i , the address offset D j calculated by the address generation module based on the index P j and address offset DNj . Among them, D i represents a coefficient the address offset relative to u0 in RAM, D j represents a coefficient the address offset relative to u0 in RAM. Therefore, there is the following relationship: and DN i is calculated from D i , DN j is calculated from D j . They are all the address offsets of specific coefficients in the dense polynomial U relative to u0 in RAM. The generation process of the address offset DN i is described by the following pseudocode. The symbol represents floor function:

[0055]

[0056] The generation process of the address offset DN j is described by the following pseudocode:

[0057]

[0058] The circuit diagram of the address generation module is as shown in Figure 6 . Since this module is completely implemented by combinational logic circuits, the above four address information can be output simultaneously. The working principle of this module is introduced below. The method to obtain the address offset D i is as follows: After the signal g is input to this module, it is sent to the shifter ① and shifted left by log2 N + 1 bits, that is, the calculation of 2gN is completed, and then it is input to the subtractor ② to subtract from the index P i of the non-zero coefficient of the sparse polynomial to get 2gN - P i , and then it is input to the "1" input port of the first multiplexer ③; the signal n and the index P i of the non-zero coefficient of the sparse polynomial are subtracted in the subtractor ④ to get n - P i , and then through the adder ⑤, n - P i + 2gN is calculated, and then it is input to the "0" input port of the first multiplexer ③. Also, D i =(n - P i + 2gN) mod n, so it is necessary to perform the modulo n operation on n - P i + 2gN: n - P i + 2gN and n are input to the first comparator ⑥. If n - P i + 2gN ≥ n, the comparison result is 1, the input of the signal selected by the first multiplexer ③ is 1, and the signal output from the "1" input port of the first multiplexer ③ is D i = 2gN - P i; Otherwise, the comparison result is 0, and the first multiplexer ③ outputs the signal of the "0" input port, i.e., D i = n - P i + 2gN.

[0059] The method for obtaining the address offset DN i is as follows: After obtaining the signal D i , it is sent to the adder ⑧ and added to N to obtain D i + N and input to the 0 input port of the second multiplexer ⑩. At the same time, D i + N is input to the subtractor ⑨ and subtracted from the signal n to obtain D i + N - n, and then input to the 1 input port of the second multiplexer ⑩; The role of {n[:log2N], log2 N′b0} is to calculate , that is, to calculate n divided by N and take the floor, and then multiply the result by N to obtain . D i and are input to the second comparator ⑦. If the comparison result is 1, the input of the signal selected by the second multiplexer ⑩ is 1, and the second multiplexer ⑩ outputs the signal of the "1" input port, i.e., DN i = D i + N - n; Otherwise, the comparison result is 0, and the second multiplexer ⑩ outputs the signal of the "0" input port, i.e., DN i = D i + N.

[0060] The method for obtaining the address offset D j is as follows: First, 2gN output by the shifter ① and N are input to the adder to obtain (2g + 1)N. Then, the result (2g + 1)N and the index P of the non-zero coefficient of the sparse polynomial j are subtracted in the subtractor to obtain (2g + 1)N - P j , and output to the 1 input port of the third multiplexer ; Then, the length n of the dense polynomial and the index P of the non-zero coefficient of the sparse polynomial j are subtracted in the subtractor to obtain n - P j . Then, n - P j and (2g + 1)n are added in the adder to obtain n - P j + (2g + 1)N, and output to the 0 input port of the third multiplexer ; n - P j + (2g + 1)N and n are input to the third comparator If n - P jIf \((2g + 1)N\geq n\), then the third comparator outputs 1 to the selection signal port of the third multiplexer The third multiplexer finally outputs the signal of the 1 input port, that is, D j \(=(2g + 1)N - P\) j ; Otherwise, the third comparator outputs 0 to the selection signal port of the third multiplexer The third multiplexer finally outputs the signal of the 0 input port, that is, D j \(=n - P\) j \(+(2g + 1)N\);

[0061] The method for obtaining the address offset DN j is: After obtaining the address offset D j , add the address offset D j and N in the adder to get D j \(+N\) and input it to the 0 input port of the fourth multiplexer . At the same time, subtract the length n of the dense polynomial from D j \(+N\) in the subtractor to get D j \(+N - n\), and then input D j \(+N - n\) to the 1 input port of the fourth multiplexer ②0; Input D j and to the fourth comparator . If , then the fourth comparator outputs 1 to the selection signal port of the fourth multiplexer . The fourth multiplexer finally outputs the signal of the 1 input port, that is, DN j \(=D\) j \(+N - n\); Otherwise, the fourth comparator outputs 0 to the selection signal port of the fourth multiplexer . The fourth multiplexer finally outputs the signal of the 0 input port, that is, DN j \(=D\) j \(+N\).

[0062] The block diagram of the multiply-accumulate unit module is as shown in Figure 4 . The register reg00 and the result register way0 (i.e., the register way0 in Figure 4 ) belong to the storage units of the 0th path. The register reg11 and the result register way1 (i.e., Figure 4The register in way1 belongs to the storage unit of the first way. The register reg01 is shared by both ways. The processing of the data of both ways is controlled by the control state machine, realizing time-sharing processing, avoiding data conflicts, and reusing the exclusive OR module, reducing resource occupancy.

[0063] The processing steps of the 0th way are as follows: The registers reg00 and reg01 input their internal data into the 0th way extraction module respectively, and then splice to obtain 2N-bit data, denoted as S[0:2N - 1] (the data of register reg00 is in the lower N bits, i.e., S[0:N - 1], and the data of register reg01 is in the higher N bits, i.e., S[N:2N - 1]). Let k0 = D i mod N, extract N-bit data starting from k0, i.e., S[k0:k0 + N - 1], and then perform an exclusive OR operation on S[k0:k0 + N - 1] and the data in the current register way0, and output the result to register way0 and update the value of way0.

[0064] The processing steps of the 1st way are as follows: The registers reg01 and reg11 input their internal data into the 1st way extraction module respectively, and then splice to obtain 2N-bit data, denoted as T[0:2N - 1] (the data of register reg01 is in the lower N bits, i.e., T[0:N - 1], and the data of register reg11 is in the higher N bits, i.e., T[N:2N - 1]). Let k1 = D j mod N, extract N-bit data starting from k1, i.e., T[k1:k1 + N - 1], and then perform an exclusive OR operation on T[k1:k1 + N - 1] and the data in the current register way1, and output the result to register way1 and update the value of way1. Among them, the splicing and extraction methods of the 0th way and the 1st way are as Figure 7 shown.

[0065] Figure 8 The following shows the algorithm flow chart of the control state machine module. Next, combined with Figure 8 each state and its transition process will be described:

[0066] (a) RESET state: After the system resets, it first enters this state, and then enters the IDLE state.

[0067] (b) IDLE state: Clear the calculation group number g; if the AXI bus master writes 0x3 to the status register, enter the PR state, otherwise keep this state.

[0068] (c) PR state: Request the data of the storage row where u0 is located at the RAM address addrU, and then read the storage row where u0 is located; request the data of the storage row where u n-1 is located at the RAM address addrU + n - 1, and then read the storage row where un-1 Then put the storage row from u0 to u (N-(n mod N)-1) N-(n mod N) bits are concatenated to u n-1 Later, we get And write to RAM address addrU+n-1u n-1 The storage row where it is located. After writing is completed, it enters the B0 state; where u0 is the coefficient of the dense polynomial U with index 0; u n-1 is the coefficient of the dense polynomial U with index n-1.

[0069] Specifically: Figure 9 , assuming N = 128, n = 17669, then n mod N = 5, request and read u0 and u respectively 17668 The storage row where u0 is located is 122 123bit is concatenated to u 17668 Later, we get {u 17664 ~u 17668 ,u0~u 122}, then write and update the original u 17668 The storage row.

[0070] (d) B0 state: clear the value i of the 0th way counter and the value j of the 1st way counter of the multiplication and addition unit, and clear the value of the register way0 and the value of the register way1 of the multiplication and addition unit. The control state machine then reads the corresponding sparse polynomial non-zero coefficient index P from the sparse polynomial non-zero coefficient index register group based on i and j. i and P j , and index P i and P j And calculate the group number g and output it to the address generation module, and receive the address offset D output by the address generation module i , Address offset DN i , Address offset D j and address offset DN j ; Then the control state machine requests the external RAM to have the RAM address addrU+D i of The storage row where Indicates that the index in the dense polynomial U under the current calculation group number g is D i Then it enters the R0 state.

[0071] (e) R0 state: the control state machine reads the RAM address as addrU+D i of The entire row of data in the storage row is written into the register reg00 of the multiplication and addition unit; then the RAM address addrU+DN is requested from the external RAM i of The storage row where it is located. Then it enters the R1 state.

[0072] (f) R1 state: The control state machine reads the entire row of data of the storage row where the RAM address is addrU + DN i of and writes it into the register reg01 of the multiply-accumulate unit. The control state machine controls the multiply-accumulate unit to perform the 0th path calculation and updates the value of the 0th path register way0 with the calculation result; the control state machine then requests the external RAM for the storage row where the RAM address is addrU + D j of where it is located ( indicating the coefficient of index D in the dense polynomial U under the current calculation group number g j ). And increments the value of the 0th path counter i by 1. Then it enters the R2 state.

[0073] (g) R2 state: The control state machine reads the entire row of data of the storage row where the RAM address is addrU + D j of and writes it into the register reg01 of the multiply-accumulate unit; the control state machine then requests the external RAM for the storage row where the RAM address is addrU + DN j of where it is located. If the value of the 0th path counter i is equal to the Hamming weight ω at this time, it means that the 0th path has completed ω additions. Next, it is necessary to write the final calculation result of the 0th path into the RAM and enter the P0 state; otherwise, it enters the R3 state.

[0074] (h) R3 state: The control state machine reads the entire row of data of the storage row where the RAM address is addrU + DN j of and writes it into the register reg11 of the multiply-accumulate unit. The control state machine controls the multiply-accumulate unit to perform the 1st path calculation and updates the value of the 1st path register way1 with the calculation result; the control state machine then requests the external RAM for the storage row where the RAM address is addrU + D i of where it is located. And increments the value of the 1st path counter j by 1. Then it enters the R0 state.

[0075] (i) P0 state: The control state machine reads the entire row of data of the storage row where the RAM address is addrU + DN j of and writes it into the register reg11 of the multiply-accumulate unit. The control state machine controls the multiply-accumulate unit to perform the 1st path and updates the value of the 1st path register way1 with the calculation result. Then the control state machine writes the value of the 0th path register way0 to the w at the RAM address addrW + 2gN 2gNOn the storage row where it is located. If and It indicates that the multiplication calculation of the entire sparse polynomial has been completed, returns to the IDLE state, and the control state machine writes 0x1 to the working status register; otherwise, it enters the P1 state.

[0076] (j) P1 state: The control state machine writes the value of the register way1 of the first path to the storage row where w is located at the RAM address addrW+(2g+1)N, and increments the value of g by 1. If (2g+1)N On the storage row where it is located, and increments the value of g by 1. If It indicates that the calculation of the entire polynomial has been completed, returns to the IDLE state; otherwise, it enters the B0 state to calculate the next set of multiplication results w 2gN ~w (2g+2)N-1 value.

[0077] Next, it will be combined with Figure 8 and Figure 10 to specifically demonstrate the operation method of sparse polynomial multiplication in HQC128. At this time, n = 17669, ω = 66, the size of the RAM storage row is 128 bits, that is, N = 128, so log2N = 7, Let the non-zero coefficient indexes in the sparse polynomial V be P0 = 227, P1 = 445,..., P 65 = 17164; Let the starting address of the dense polynomial U in the RAM be addrU = 0, and the starting address of the result polynomial W in the RAM be addrW = 17792. The calculation steps are as follows:

[0078] (1) After the state machine is reset, it enters the IDLE state. If the AXI bus host writes 0x3 to the status register, it enters the PR state; otherwise, it remains in the IDLE state. In the PR state, it requests the storage row where u0 is located at the RAM address 0, and obtains 128 bits of u0~u 127 ; Requests the storage row where u 17668 is located at the RAM address 17668, and the read result is {u 17664 ~u 17668 , 123′b0}. Then splice to get {u 17664 ~u 17668 , u0~u 122} and write it to the storage row where u 17668 is located. At this time, the values of the RAM addresses 17668~17791 change from the original 0 to u0~u 122 , and the values of the addresses 17664~17668 remain unchanged. If it has been written to the RAM, it enters the B0 state to calculate the multiplication results of the two paths of polynomials in the 0th group.

[0079] (2) In the B0 state, set i, j, way0, and way1 to zero. At this time, P i = P0 = 227, calculate D i = (n - P i + 2gN) mod n = 17442, then request the storage row at RAM address 17442 and enter the R0 state.

[0080] (3) In the R0 state, read the storage row at RAM address 17442 to obtain 128 bits of u 17408 ~ u 17535 and write them into register reg00. Since So DN i = D i + N = 17570, then request the storage row at RAM address 17570, and then enter the R1 state.

[0081] (4) In the R1 state, read the storage row at RAM address 17570, obtain 128 bits of u 17536 ~ u 17663 and write them into register reg01, then concatenate to get S[0:255]. Among them, the data in reg00 is in the lower N bits, that is, S[0:127] = {u 17408 ~ u 17535}; the data in reg01 is in the higher N bits, that is, S[128:255] = {u 17536 ~ u 17663}, so S[0:255] = {u 17408 ~ u 17663}. Also, when introducing the multiply-accumulate unit, it is mentioned that k0 = D i mod N = 34, so extract the value of S[34:161], obtain 128-bit data of u 17442 ~ u 17569 and send it to the 0th path of the multiply-accumulate unit. At this time, way0 is equal to 0, so after performing an exclusive OR operation with 0, update the value of register way0 to get way0 = {u 17442 ~ u 17569}, completing the first exclusive OR operation. At this time, j = 0, so D j = (n - P j + (2g + 1)N) mod n = 17570, request the storage row of u 17570 at RAM address 17570. Increment i by 1 and enter the R2 state.

[0082] (5) In the R2 state, first read the storage row of u 17570 at RAM address 17570 to obtain u 17536~u 17663 128 bits are written into register reg01. So DN j =D j +N=17698, requesting the storage row at RAM address 17698. At this time, i is less than ω, and the R3 state is entered.

[0083] (6) In the R3 state, first read the storage row at RAM address 17698 to obtain {u 17664 ~u 17668 ,u0~u 122} and write it into register reg11, and then concatenate to get T[0:255], where reg01 data is in the lower N bits, that is, T[0:127] = {u 17536 ~u 17663}; reg11 data is in the high Nbit, that is, T[128:255] = {u 17664 ~u 17668 ,u0~u 122}, so T[0:255]={u 17536 ~u 17668 ,u0~u 122}. When introducing the multiplication and addition unit, it is mentioned that k1=D j mod N = 34, so extract the value of T[34:161] and obtain {u 17570 ~u 17668 ,u0~u 28} is sent to the first way of the multiplication and addition unit. At this time, way1 is equal to 0, so the value of register way1 is updated after XOR operation with 0, and the result is way1 = {i 17570 ~i 17668 ,i0~u 28}, complete the second XOR operation, so far completed Figure 10 The calculation of the corresponding column of P0=227 in group 0. Since i=1, P i =P1=445,D i =(nP i +2gN)modn=17224, so the requested RAM address is addrU+D i =17224 The storage row where the data is located. Add 1 to j and enter the R0 state.

[0084] (7) In the R0 state, read the memory row at RAM address 17224 and write reg00 = {u 17152 ~u 17279},because So DN i= 17352, request the memory row at RAM address 17352 and enter the R1 state; in the R1 state, read the memory row at RAM address 17352 and write it into reg01 = {u 17280 ~u 17407}, concatenate S[0:255] = {u 17152 ~u 17407}, k0 = D i mod N = 72, so extract the value of S[72:199], obtain the 128-bit data of u 17224 ~u 17351 and send it to the 0th path of the multiply-accumulate unit. At this time, way0 = {u 17442 ~u 17569}, perform an exclusive OR operation and update the value of the register way0 to obtain Complete the 3rd exclusive OR operation. At this time, j = 1, so P j = 445, D j = (n - P j + (2g + 1)N) mod n = 17352, so request the memory row where u 17352 is located at RAM address 17352, increment i by 1, and enter the R2 state. In the R2 state, read the memory row at RAM address 17352 and write it into reg01 = {u 17280 ~u 17407}, because So DN i = 17480, request the memory row at RAM address 17480. Since i is less than ω, enter the R3 state. In the R3 state, read the memory row at RAM address 17480 and write it into reg11 = {u 17408 ~u 17535}, concatenate T[0:255] = {u 17280 ~u 17535}, k1 = D j mod N = 72, so extract the value of T[72:199], obtain the 128-bit data of u 17352 ~u 17479 and send it to the 1st path of the multiply-accumulate unit. At this time, way1 = {u 17570 ~u 17668 , u0~u 28}, perform an exclusive OR operation and update the value of the register way1 to obtain Complete the 4th exclusive OR operation. Thus, complete the calculation of the column corresponding to P1 = 445 in the 0th group in Figure 10 . And so on, until the calculation of the 0th group of P 65Calculation of the column corresponding to 17164, and then enter the P0 and P1 states in sequence. Write the calculation results w0 to w stored in the register way0 127 to 128 bits at RAM addresses from addrW to addrW + 127. Write the calculation results w 128 ~w 255 stored in the register way1 to 128 bits at RAM addresses from addrW + 128 to addrW + 255. Then enter the B0 state to perform the calculation of the first two-way group.

[0085] (8) The calculation of the first two-way group is the same as that of the 0th group. It also starts from the B0 state and sequentially completes the calculations of P0 = 227, P1 = 445,..., P 65 = 17164 for the corresponding columns. Finally, write the calculation results w 256 ~w 383 stored in the register way0 to 128 bits at RAM addresses from addrW + 256 to addrW + 383. Write the calculation results w 384 ~w 511 stored in the register way1 to 128 bits at RAM addresses from addrW + 384 to addrW + 511. Then enter the B0 state to perform the calculation of the second two-way group until the calculation of the 69th group is completed and the calculation results w 17664 ~w 17668 are written into the RAM.

[0086] The above embodiments are written with System Verilog for the implementation code and simulated with VCS, Figure 11 showing the simulation waveform diagram of some signals of this embodiment viewed with Verdi. The selected signals are: clk is the clock signal, status is the status register, cout_g is the current calculation group number g, cout_i is the value i of the 0th-way counter, cout_j is the value j of the 1st-way counter, addrU_bit is the polynomial U starting address register addrU, addrW_bit is the polynomial W starting address register addrW, sram_addr_bit is the read / write RAM address signal, sram_rdata_i is the read RAM data signal, sram_wdata_o is the write RAM data signal, and sram_we_o is the write RAM enable signal.

[0087] From the first part of the waveform, it can be seen that the starting address addrU of the polynomial U is set to 0x0, and the starting address addrW of the polynomial W is set to 17792. At the cursor M1, the value of the register status becomes 0x3, and the accelerator first requests the storage row at RAM address 17668 and reads {u17664 ~u 17668 ,123′b0}; then request the storage row at RAM address 0 and read u0~u 127 ; Then concatenate to get {u 17664 ~u 17668 ,u0~u 122} and write it to the storage row where RAM address 17668 is located. At cursor M4, the current group number g is equal to 0, and the calculation of the two paths corresponding to column P0 in group 0 is started, requesting RAM address addrU+D in turn. i =17442, addrU+DN i =17570 and complete the first XOR operation, request addrU+D j =17570, addrU+DN j =17698 and complete the second XOR operation. At cursor M5, start the calculation of the two paths corresponding to column P1 in group 0, and request the RAM address addrU+D in turn. i =17224, addrU+DN i =17352 and complete the third XOR operation, request addrU+D j =17352, addrU+DN j =17480 and complete the fourth XOR operation. At cursor M6, start the calculation of the two paths corresponding to column P2 in group 0, and so on. Until cursor M2, the values of i and j both become 66, indicating that the two-path calculation of group 0 is completed, and are written to the storage rows where addresses addrW+2gN=17792 and addrW+(2g+1)N=17920 are located in sequence, and then the current group number g becomes 1, and the two-path calculation of group 1 and other groups is started. Until cursor M8, the calculation of g=69 is completed and written to the storage row where address addrW+2gN=35456 is located. At this point, all the calculation results of groups 0 to 69 are written into RAM, and the calculation is completed.

[0088] Figure 12 This is an integration method of the accelerator of the present invention in SoC, which adopts the two-level bus architecture of AXI and APB, high-speed peripherals, instruction RAM, data RAM, and accelerator are mounted on the AXI bus, low-speed peripherals are mounted on the APB bus, and the APB bus is bridged to the AXI bus; the processor core can directly access the instruction RAM and data RAM, and can also initiate access to the peripherals as the host of the AXI bus through a bridge circuit; the accelerator is mounted on the AXI bus and accesses the data RAM through the bridge circuit.

[0089] The above-described embodiments merely represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent for the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.

Claims

1. A sparse polynomial multiplication accelerator applied to the HQC algorithm, characterized in that It includes a sparse polynomial non-zero coefficient index register group, a status register group, a multiply-accumulate unit, a control state machine, and an address generation module; The sparse polynomial non-zero coefficient index register group is used to store the indexes of the non-zero coefficients of the externally input sparse polynomial, and provide the indexes to the control state machine, which will then input the obtained indexes into the address generation module; The status register group is used to configure and represent the working state of the accelerator according to the external or control state machine's state input, and store the Hamming weight of the externally input sparse polynomial, the starting address of the dense polynomial in the external RAM, the length of the dense polynomial, and the starting address of the result polynomial in the external RAM, and provide the information stored in the status register group to the address generation module and the control state machine; The address generation module is used to calculate the address information based on the index input by the control state machine and the input of the status register group, and feedback the calculated address information to the control state machine; The control state machine reads the coefficients of the dense polynomial from the external RAM based on the address information and the starting address of the dense polynomial in the external RAM, and inputs them into the multiply-accumulate unit for calculation. After the multiply-accumulate unit completes the multiplication calculation, the control state machine writes the calculation result of the multiply-accumulate unit to the external RAM based on the starting address of the result polynomial in the external RAM.

2. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 1, wherein The set \(P\) of indices of non-zero coefficients in the non-zero coefficient index register group of the sparse polynomial is \(P = \{P_0, P_1, \ldots, P\) ω-1 \} \in \{0, 1, \ldots, n - 1\}, where \(\omega\) is the number of non-zero coefficients in the sparse polynomial; \(P_0\) is the index of the first non-zero coefficient, \(P_1\) is the index of the second non-zero coefficient; and \(n\) is the length of the dense polynomial.

3. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 2, characterized in that, The status register group includes a working state register, a Hamming weight register, a dense polynomial starting address register, a polynomial length register, and a result polynomial starting address register; The working state register is used to configure and represent the working state of the accelerator according to the external state input or the control state machine's state input. When 0x3 is written to the working state register externally, the control state machine will start the polynomial multiplication calculation. After the calculation is completed, the control state machine will write 0x1 to the working state register, indicating that the polynomial multiplication calculation is completed; The Hamming weight register is used to store the Hamming weight of the externally input sparse polynomial; The dense polynomial starting address register is used to store the starting address of the externally input dense polynomial in the RAM; The polynomial length register is used to store the length of the externally input dense polynomial; The result polynomial starting address register is used to store the starting address of the result polynomial calculated by the accelerator in the RAM.

4. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 3, wherein The multiply-accumulate unit includes a register reg00, a register reg01, a register reg11, a 0th path extraction module, a 1st path extraction module, a 0th path register way0, and a 1st path register way1; When the control state machine writes data to register reg00 and register reg01, the data written to register reg00 and register reg01 are respectively input to the 0th extraction module. The 0th extraction module splices the received data to obtain 2N-bit data, denoted as S[0:2N-1]. Let k0 = D i mod N, D i is the address offset calculated by the address generation module based on the index P i N is the amount of data stored in one row of the external RAM. The 0th extraction module extracts N-bit data starting from k0 from the spliced 2N-bit data and outputs it, denoted as S[k0:k0+N-1]. Then, the N-bit data output by the 0th extraction module is XORed with the data in the 0th register way0 at the current moment and output to the 0th register way0 to update the value in the 0th register way0; among them, the data written to register reg00 and register reg01 are both N-bit. When the 0th extraction module performs splicing, the data input by register reg00 is in the low N bits, that is, S[0:N-1], and the data input by register reg01 is in the high N bits, that is, S[N:2N-1]; at the initial moment, the data in the 0th register is N-bit 0s; When the control state machine writes data to register reg01 and register reg11, the data written to register reg01 and register reg11 are respectively input into the first extraction module. The first extraction module splices the received data to obtain 2N-bit data, denoted as T[0:2N-1]. Let k1 = D j mod N, D j is the address offset calculated by the address generation module based on the index P j The first extraction module extracts N-bit data starting from k1 from the spliced 2N-bit data and outputs it, denoted as T[k1:k1+N-1]. Then, the N-bit data output by the first extraction module is XORed with the data in the first register way1 at the current moment and output to the first register way1 to update the value in the first register way1; among them, when the first extraction module performs splicing, the data input by register reg01 is in the lower N bits, that is, T[0:N-1], and the data input by register reg11 is in the higher N bits, that is, T[N:2N-1]; at the initial moment, the data in the first register is N bits of 0.

5. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 4, characterized in that, The information provided by the status register group to the control state machine includes the Hamming weight of the sparse polynomial, the starting address of the dense polynomial in the RAM, the length of the dense polynomial, and the starting address of the result polynomial in the RAM; Before the control state machine starts to control the multiply-accumulate unit to perform multiplication calculation, it will first configure and initialize the calculation group number g, that is, set the calculation group number g to 0, and at the same time output the calculation group number g to the address generation module; When the control state machine starts to control the multiply-accumulate unit to perform multiplication, the control state machine configures and initializes the value i of the counter on the 0th path and the value j of the counter on the 1st path of the multiply-accumulate unit, that is, both the value i of the counter on the 0th path and the value i of the counter on the 1st path are set to 0; then the control state machine reads the index P of the corresponding non-zero coefficient of the sparse polynomial from the sparse polynomial non-zero coefficient index register group based on the value i of the counter on the 0th path and the value j of the counter on the 1st path i and P j , and outputs the obtained index P of the non-zero coefficient of the sparse polynomial i and P j to the address generation module.

6. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 5, wherein The address generation module receives the length of the polynomial input by the status register group and the calculation group number g input by the control state machine, and the index P of the non-zero coefficients of the sparse polynomial i and P j , and outputs four address information: The address generation module calculates the address offset D i based on the index P i , the address offset DN i , the address offset D j calculated by the address generation module based on the index P j and the address offset DN j .

7. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 6, wherein Method for obtaining address offset D i is as follows: First, shift the received computing group number g to the left by log2N + 1 bits, denote the result as 2gN, and then subtract the index P of the non-zero coefficient of the sparse polynomial from the result 2gN i to get 2gN - P i , and output it to the 1 input port of the first multiplexer; Then subtract the length n of the dense polynomial from the index P of the non-zero coefficients of the sparse polynomial i to obtain n - P i . Then add n - P i to the result 2gN to get n - P i + 2gN, and output it to the 0 input port of the first multiplexer; Input n - P i + 2gN and n into the first comparator. If n - P i + 2gN ≥ n, the first comparator outputs 1 to the selection signal port of the first multiplexer, and the first multiplexer finally outputs the signal of the 1 input port, that is, D i = 2gN - P i ; Otherwise, the first comparator outputs 0 to the selection signal port of the first multiplexer, and the first multiplexer finally outputs the signal of the 0 input port, that is, D i = n - P i + 2gN; Method for obtaining address offset DN i is as follows: After obtaining address offset D i , add address offset D i to N to obtain D i +N and input it to the 0 input port of the second multiplexer. At the same time, subtract the length n of the dense polynomial from D i +N to obtain D i +N-n, and then input D i +N-n to the 1 input port of the second multiplexer; Calculate the floor of n divided by N, and then multiply the result by N to get Input D i and into the second comparator. If the second comparator outputs a selection signal to the selection signal port of the second multiplexer, and the second multiplexer finally outputs the signal of the 1 input port, that is, DN i = D i + N - n; otherwise, the second comparator outputs 0 to the selection signal port of the second multiplexer, and the second multiplexer finally outputs the signal of the 0 input port, that is, DN i = D i + N.

8. The sparse polynomial multiplication accelerator applied to the HQC algorithm according to claim 7, wherein Method for obtaining address offset D j is as follows: First, record the result as the sum of 2gN and N, obtaining (2g + 1)N. Then subtract the index P of the non-zero coefficient of the sparse polynomial from the result (2g + 1)N j to get (2g + 1)N - P j , and output it to the 1 input port of the third multiplexer; Then subtract the length n of the dense polynomial from the index P of the non-zero coefficients of the sparse polynomial j to obtain n - P j . Then add n - P j to (2g + 1)N to get n - P j +(2g + 1)N, and output it to the 0 input port of the third multiplexer; Input n - P j +(2g + 1)N and n into the third comparator. If n - P j +(2g + 1)N ≥ n, then the third comparator outputs 1 to the selection signal port of the third multiplexer, and the third multiplexer finally outputs the signal of the 1 input port, that is, D j =(2g + 1)N - P j ; Otherwise, the third comparator outputs 0 to the selection signal port of the third multiplexer, and the third multiplexer finally outputs the signal of the 0 input port, that is, D j =n - P j +(2g + 1)N; Method for obtaining address offset DN j is as follows: After obtaining address offset D j , add address offset D j to N to get D j +N and input it to the 0 input port of the fourth multiplexer. At the same time, subtract the length n of the dense polynomial from D j +N to get D j +N - n, and then input D j +N - n to the 1 input port of the fourth multiplexer; Input D j and to the fourth comparator. If , then the fourth comparator outputs 1 to the selection signal port of the fourth multiplexer, and the fourth multiplexer finally outputs the signal of the 1 input port, that is, DN j = D j +N - n; Otherwise, the fourth comparator outputs 0 to the selection signal port of the fourth multiplexer, and the fourth multiplexer finally outputs the signal of the 0 input port, that is, DN j = D j +N.

9. A sparse polynomial multiplication calculation method for the sparse polynomial multiplication accelerator applied to the HQC algorithm described in claim 8, characterized in that, It includes: 1) The control state machine initializes the calculation group number g and sets the calculation group number g to 0; 2) The sparse polynomial non-zero coefficient index register group receives and stores the indices of the non-zero coefficients of the externally input sparse polynomial; The Hamming weight register receives the Hamming weight of the externally input sparse polynomial, and the dense polynomial starting address register receives the starting address of the externally input dense polynomial in the RAM; The polynomial length register receives the length of the externally input polynomial; The result polynomial starting address register receives the starting address of the result polynomial in the RAM; The working status register receives the externally written 0x3 to enable the control state machine to start polynomial multiplication calculation; 3) The control state machine requests and reads the dense polynomial coefficient u0 and the dense polynomial coefficient u from the external RAM n-1 from the storage row where they are located. The size of one storage row is N bits. Concatenate the N-(n mod N) bits from u0 to the dense polynomial coefficient u (N-(nmodN)-1) to the back of u n-1 to obtain and write it to the storage row where u n-1 is located; 4) Clear the values i of the 0th counter and the values j of the 1st counter of the multiply-accumulate unit, and at the same time clear the values of the 0th register and the 1st register of the multiply-accumulate unit; then the control state machine reads the indices P of the corresponding sparse polynomial non-zero coefficients from the sparse polynomial non-zero coefficient index register group based on i and j i and P j , and index P i and P j as well as calculate the group number g and output them to the address generation module, and receive the address offset D i , address offset DN i , address offset D j and address offset DN j ; then the control state machine requests the storage row where is located from the external RAM; 5) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg00 of the multiply-accumulate unit, and requests the storage row where it is located; 6) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg01 of the multiply-accumulate unit. The control state machine controls the multiply-accumulate unit to perform the 0th path calculation and updates the value of the 0th path register way0; the control state machine then requests the storage row where it is located, and increments the value of the 0th path counter i by 1; 7) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg01 of the multiply-accumulate unit. Then the control state machine requests the storage row where it is located. If i is equal to ω at this time, go to step 9); otherwise, continue with step 8). 8) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg11 of the multiply-accumulate unit. The control state machine controls the multiply-accumulate unit to perform the first-way calculation and updates the value of the first-way register way1; the control state machine then requests the storage row where it is located, increments the value of the first-way counter j by 1, and then proceeds to step 5); 9) The control state machine reads the entire row data of the storage row where it is located and writes it into the register reg11 of the multiply-accumulate unit. The control state machine controls the multiply-accumulate unit to perform the first-way calculation and updates the value of the first-way register way1; then writes the value of the zero-way register way0 to the external RAM w 2gN on the storage row where it is located; if at this time and then return to step 1), otherwise proceed to step 10); 10) The control state machine writes the value of the first path register way1 to the external RAM w (2g+1)N on the storage row where it is located, and increments the value of the calculation group number g by 1; if it indicates that the multiplication calculation of the entire sparse polynomial is completed, and the control state machine writes 0x1 to the working state register; otherwise, return to step 4) to continue the multiplication calculation of the sparse polynomial.

10. An SoC integrating the sparse polynomial multiplication accelerator for the HQC algorithm according to any one of claims 1-8.

Citation Information

Cited By

  • Sparse adder circuit suitable for HQC algorithm

    CN121614111A

  • A sparse adder circuit suitable for HQC algorithm

    CN121614111B