Fast modular multiplication system based on base transformation
By adopting a fast mode multiplication system based on transform basis in the grid cryptographic system, the problems of many reduction iterations and complex connections are solved, and more efficient hardware resource usage and calculation speed are achieved.
Patent Information
- Application Number
- PCT/CN2024/136124
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-12
AI Technical Summary
In the prior art, the number of reduction iterations of grid cryptographic systems is large and the connection is complex, resulting in large hardware resource consumption and high computing delay.
A fast modular multiplication system based on transform basis is used, which includes a pre-computing layer, a polynomial modular layer, an iterative reduction layer, and a phase reduction layer. The analog and input data are converted into X-phase through transform basis processing, and the grouping multiplication and recombination reduction processing are performed using adders and multipliers to shorten the coefficient bit width of the polynomial and further optimized through the iterative reduction layer.
It effectively reduces the use of circuit hardware, saves the area of the circuit, and has advantages in area and speed, making it more efficient than traditional mode multiplication systems.
Smart Images

Figure CN2024136124_12062025_PF_FP_ABST
Abstract
Description
A Fast Modular Multiplication System Based on Transformation Bases Technical Field The present invention relates to the technical fields of computers and integrated circuits, and particularly to a fast modular multiplication system based on transformation bases. Background Art With the rapid development of quantum computing technology, the traditional public-key cryptosystem that is currently widely used and based on the problems of large integer factorization and discrete logarithm has the risk of being broken and invalidated by quantum computers, which will seriously damage the security, confidentiality, and integrity of digital communications. In recent years, the cryptographic community has actively studied a new public-key cryptosystem that can resist quantum computing attacks and is thus called "post-quantum cryptography". Among them, lattice-based encryption algorithms, due to advantages such as strong security, are one of the most concerned post-quantum cryptos, and a large number of studies on the optimization of this encryption algorithm have been carried out in the academic and industrial circles. In lattice cryptography, the modular multiplication operation is one of the most core operations, occupying most of the computing time and hardware resources of the cryptosystem. Therefore, the optimization of modular multiplication can effectively improve the efficiency of the entire cryptosystem. Generally, there are two ways to implement modular multiplication: The first way is to take "multiply first, then reduce", where multiplication and reduction are relatively independent, which is called non-interleaved modular multiplication. Among them, the reduction process is relatively complex and is the focus of optimizing modular multiplication. Common reduction algorithms include Montgomery reduction algorithm and Barrett reduction algorithm. In addition, there is a "lookup table - accumulation" method in the prior art, which is a very effective method for modular reduction in the finite field of lattice cryptography; The second way is to take "multiply while reducing", combining multiplication and reduction through an interleaved method, which is called interleaved modular multiplication. Common algorithms include Montgomery modular multiplication algorithm. The reduction method based on "lookup table - accumulation" is generally described and summarized as follows: For a fixed modulus q, its bit width can be expressed as Then q can be expressed as follows: q = 2 N -δ (1 ≤ δ < 2 N-1 ) Furthermore, we get 2 N ≡ δ (mod q) For the number to be reduced z[D - 1:0] (D > N), we have: z ≡ z[D - 1:N]·2 N + z[N - 1:0] ≡ z[D - 1:N]·δ + z[N - 1:0] ≡ z′[D′ - 1:N′]·2 N + z'[N' - 1:0] + z[N - 1:0] ≡z′[D′-1:N′]·δ+z′[N'-1:0]+z[N-1:0] ≡…(mod q); The modulus q applicable to the "look-up table - accumulation" reduction method generally has the following two characteristics based on the above expression: (1) The value of δ is small, that is, the bit width is small; (2) The Hamming weight of δ is low, that is, the number of non-zero bit positions is small. Therefore, this method can quickly reduce the bit width to achieve the purpose of modulo operation. Compared with the traditional Barrett algorithm and Montgomery algorithm, it can save additional multiplication overhead, reduce the area, and reduce the delay, which is a very effective reduction method. However, the limitations of this method are also obvious. In less ideal situations (larger or higher Hamming weight), the reduction will face problems such as a large number of iteration times and complex wiring, resulting in an increase in both area and time complexity. Therefore, there are very few moduli that can effectively adopt this method in practice. Summary of the Invention The present application provides a fast modular multiplication system based on a transformation basis to solve the problems of a large number of reduction iterations and complex wiring in the lattice cryptosystem in the prior art. The fast modular multiplication system includes: A pre-computation layer for performing transformation basis processing on the input first modular multiplication input A and second modular multiplication input B, thereby converting the first modular multiplication input A and the second modular multiplication input B from binary to X - base. The first modular multiplication input A and the second modular multiplication input B are obtained based on the modulus q; A polynomial modular multiplication layer including a plurality of adders and multipliers; the polynomial modular multiplication layer is used to perform grouped multiplication processing and recombination reduction processing on the intermediate value coefficients of the first modular multiplication input A and the second modular multiplication input B that have completed the transformation basis processing in sequence through the built-in adders and multipliers, thereby reducing the degree of the polynomial including the first modular multiplication input A and the second modular multiplication input B; An iterative reduction layer including a plurality of sub-iterative reduction layers, and each sub-iterative reduction layer has a plurality of adders; the iterative reduction layer is used to perform a plurality of mapping wiring processing and accumulation processing on the polynomial including the first modular multiplication input A and the second modular multiplication input B that has completed the recombination reduction processing in sequence through the built-in adders, thereby shortening the coefficient bit width of the (n - 1)th polynomial of X including the first modular multiplication input A and the second modular multiplication input B; A base restoration layer for converting the polynomial including the first modular multiplication input A and the second modular multiplication input B that has completed the last accumulation processing from X - base to binary to obtain the actual modular multiplication result. Preferably, the pre-computation layer is further used for: Set the modulus \(q\), the bit width \(N\) of the modulus \(q\), the first modular multiplication input \(A\), and the second modular multiplication input \(B\) in the modular multiplication operation; Convert the first modular multiplication input \(A\) and the second modular multiplication input \(B\) into polynomials of \(X\) according to the transformation basis formula; the transformation basis formula is: \(X = 2\) α \(^{-1}\), where \(\alpha\) is a positive integer and \(\alpha\lt N\); the polynomial of \(X\) is \(q(X)=kX\) n \(-\delta(X)\); Obtain the first target polynomial according to the polynomial of \(X\); the first target polynomial is composed of the initial polynomials of the first modular multiplication input \(A\) and the second modular multiplication input \(B\), respectively: The initial polynomial of the modulus \(q\), \(kX\) n \(\bmod q=\delta(X)\); The initial polynomial of the first modular multiplication input \(A\) The initial polynomial of the second modular multiplication input \(B\) where \(\delta(X)\) is a polynomial about \(X\), \(\vert r\) i \(\vert(i = 0,1,2…)\), and are the coefficients of the polynomials \(A(X)\) and \(B(X)\) respectively, and \(i\) is a constant. Preferably, the polynomial multiplication layer further includes a grouped multiplication layer and a recombination reduction layer; The grouped multiplication layer includes a number of adders and multipliers, and the recombination reduction layer includes a number of adders; The grouped multiplication layer is used to perform grouped multiplication processing on the first target polynomial to obtain a second target polynomial; The recombination reduction layer is used to perform recombination reduction processing on the second target polynomial to obtain a polynomial of degree \((n - 1)\) of \(X\). Preferably, the grouped multiplication layer is further used for: Convert the calculation relationship in the first target polynomial from constant multiplication to polynomial multiplication and convert individual multiplication to grouped multiplication through the built-in adders and multipliers to obtain the second target polynomial. Preferably, the recombination reduction layer is further used for: Perform recombination and merging processing on the second target polynomial through the built-in adders to obtain a recombined polynomial; Perform reduction processing on the recombined polynomial to obtain a reduced polynomial; Substitute the reduced polynomial into the initial polynomial of the modulus \(q\) multiple times to obtain the polynomial of degree \((n - 1)\) of \(X\). Preferably, the second target polynomial is The recombinant polynomial is: The reduction polynomial is: The polynomial of degree (n - 1) of X is: Wherein, u i is the new coefficient of the polynomial obtained by accumulating several terms, and u i has a bit width of bits, and c i is a coefficient and is 2M bits; n is a constant; is M bits; is a coefficient and is M bits, and k is the coefficient of the highest-degree term. Preferably, the sub-iteration reduction layer includes a mapping wiring layer and an accumulation layer, and the accumulation layer includes a plurality of adders; The mapping wiring layer is used to perform mapping wiring processing on the polynomial of degree (n - 1) of X; The accumulation layer is used to perform accumulation processing on the polynomial of degree (n - 1) of X that has completed one-time mapping wiring processing. Preferably, the iteration reduction layer is further used for: Performing a plurality of mapping wiring processes and accumulation processes on the polynomial of degree (n - 1) of X that has completed recombinant reduction processing in sequence through the built-in adder; When the coefficient bit widths of the polynomial of degree (n - 1) of X are all reduced to the required range, output the third target polynomial. Preferably, the iteration reduction layer is further used for: Converting the polynomial of degree (n - 1) of X that has completed recombinant reduction processing into a polynomial in the form of u i The polynomial in the form of u i is Substituting the polynomial in the form of u i into the initial polynomial of the modulus q and the polynomial of degree (n - 1) of X for reduction multiple times to obtain the third target polynomial, and the third target polynomial is Wherein, z i is a coefficient. Preferably, the radix restoration layer is further used for: Perform a number of addition and shift operations on the third target polynomial to obtain the actual modular multiplication result. This application provides a fast modular multiplication system based on a transformed base. The fast modular multiplication system includes a pre-computation layer for performing a transformed base process on an input modulus q to transform the modulus q from binary to base-X; a polynomial modular multiplication layer including a number of adders and multipliers; the polynomial modular multiplication layer is used to perform grouped multiplication and recombination reduction processes on the modulus q that has completed the transformed base process in sequence through the built-in adders and multipliers to reduce the degree of the polynomial containing the modulus q; an iterative reduction layer including a number of sub-iterative reduction layers, each of the sub-iterative reduction layers having a number of adders; the iterative reduction layer is used to perform a number of mapping wiring and accumulation processes on the polynomial containing the modulus q that has completed the recombination reduction process in sequence through the built-in adders to shorten the coefficient bit width of the (n-1)th degree polynomial containing the base-X modulus; a base restoration layer for transforming the polynomial containing the modulus q that has completed the last accumulation process from base-X to binary to obtain the actual modular multiplication result. This application reduces the use of circuit hardware, saves the circuit area, and has more advantages in terms of area and speed compared with traditional modular multiplication systems through the above fast modular multiplication system. Description of the Drawings FIG. 1 is a schematic diagram of a fast modular multiplication system based on a transformed base according to this application; FIG. 2 is a logic flow chart of a fast modular multiplication system based on a transformed base according to this application; FIG. 3 is a schematic diagram of the process of performing multiple modular multiplications by a fast modular multiplication system based on a transformed base according to this application. Detailed Embodiments It should be noted that in this application, the concept of each "layer" can also be referred to as a "unit" or "module", which should be regarded as a structure with a certain function composed of several hardware components. When describing the specific composition of different layers, in addition to referring to the hardware included in each layer structure, it also includes the functions of the hardware combined with the methods executed by the software configured according to the hardware. In this embodiment, the main hardware components of each layer of the hardware structure include adders and multipliers. Among them, an adder is an electronic circuit used to perform binary addition operations, capable of processing two or more addition operations. According to the number of bits, adders can be divided into one-bit adders, multi-bit adders, etc. In practical applications, multiple adders can also be combined according to preset logical operations to form an adder combination to achieve higher-bit addition operations. A multiplier is an electronic circuit or device used to perform the multiplication operation of two binary numbers. The basis of a multiplier is the adder structure, and it can be essentially regarded as an algorithm of "shift and add". According to the number of bits processed, multipliers can be divided into one-bit multipliers, multi-bit multipliers, etc. In practical applications, multiple multipliers can be combined to form a multi-bit multiplier to achieve higher-bit multiplication operations. Referring to FIGS. 1 and 2, it can be seen that this embodiment provides a fast modular multiplication system based on a transformation base. The fast modular multiplication system includes: A pre-computation layer 100, which is used to perform transformation base processing on the input first modular multiplication input A and second modular multiplication input B, thereby converting the first modular multiplication input A and the second modular multiplication input B from binary to X - base. The first modular multiplication input A and the second modular multiplication input B are obtained based on the modulus q. Specifically, in this embodiment, different from the traditional binary representation method, in this solution, before modular multiplication, the modulus and the input need to be subjected to transformation base processing. Through the pre-computation layer 100, the modulus q, the bit width N of the modulus q, the first modular multiplication input A and the second modular multiplication input B in the modular multiplication operation are set, and the first modular multiplication input A and the second modular multiplication input B are converted into polynomials of X according to the transformation base formula; the transformation base form is: X = 2 α α - 1, where α is a positive integer and α < N; the polynomial of X is q(X) = kX n n - δ(X), where n is the order of q(X) after X transformation, and k is the coefficient of the highest - order term; and a first target polynomial is obtained according to the polynomial of X, thereby performing base conversion on the first modular multiplication input A and the second modular multiplication input B. The pre - computation layer 100 can implement the above functions through a structure composed of low - latency adders, thereby improving the operation speed; in order to further improve the operation efficiency, low - latency adders can be combined to reduce the critical path length of the multiplier during modular multiplication. It should be noted that when selecting the value of α, two principles are followed: making the "Hamming weight" of the coefficients of δ(X) as small as possible, that is, |r i | (i = 0, 1, 2 …) is as small as possible and is 0, so as to reduce the wiring complexity; making the binary bit width of δ(X) as small as possible, so as to reduce the subsequent number of iterations. The specific process is as follows: The inputs for base conversion are two binary numbers to be modular - multiplied, that is: After 100 layers, they are respectively converted into n numbers (the n coefficients of the initial polynomial); among them, The initial polynomial kX of the modulus q n mod q = δ(X); then, The initial polynomial of the first modular multiplication input A The initial polynomial of the second modular multiplication input B Among them, δ(X) is a polynomial about X, |r i | (i = 0, 1, 2…), and are the coefficients of the polynomials A(X) and B(X) respectively, and i is a constant. It should be noted that the bit widths of the first modular multiplication input A and the second modular multiplication input B are the same as that of the modulus q, both being N. It should be noted that the modulus q can be determined according to requirements. Exemplarily, for example, the modulus q formulated for data in fields such as bank encryption systems and exchange encryption can be different. The fast modular multiplication system further includes: A polynomial modular multiplication layer 200, which includes a number of adders and multipliers; the polynomial modular multiplication layer 200 is used to perform grouped multiplication processing and recombination reduction processing on the intermediate value coefficients of the first modular multiplication input A and the second modular multiplication input B that have completed the transform base processing through the built-in adders and multipliers, so as to reduce the degree of the polynomials including the first modular multiplication input A and the second modular multiplication input B. Specifically, in this embodiment, the fast FIR algorithm is combined with the polynomials after the radix conversion through the polynomial modular multiplication layer 200, so as to reduce the number of multipliers used, thereby reducing the area of the circuit. Among them, the polynomial modular multiplication layer 200 further includes a grouped multiplication layer 210 and a recombination reduction layer 220; the grouped multiplication layer 210 includes a number of adders and multipliers, and the recombination reduction layer 220 includes a number of adders. In this embodiment, the grouped multiplication layer 210 performs grouped multiplication processing on the first target polynomial to obtain a second target polynomial. Specifically, the grouped multiplication layer 210 converts the calculation relationship in the first target polynomial from constant multiplication to polynomial multiplication and converts individual multiplications to grouped multiplications through the built-in adders and multipliers to obtain the second target polynomial. Among them, the multiplier is used to perform the multiplication operation of the polynomial, and the adder is used to perform the calculation of the intermediate result of the multiplication operation. In this embodiment, the second target polynomial is subjected to recombination reduction processing through the recombination reduction layer 220 to obtain an (n - 1)-th degree polynomial of X. Specifically, the recombination reduction layer 220 performs recombination and merging processing on the second target polynomial through an internal adder to obtain a recombination polynomial, and performs reduction processing on the recombination polynomial to obtain a reduced polynomial. Finally, the reduced polynomial is substituted into the initial polynomial of the modulus q multiple times to obtain the (n - 1)-th degree polynomial of X. Among them, the adder can not only calculate the result of the current bit, but also transmit the carry signal to the next bit to achieve continuous addition operations. In addition to basic addition operations, the adder can also implement other logic operations through different combinations of logic gates, such as exclusive OR, NAND, etc. These operations may be used for specific data conversion or verification in the recombination reduction processing. Among them, the second target polynomial is The recombination polynomial is: The reduced polynomial is: The (n - 1)-th degree polynomial of X is: Among them, u i is the new coefficient of the polynomial obtained by accumulating several terms. Denote u i The bit width is bits, and c i is a coefficient and is 2M bits; n is a constant; is M bits; is a coefficient and is M bits. Among them, in this embodiment, the fast FIR algorithm is adopted for this polynomial multiplication, which can effectively reduce the number of multipliers. The fast FIR algorithm is an algorithm used to improve the computational efficiency in digital signal processing. Its core idea is to decompose and recombine the FIR filter unit, at the cost of increasing the number of additions, in exchange for reducing the number of multiplications. Since the multiplier is more complex than the adder, reducing the number of multiplications can effectively improve the computational efficiency of the convolution operation. The fast FIR algorithm is also applicable to polynomial multiplication operations. The traditional multiplication of two (N - 1)-th order polynomials requires N 2 multiplications, while using the 2-parallel and 3-parallel fast FIR algorithms theoretically only requires 3 / 4N 2 and 2 / 3N 2 multiplications respectively, saving nearly 25% and 33.3% of the hardware overhead; for example, if N = 2, the fast FIR with 2-parallel can expand the expression in the following way: (a 1 X + a 0 )(b 1 X + b 0 ) = a 1 b 1 (X 2 ) + [(a 1 + a 0 )(b 1 + b 0 ) - a 1 b 1 - a 0 b 0 (X) + a 0 b 0 ; The required number of multiplications is reduced from the original 4 to 3, at the cost of increasing the number of adders from 1 to 4. Specifically, the essence of the polynomial modular multiplication layer 200 is to rewire the output of the previous multiplier and then use it as the input of n adders. The outputs of these n adders are denoted as u i (i = 0, 1, 2…n - 1), with a bit width of bits, and Specifically, the result obtained by the above modular multiplication is not fully reduced. To further reduce it, we can also represent u with a polynomial of X i : Among them, Substitute u i back into the original formula and repeatedly use the initial polynomial of the modulus q and the (n - 1)th polynomial of X. The essence of this process is to reorganize and merge the bits of the coefficients. After a certain mapping and wiring, and accumulation, a new round of coefficients is obtained to achieve the reduction purpose. Iterate the above process. Our ultimate goal is to reduce the bit width of the coefficients of the (n - 1)th polynomial of X with respect to the modulus to the required range, that is, α. The fast modular multiplication system further includes: An iterative reduction layer 300, which includes several sub-iterative reduction layers, and each sub-iterative reduction layer has several adders; the iterative reduction layer 300 is used to perform several mapping and wiring processes and accumulation processes on the polynomials containing the first modular multiplication input A and the second modular multiplication input B that have completed the reorganization and reduction process through the built-in adders, so as to shorten the bit width of the coefficients of the (n - 1)th polynomial of X containing the modulus, thereby simplifying the representation form of the polynomial. Specifically, in this embodiment, the bit width of the coefficients of the (n - 1)th polynomial of X containing the modulus is shortened by the iterative reduction layer 300. Among them, the sub-iteration reduction layer includes a mapping wiring layer 310 and an accumulation layer 320. The accumulation layer 320 includes a number of adders. The (n - 1)-th degree polynomial of X is subjected to mapping wiring processing through the mapping wiring layer 310. The (n - 1)-th degree polynomial of X that has completed one mapping wiring processing is subjected to accumulation processing through the accumulation layer 320. Among them, in this embodiment, the iteration reduction layer 300 is also used to sequentially perform a number of mapping wiring processes and accumulation processes on the (n - 1)-th degree polynomial of X that has completed the recombination reduction process through the built-in adder. When the coefficient bit widths of the (n - 1)-th degree polynomial of X are all reduced to the required range, the third target polynomial is output, thereby completing the reduction of the coefficient bit widths of the (n - 1)-th degree polynomial of X. The specific method for the iteration reduction layer 300 to complete the reduction of the coefficient bit widths of the (n - 1)-th degree polynomial of X is as follows: Convert the (n - 1)-th degree polynomial of X that has completed the recombination reduction process into a polynomial in the form of u i The polynomial in the form of u i is And substitute the polynomial in the form of u i into the initial polynomial of the modulus q and the (n - 1)-th degree polynomial of X for reduction multiple times to obtain the third target polynomial. The third target polynomial is Among them, z i is a coefficient with a bit width of α. The fast modular multiplication system further includes: A radix reduction layer 400, which is used to convert the polynomial containing the first modular multiplication input A and the second modular multiplication input B that has completed the last accumulation process from the X radix to the binary radix to obtain the actual modular multiplication result. Specifically, in this embodiment, through the radix reduction layer 400, several addition and shift operations are performed on the third target polynomial to obtain the actual modular multiplication result. Exemplarily, the following is an overall embodiment of this solution: Let the modulus q[22:0] = 8380417, and select a new radix X = 2 11 - 1, then q = 2X 2 - 1, and further: 2X 2 mod q = 1; Among them, the modulus q is the modulus selected by the lattice cryptography CRYSTALS-Dilithium, and this cipher has been standardized. Let the modular multiplication inputs be A[22:0] and B[22:0] respectively. For the convenience of subsequent reduction, we use the following polynomials containing X to represent A and B: A = a 1 2X + a 0 ; B = b 1 2X + b 0 ; where a 1 , a 0 , b 1 , b 0 are all 12-bit wide. Let the multiplication result be C[45:0]: C = A·B = (a 1 2X + a 0 )(b 1 2X + b 0 ); Using the fast FIR algorithm to expand and sort out this formula, we get: C = 2a 1 b 1 (2X 2 ) + [(a 1 + a 0 )(b 1 + b 0 ) - a 1 b 1 - a 0 b 0 (2X) + a 0 b 0 ; Using formula (3) to perform the first reduction on C: C mod q ≡ [(a 1 + a 0 )(b 1 + b 0 ) - a 1 b 1 - a 0 b 0 (2X) + 2a 1 b 1 + a 0 b 0 ; where, [(a 1 + a 0 )(b 1 + b 0 ) - a 1 b 1 - a 0 b 0 has a maximum bit width of 25 bits, (2a 1 b 1 + a0 b 0 ) The maximum bit width is 26 bits, which can be represented by u 1 [24:0] and u 0 [25:0] respectively, and then we have: C mod q ≡ u 1 [24:0](2X) + u 0 ; After re - arranging and organizing the coefficients of C, we get: C mod q ≡ (u 1 [24:15]·2 15 + u 1 [14:0])(2X) + u 0 [25:14]·2 14 + u 0 [13:0] ≡ (u 1 [24:15](16X + 16) + u 1 [14:0])(2X) + u 0 [25:14](8X + 8) + u 0 [13:0] ≡ (16u 1 [24:15] + u 1 [14:0] + 4u 0 [25:14])(2X) + 16u 1 [24:15] + 8u 0 [25:14] + u 0 [13:0]; By the same token, let v 1 [15:0] = 16u 1 [24:15] + u 1 [14:0] + 4u 0 ; v 0 [15:0] = 16u 1 [24:15] + 8u 0 [25:14] + u 0 ; We can get: C mod q ≡ v 1 [15:0](2X) + v 0 [15:0] ≡ (v 1 [15:11](X + 1) + v 1 [10:0])(2X) + v 0[15:12](2X + 2) + v 0 [11:0] ≡(v 1 [15:11] + v 1 [10:0] + v 0 [15:12])(2X) + v 1 [15:11] + 2v 0 [15:12] + v 0 [11:0] ≡w 1 [11:0](2X) + w 0 [12:0] ≡(w 1 [11:0] + w 0
[0012] )(2X) + 2w 0
[0012] + w 0 [11:0]. The advantages of this embodiment are as follows: The modulus that cannot be processed by the traditional "look-up table - accumulation" method is subjected to transformation base processing, making this solution more universal. Since the data is converted into polynomial expression after the transformation base, it can be combined with the fast FIR algorithm, and the number of multiplications can be further reduced through reuse, reducing the area of the circuit. An architecture of interleaved modular multiplication and an interleaved architecture of grouped modular multiplication are proposed. Compared with the traditional method in the reduction part, the use of multipliers is avoided, and it is composed of cascaded multi-layer adders, greatly reducing the hardware resources and calculation latency. In addition, in the hardware structure composed of the above-mentioned pre-computation layer 100, polynomial modular multiplication layer 200, iterative reduction layer 300, and radix reduction layer 400, in the implementation of specific data calculations, the steps executed by the pre-computation layer 100 and the radix reduction layer 400 are equivalent to converting the polynomial operation from the binary domain to the X - ary domain, that is, if the data needs to perform multiple modular multiplication operations, only the operations of the polynomial modular multiplication layer 200 and the iterative reduction layer 300 need to be executed multiple times, and the pre-computation layer 100 and the radix reduction layer 400 only need to complete one operation respectively (as shown in Figure 3), which can further reduce the latency of multiple modular multiplications. The fast modular multiplication system provided by the embodiment of this application achieves the original calculation effect and calculation accuracy, and at the same time has the effect of simplifying the hardware structure. Therefore, compared with the traditional modular multiplication system structure, it has advantages in circuit area and processing speed.
Claims
1. A fast modular multiplication system based on a transformation basis, characterized in that: The fast modular multiplication system comprises: A pre-computation layer, the pre-computation layer is used to perform base conversion processing on the input first modular multiplication input A and the second modular multiplication input B, so as to convert the first modular multiplication input A and the second modular multiplication input B from binary to X-ary, and the first modular multiplication input A and the second modular multiplication input B are obtained based on the modulus q; A polynomial modular multiplication layer, the polynomial modular multiplication layer comprising a plurality of adders and multipliers; the polynomial modular multiplication layer is used to sequentially perform group multiplication processing and reorganization reduction processing on the intermediate value coefficients of the first modular multiplication input A and the second modular multiplication input B that have completed the transformation basis processing through built-in adders and multipliers, so as to reduce the polynomial degree of the first modular multiplication input A and the second modular multiplication input B; an iterative reduction layer, the iterative reduction layer comprising a plurality of sub-iterative reduction layers, each of the sub-iterative reduction layers having a plurality of adders; the iterative reduction layer is used to sequentially perform a plurality of mapping, wiring and accumulation processes on the polynomial that has completed the reorganization and reduction process and includes the first modular multiplication input A and the second modular multiplication input B through the built-in adders, so as to shorten the coefficient bit width of the (n-1)-order polynomial of X including the first modular multiplication input A and the second modular multiplication input B; A base restoration layer is used to convert the polynomial containing the first modular multiplication input A and the second modular multiplication input B, which has completed the last accumulation process, from base X to binary to obtain an actual modular multiplication result.
2. A fast modular multiplication system based on a transformation basis according to claim 1, characterized in that: The pre-computation layer is also used to: Setting the modulus q, the bit width N of the modulus q, the first modular multiplication input A and the second modular multiplication input B in the modular multiplication operation; The first modular multiplication input A and the second modular multiplication input B are converted into polynomials of X according to the transformation basis formula; the transformation basis formula is: X=2 α -1, where α is a positive integer and α<N; the polynomial of X is q(X)=kX n -δ(X); A first target polynomial is obtained according to the polynomial of X; the first target polynomial is composed of initial polynomials of the first modular multiplication input A and the second modular multiplication input B, which are respectively: Initial polynomial kX modulo q n modq = δ(X); The first modular multiplication is the initial polynomial of input A. The second modular multiplication input B's initial polynomial Among them, δ(X) is a polynomial about X, and are the coefficients of polynomials A(X) and B(X) respectively, and i is a constant.
3. A fast modular multiplication system based on a transformation basis according to claim 2, characterized in that: The polynomial modular multiplication layer also includes a group multiplication layer and a reorganization reduction layer; The group multiplication layer includes a plurality of adders and multipliers, and the reorganization reduction layer includes a plurality of adders; The group multiplication layer is used to perform group multiplication processing on the first target polynomial to obtain a second target polynomial; The reorganization reduction layer is used to perform reorganization reduction processing on the second target polynomial to obtain a (n-1)th degree polynomial of X.
4. A fast modular multiplication system based on a transformation basis according to claim 3, characterized in that: The grouped multiplication layer is also used to: The calculation relationship in the first target polynomial is converted from constant multiplication to polynomial multiplication, and the single multiplication is converted to group multiplication through the built-in adder and multiplier, so as to obtain the second target polynomial.
5. A fast modular multiplication system based on a transformation basis according to claim 4, characterized in that: The reorganization-reduce layer is also used to: Reorganizing and merging the second target polynomial through a built-in adder to obtain a reorganized polynomial; Simplifying the reorganized polynomial to obtain a simplified polynomial; Substituting the reduced polynomial into the initial polynomial of the modulus q multiple times, a (n-1)th degree polynomial of X is obtained.
6. A fast modular multiplication system based on a transformation basis according to claim 5, characterized in that: The second target polynomial is The reorganization polynomial is: The reduced polynomial is: The (n-1)-order polynomial of X is: Among them, u i is the new coefficient of the polynomial obtained by adding several items, denoted by u i The bit width is bits, and c i is a coefficient and is 2M bits; n is a constant; is M bits; is a coefficient and is M bits, k is the highest-order coefficient.
7. A fast modular multiplication system based on a transformation basis according to claim 6, characterized in that: The sub-iterative reduction layer includes a mapping wiring layer and an accumulation layer, and the accumulation layer includes a plurality of adders; The mapping wiring layer is used to perform mapping wiring processing on the (n-1)th degree polynomial of X; The accumulation layer is used to perform accumulation processing on the X-order polynomial that has completed one mapping and routing processing.
8. A fast modular multiplication system based on a transformation basis according to claim 7, characterized in that: The iterative reduction layer is also used to: The (n-1)-order polynomial of X that has completed the reorganization and reduction process is sequentially subjected to several mapping, wiring and accumulation processes through a built-in adder; When the bit widths of the coefficients of the (n-1)th degree polynomial of X are reduced to a required range, a third target polynomial is output.
9. A fast modular multiplication system based on a transformation basis according to claim 8, characterized in that: The iterative reduction layer is also used to: Convert the (n-1)-degree polynomial of X that has completed the reorganization reduction process into u i A polynomial of the form, the u i The polynomial of the form is The u i The polynomial of the form is repeatedly substituted into the initial polynomial of the modulus q and the (n-1)-order polynomial of X for simplification to obtain the third target polynomial, which is: Among them, z i is the coefficient.
10. A fast modular multiplication system based on a transformation basis according to claim 9, characterized in that: The binary reduction layer is also used for: Performing several addition and shift operations on the third target polynomial to obtain the actual modular multiplication result.
Citation Information
Patent Citations
System and method for efficient basis conversion
CA2265389A1
Polynomial modular multiplication coprocessor based on lattice-based cryptosystem
CN104065478A
Montgomery modular multiplication based Tate pairing algorithm and hardware structure therefor
CN105068784A
High-speed modular multiplier based on post-quantum cryptography of homologous curve and modular multiplication method of high-speed modular multiplier
CN110908635A
Fast modular multiplication system based on transformation base
CN117742663A
Cited By
Modular multiplier based on lookup table and modular multiplication operation method
CN120723202A