Method and device for determining processing parameters for modulo multiplication circuit and rounding multiplication circuit
By dividing the multiplier into low- and high-bit segments, the transformation matrix of Karatsuba algorithm optimizes the modular and rounding multiplication circuit, the problem of high mode multiplication operation is solved and more efficient encryption calculation is achieved.
Patent Information
- Application Number
- CN202510442775.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the modular multiplication operation has the problem of high calculation cost in encryption calculation, especially in the implementation of division, which affects the efficiency of encryption calculation.
By dividing the multiplier into low-bit segments and high-bit segments, the transform matrix of Karatsuba algorithm optimizes the modulus multiplication and rounding multiplication circuit, and the low-bit evaluation matrix and the high-bit evaluation matrix are used to process the fragment parts of the multiplier respectively, and the modulus and rounding operations are optimized.
The calculation cost of the modulus multiplication and rounding multiplication circuit is reduced, the efficiency of encryption calculation is improved, and the area and calculation cost of the hardware circuit are reduced.
Smart Images

Figure CN120335764A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of hardware design for cryptographic computing, and particularly to methods and devices for determining processing parameters for a rounding multiplication circuit and a modulo multiplication circuit. Background Art
[0002] With the continuous improvement of people's awareness of data security and privacy protection, privacy computing, as a new type of secure computing mode, is gradually becoming one of the mainstream methods for data processing and analysis. Privacy computing realizes the protection of data privacy and security by directly computing on encrypted or anonymized data without exposing the original data, and has broad application prospects. At present, privacy computing technology has been widely applied in fields such as artificial intelligence, finance, and healthcare, and has become an important supporting technology for the future digital society.
[0003] Privacy computing relies on the use of various encryption algorithms, including RSA encryption algorithm, elliptic curve ECC encryption algorithm, homomorphic encryption algorithm, and so on. In particular, the homomorphic encryption algorithm has become one of the mainstream technologies in the field of privacy computing, and can perform various common computing operations in the encrypted state while ensuring the correctness of the result and the privacy of the data. Among the above various encryption algorithms, arithmetic operations are often carried out in a certain modulo space, and modulo multiplication operation, as a basic operation and primitive, has an important impact on the performance of cryptographic computing. At present, some hardware solutions have been proposed to accelerate the performance of modulo multiplication operation.
[0004] It is hoped that there will be an improved solution to further optimize the implementation process of basic arithmetic operations related to cryptographic computing. Summary of the Invention
[0005] One or more embodiments of this specification describe a method and a device for determining processing parameters for a modulo multiplication circuit and a rounding multiplication circuit, which can save the cost of parameter design and configuration for the modulo multiplication circuit and the rounding multiplication circuit.
[0006] According to a first aspect, there is provided a method for determining processing parameters of a modulo multiplication circuit, including:
[0007] For a target number k of items for dividing a multiplier into multiplier segments, divide it into the sum of a first number p and a second number q, where the second number q corresponds to the low-order segment part;
[0008] Obtain a q-order full-amount evaluation matrix corresponding to the second number q;
[0009] Determine an additional matrix, which is used to calculate the cross product between the first segment part corresponding to the first number p of two multipliers and the second segment part corresponding to the second number q;
[0010] Determine a k - order low - order evaluation matrix corresponding to the target number of terms k according to the q - order full - quantity evaluation matrix and the additional matrix; an s - order low - order / full - quantity evaluation matrix corresponding to any number of terms s is used to apply to the multiplier vectors composed of s multiplier segments of two multipliers to obtain two intermediate vectors, and the combination of the elements of the two intermediate vectors is used to obtain the low - order / full - quantity segment of the product; wherein, the k - order low - order evaluation matrix is used as the processing parameter of the modulo - multiplication circuit for taking the modulo of r to the k - th power.
[0011] In one embodiment, the second number of terms q is an integer greater than k / 2.
[0012] According to the first implementation manner, determining the additional matrix specifically includes: determining a first matrix, which is used to calculate the p - order modular multiplication of the first segment part of the first multiplier and the second segment part of the second multiplier; determining a second matrix, which is used to calculate the p - order modular multiplication of the second segment part of the first multiplier and the first segment part of the second multiplier.
[0013] Further, in one example, q all - zero columns can be added to the right of the p - order low - order evaluation matrix corresponding to the first number of terms p as the first matrix; q all - zero columns can be added to the left of the p - order low - order evaluation matrix as the second matrix.
[0014] According to an embodiment in the first implementation manner, determining the k - order low - order evaluation matrix corresponding to the target number of terms k specifically includes: sequentially stacking the basic matrix, the first matrix, and the second matrix to obtain a first evaluation matrix acting on the first multiplier; the basic matrix is an augmented matrix of the q - order full - quantity evaluation matrix; sequentially stacking the basic matrix, the second matrix, and the first matrix to obtain a second evaluation matrix acting on the second multiplier.
[0015] According to the second implementation manner, determining the additional matrix specifically includes: enumerating the segment indices of the multiplier segments that meet the target conditions, and filling the matrix elements according to the segment indices to obtain the additional matrix, where the target conditions include that if a row has a single segment index, the segment index is greater than or equal to q; if a row has a combination of two multiplier segments, the first segment index i in the combination is less than the second segment index j, the sum of the first segment index i and the second segment index j is less than or equal to k - 1, and the second segment index is greater than or equal to q.
[0016] According to an embodiment in the second implementation manner, determining the k - order low - order evaluation matrix corresponding to the target number of terms k specifically includes: sequentially stacking the basic matrix and the above - mentioned additional matrix to obtain the k - order low - order evaluation matrix, where the basic matrix is an augmented matrix of the q - order full - quantity evaluation matrix.
[0017] In a possible implementation, the above method further includes: determining a first cost value according to the effective multiplication times of modulo multiplication operation using the p-term Karatsuba algorithm; and determining a second cost value based on the algebraic operation of p; if the first cost value is less than the second cost value, then adopt the first implementation manner, otherwise, adopt the second implementation manner.
[0018] According to a second aspect, there is provided a method for processing parameters of a rounding multiplication circuit, including:
[0019] For a target number of terms k used to divide a multiplier into multiplier segments, divide it into the sum of a first number of terms p and a second number of terms q, where the first number of terms p corresponds to the high-order segment part;
[0020] Obtain a p-order full evaluation matrix corresponding to the first number of terms p;
[0021] Determine an additional matrix, which is used to calculate the cross product between the first segment part corresponding to the first number of terms p of two multipliers and the second segment part corresponding to the second number of terms q;
[0022] According to the p-order full evaluation matrix and the additional matrix, determine a k-order high-order evaluation matrix corresponding to the target number of terms k; an s-order high-order / full evaluation matrix corresponding to any number of terms s is used to be applied to multiplier vectors composed of s multiplier segments of two multipliers respectively to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the high-order / full segment of the product; wherein, the k-order low-order evaluation matrix is used as the processing parameter of the rounding multiplication circuit for rounding to the kth power of r.
[0023] According to a third aspect, there is provided a device for determining processing parameters of a modulo multiplication circuit, including:
[0024] A first division unit configured to divide a target number of terms k used to divide a multiplier into multiplier segments into the sum of a first number of terms p and a second number of terms q, where the second number of terms q corresponds to the low-order segment part;
[0025] A first obtaining unit configured to obtain a q-order full evaluation matrix corresponding to the second number of terms q;
[0026] A first determining unit configured to determine an additional matrix, which is used to calculate the cross product between the first segment part corresponding to the first number of terms p of two multipliers and the second segment part corresponding to the second number of terms q;
[0027] A first calculation unit, configured to determine a k-order low-order evaluation matrix corresponding to a target number of terms k according to the q-order full-scale evaluation matrix and the additional matrix; an s-order low-order / full-scale evaluation matrix corresponding to any number of terms s, which is used to be applied to multiplier vectors formed by s multiplier segments of two multipliers respectively to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain a low-order / full-scale segment of the product; wherein, the k-order low-order evaluation matrix is used as a processing parameter of a modulo multiplication circuit for taking the modulo of r to the kth power.
[0028] According to a fourth aspect, there is provided an apparatus for determining processing parameters of a rounding multiplication circuit, including:
[0029] A second partitioning unit, configured to partition a target number of terms k for partitioning a multiplier into multiplier segments into a sum of a first number of terms p and a second number of terms q, where the first number of terms p corresponds to a high-order segment part;
[0030] A second obtaining unit, configured to obtain a p-order full-scale evaluation matrix corresponding to the first number of terms p;
[0031] A second determining unit, configured to determine an additional matrix, which is used to calculate a cross product between a first segment part corresponding to the first number of terms p of two multipliers and a second segment part corresponding to the second number of terms q;
[0032] A second calculation unit, configured to determine a k-order high-order evaluation matrix corresponding to a target number of terms k according to the p-order full-scale evaluation matrix and the additional matrix; an s-order high-order / full-scale evaluation matrix corresponding to any number of terms s, which is used to be applied to multiplier vectors formed by s multiplier segments of two multipliers respectively to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain a high-order / full-scale segment of the product; wherein, the k-order low-order evaluation matrix is used as a processing parameter of a rounding multiplication circuit for taking the integer part of r to the kth power.
[0033] According to a fifth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in the first aspect or the second aspect.
[0034] According to a sixth aspect, there is provided a computing device, including a memory and a processor, characterized in that an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect or the second aspect is implemented.
[0035] In the embodiments of this specification, for a target order to be processed, it is decomposed into a high-order segment part and a low-order segment part of a lower order. Based on the full evaluation matrix of the lower order and an additional matrix used to calculate the cross product of the high-order part and the low-order part of two multipliers, a transformation matrix of the high-order target order is quickly obtained as the processing parameter of the modulo multiplication circuit / rounding multiplication circuit. Further, strategies for independently processing two multipliers and a merging processing strategy are proposed to determine the additional matrix and the final target order transformation matrix. Furthermore, in some examples, the strategy selection can also be considered based on the calculation cost, so as to efficiently obtain a high-order transformation matrix with a lower calculation cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 Illustrates the algorithm process of modular multiplication using Barrett reduction;
[0038] Figure 2 Illustrates the schematic diagram of the principle of the 2-term Schoolbook algorithm;
[0039] Figure 3 Illustrates the digit diagram in the operation;
[0040] Figure 4 Illustrates the schematic diagram of a modular multiplication circuit according to an embodiment;
[0041] Figure 5 Illustrates the flowchart of the method for determining the processing parameter of the modulo multiplication circuit according to an embodiment;
[0042] Figure 6 Illustrates the method for determining the k-order low-order evaluation matrix according to an embodiment;
[0043] Figure 7 Illustrates the flowchart of the method for determining the processing parameter of the rounding multiplication circuit according to an embodiment;
[0044] Figure 8 Illustrates the schematic diagram of the structure of the modulo multiplication parameter determination device according to an embodiment;
[0045] Figure 9 Illustrates the schematic diagram of the structure of the rounding multiplication parameter determination device according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The solution provided in this specification will be described below with reference to the accompanying drawings.
[0047] To achieve privacy computing, various encryption algorithms are adopted. In particular, homomorphic encryption has become one of the mainstream technologies in the field of privacy computing because it can perform corresponding mathematical calculations on data in an encrypted state. The operations of homomorphic encryption involve a large number of four arithmetic operations under modular arithmetic. The main difficulty of modular arithmetic lies in that division calculation is involved in the process of calculating the quotient value, and the implementation cost of division is relatively high. For this reason, the Barrett reduction algorithm is proposed. This algorithm achieves an effect similar to division through multiplication and shift operations, thereby improving the efficiency of modular arithmetic.
[0048] Modular multiplication is a commonly used operation in encryption systems, that is, after multiplying two integers X and Y, a modular operation is performed on a modulus M, that is, X * Y mod M. Modular multiplication can be implemented through a hardware circuit using the Barrett reduction idea.
[0049] Figure 1 The algorithm process of modular multiplication using Barrett reduction is shown. As Figure 1 shown, the inputs of this algorithm include the multiplicands X, Y, the modulus M, and the pre-computed value μ corresponding to the modulus M. Here, it is assumed that both the multiplicands X and Y are not greater than the modulus M corresponding to the modular space. And the value range of this modulus M can be expressed as [2 β-1 , 2 N , where β is a very small constant, for example, 2; N is the maximum bit width supported by the system. In this way, the modulus M can take various values within the acceptable bit width N of the system. The parameters α and t involved in the pre-computed value can be set as needed. Specifically, α and β can be appropriately set so that the Z generated in the 4th line of the algorithm is within the range of [0, 2M), so that in lines 5 - 6, at most one correction operation is required.
[0050] It can be seen that in the algorithm process of modular multiplication using Barrett reduction, the following several operations are relatively high-cost core operations: (1) Conventional multiplication calculation, such as the calculation X * Y shown in the 1st line of the algorithm; (2) Rounding multiplication calculation, that is, calculating the truncated rounding result of the product of two numbers multiplied relative to the target value in the form of a power of 2, such as the calculation shown in the 3rd line of the algorithm where represents rounding down; (3) Modular multiplication calculation, that is, calculating the modular result of the product of two numbers multiplied relative to the target value in the form of a power of 2, such as the calculation involved in the 4th line of multiplication qM mod 2 N+1 .
[0051] Without loss of generality, conventional multiplication can be represented as AB, and rounding multiplication can be represented as Express the modular multiplication as AB mod r k , where r = 2 m , and where represents the ceiling function. That is, the target value in the form of a power of 2 (when applied to the Figure 1 algorithm, here N may be different from the bit width corresponding to the modulus M) is divided into k segments, each segment with a bit width of m and a base of r for each segment.
[0052] To accelerate the modular multiplication operation, the operation processes and calculation circuits of the above-mentioned conventional multiplication, ceiling multiplication, and modular multiplication can be optimized and improved respectively.
[0053] For the conventional multiplication, the Karatsuba algorithm can be used for acceleration. According to the k-term Karatsuba algorithm, first, the multiplier A (and multiplier B) with a bit width of N is divided into k segments, or called k words. The bit width of each segment A i is Thus, the multiplier A can be expressed as where r = 2 m is the base for each segment,
[0054] Thus, the product result C of AB can be calculated by the following formula:
[0055]
[0056] Taking k = 2 as an example, the process of calculating the product C using the Schoolbook algorithm is shown in the following formula:
[0057] C = AB = A1B1r 2 +(A1B0 + A0B1)r + A0B0 (2)
[0058] Figure 2 Fig. shows the schematic diagram of the principle of the 2-term Schoolbook algorithm.
[0059] It can be found that if strictly following the calculation process of formula (2) and calculating A i B j and then summing them up, 4 multiplication operations are required. However, for the second term in formula (2), what matters is the product sum A1B0 + A0B1, rather than the individual products. Therefore, the following formula can be used to recover the product sum:
[0060] A0B1 + A1B0 = -(A0 - A1)(B0 - B1) + A1B1 + A0B0 (3)
[0061] Since the result of A1B1 and A0B0 can be reused in formula (3), by calculating the second term in formula (2) using formula (3), the number of multiplication operations required in the calculation process of formula (2) can be reduced to 3 times, thus realizing the optimization of the Karatsuba algorithm.
[0062] Still taking k = 2 as an example below, the optimization process of the Karatsuba algorithm is divided into three stages: the evaluation stage, the multiplication stage, and the interpolation stage. In the preparatory process, the multiplier can be represented as a vector composed of each multiplier fragment (or word), called the multiplier vector. For example, in the case of k = 2, the two multipliers can be respectively represented as the multiplier vectors: A = [A0, A1] T , B = [B0, B1] T .
[0063] In the evaluation stage, the evaluation matrix E is used to act on the multiplier vector to obtain the corresponding generated vector. The elements in the generated vector are combinations of multiplier fragments and are used as the basis for subsequent multiplications. For example, when k = 2, the second-order evaluation matrix E (2) can take the following form:
[0064]
[0065] Acting this evaluation matrix on the multiplier vector corresponding to multiplier A, the generated vector e corresponding to multiplier A can be obtained as follows A :
[0066]
[0067] Similarly, the generated vector corresponding to multiplier B can also be obtained.
[0068] Then, in the multiplication stage, the generated vector e A corresponding to multiplier A and the generated vector e B corresponding to multiplier B are multiplied bit by bit to obtain the intermediate vector, that is: e = e A ⊙ e B , where ⊙ is the Hadamard operator, indicating bit-by-bit multiplication. In the case of k = 2, the obtained intermediate vector e is as follows:
[0069]
[0070] Next, in the interpolation stage, the interpolation matrix I is multiplied by the intermediate vector, and the resulting vector shows each fragment of the product result. In the case of k = 2, the second-order interpolation matrix I (2) can take the following form:
[0071]
[0072] In this way, the obtained result vector is as follows:
[0073]
[0074] The radix vector R = [r 0 , r 1 , r 2 T can act on the above result vector, that is, calculate R T ·C, so as to recover the value of the product result C.
[0075] Combining the above three stages, the product calculation process under the k-term Karatsuba algorithm can be summarized as follows:
[0076] C = Karatsuba(A, B) = R T ·I (k) ·((E (k) ·A) ☉ (E (k) ·B)) (9)
[0077] According to formula (9), pre-construct the transformation matrix under the k-term Karatsuba algorithm, including the evaluation matrix E (k) and the interpolation matrix I (k) . Use the evaluation matrix E (k) to act on the vector composed of the multiplier segments of the multiplier A and the vector composed of the multiplier segments of the multiplier B respectively, multiply the two obtained vectors bit by bit to obtain an intermediate vector; then multiply the intermediate vector by the interpolation matrix I (k) . The resulting result vector C then shows each segment of the product.
[0078] In practice, the construction of the evaluation matrix E (k) and the interpolation matrix I (k) is not unique, and different matrix forms may bring different computational costs, such as different numbers of multiplications.
[0079] The above optimization idea based on the k-term Karatsuba algorithm can also be applied to rounding multiplication and modular multiplication AB mod r k .
[0080] It can be understood that according to formula (1), under the k-term Karatsuba algorithm, the product C of AB contains segments from C0 to C 2k-2 . The segment subscripts or indices correspond to the orders of r. For modular multiplication AB mod r k , the segments of the product C with orders higher than k have no influence on the result, so only the product segments of the lower-order part within the k-th order need to be considered. For this purpose, the use of the transformation matrix (E (k) , I (k))Idea for implementing the k-term Karatsuba algorithm and constructing the low-order transformation matrix Through a similar operation, the low-order part C in the product result is obtained L . Specifically, the k-term Karatsuba low-order algorithm for modular multiplication can be implemented as follows:
[0081]
[0082]
[0083] where R L = [r 0 , r 1 , …, r k-1 T .
[0084] Then take the modulus of C L with r k . It can be seen that the process of obtaining the low-order product part using the low-order transformation matrix is similar to the process of implementing conventional multiplication: using the low-order evaluation matrix to act on the vectors composed of the multiplier segments of multiplier A and the vectors composed of the multiplier segments of multiplier B respectively, multiplying the two obtained vectors bit by bit to get the intermediate vector e L ; then multiplying this intermediate vector e by the low-order interpolation matrix L , and the resulting result vector C L shows the low-order part of the product.
[0085] Still taking k = 2 as an example to describe an example process. In this example, the low-order evaluation matrix and the low-order interpolation matrix can take the following forms respectively:
[0086]
[0087] Thus, the obtained result is:
[0088]
[0089] The in the form of formula (11) can further simplify the circuit calculation because when calculating the modulus of C L with r 2 later, only the last m bits in the terms of e L,1 and e L,2 need to be considered, that is, only e L,1 mod r and e L,2 mod r need to be considered. Its principle can be shown by the digit diagram of Figure 3 .
[0090] It is understood that r = 2 m is the radix or size of a segment, and m is the bit width of a segment. For any e L,i is obtained by multiplying the additive combination of multiplier segment A i by the additive combination of multiplier segment B j . The maximum bit width of each additive combination is m + 1 (carry may occur during addition). Therefore, the maximum bit width of e L,i is 2m + 2, slightly exceeding the size of two segments. As shown in Figure 3 , e L,i r is obtained by shifting e L,i to the left by one segment. When calculating C L mod r 2 , the part exceeding the size of r 2 has no effect on the result. As shown by the shaded part in the figure, only the white part falling within the range of r 2 has an effect on the result. It can be seen that only the last m bits of e L,i r fall within the range of r 2 .
[0091] Therefore, the modulo operation of C L can be calculated as follows:
[0092] C L mod r 2 = e L,0 (r + 1)+(e L,1 mod r)r-(e L,2 mod r)r mod r 2 (13)
[0093] Since only the last m bits of e L,1 and e L,2 need to be considered, their calculation can be implemented by a half - multiplication circuit. The half - multiplication circuit includes a high - order half - multiplication circuit that only calculates the high - order part and a low - order half - multiplication circuit that only calculates the low - order part. In formula (13), the low - order half - multiplication circuit can be used to calculate e L,1 and e L,2 , obtaining e L,1 mod r and e L,2 mod r. Compared with the full - multiplication circuit that calculates all bits, the half - multiplication circuit can save circuit area and calculation cost. The calculation cost of the half - multiplication circuit can be considered as 0.5 times the full - multiplication calculation. Correspondingly, calculating the modulo multiplication through formula (13) only requires the calculation cost of 1 + 2 * 0.5 = 2 full - multiplication operations. Compared with the 3 multiplications of the 2 - term Karatsuba algorithm for implementing conventional multiplication, the cost is further reduced.
[0094] For simple and efficient implementation of the operation, preferably, any element in the low-order transformation matrix is selected from 0, 1, and -1.
[0095] In contrast to modular multiplication, for integer multiplication , the segments in the product C with orders lower than k have no influence on the result. Therefore, only the product segments of the high-order part with orders above k need to be considered. For this purpose, similarly referring to the idea of using the transformation matrix (E (k) , I (k) ) to implement the k-term Karatsuba algorithm, a high-order transformation matrix is constructed Through similar operations, the high-order part in the product result is obtained, that is, the high-order product result C U . Specifically, the k-term Karatsuba high-order algorithm for modular multiplication can be implemented as follows:
[0096]
[0097] where R U = [r k-1 , r k , …, r 2k-2 T .
[0098] Then, calculate C U / r k by shifting. It can be seen that the process of obtaining the high-order product result using the high-order transformation matrix is similar to the process of implementing conventional multiplication: using the high-order evaluation matrix to act on the vector composed of the multiplier segments of the multiplier A and the vector composed of the multiplier segments of the multiplier B respectively, multiplying the two obtained vectors bit by bit to get the intermediate vector e U ; then multiplying this intermediate vector e by the high-order interpolation matrix U . The resulting result vector C U shows the high-order part of the product.
[0099] Similar to optimizing modular multiplication using an appropriate form of , the process of integer multiplication can also be optimized by using an appropriate form of the high-order transformation matrix , including using a high-order half-multiplication circuit to calculate partial product terms, thereby reducing the calculation cost. And, similar to the low-order transformation matrix, preferably, any element in the high-order transformation matrix is selected from 0, 1, and -1.
[0100] Figure 4 shows a schematic diagram of a modular multiplication circuit according to an embodiment. This modular multiplication circuit uses the idea of the k-term Karatsuba algorithm to implement Figure 1 The calculation process of modular multiplication through Barrett reduction. As Figure 4 shown, the modular multiplication circuit includes a circuit 11 for conventional multiplication, a circuit 12 for truncated multiplication, and a circuit 13 for modular multiplication. Specifically, the inputs of circuit 11 include multiplicands X and Y. In this circuit, with the evaluation matrix E (k) as the circuit parameter, addition operations are respectively performed on the multiplier segments corresponding to multiplicand X and multiplicand Y to obtain the generated vectors e A and e B for each element in. After compression and combination of the generated vectors, with the interpolation matrix I (k) as the parameter, the result T of XY multiplication as shown in the Figure 1 algorithm can be obtained.
[0101] In the following text, for the sake of distinction, the transformation matrices (E (k) , I (k) ) in the k-term Karatsuba algorithm for implementing conventional multiplication are also called the full-scale evaluation matrix and the full-scale interpolation matrix.
[0102] Circuit 12 implements truncated multiplication using the KaratsubaUpper algorithm shown in formula (14). The inputs of circuit 12 include the pre-computed value μ and T obtained by shifting T H (as shown in line 2 of the Figure 1 algorithm). In this circuit 12, with the aforementioned high-order evaluation matrix as the circuit parameter, addition operations and corresponding encoding operations are respectively performed on the multiplier segments of the two inputs. After compression and combination of the generated vectors, with the high-order interpolation matrix as the parameter for combination and compression, the result of truncated multiplication shown in line 3 of the algorithm
[0103] is obtained. Circuit 13 implements modular multiplication using the KaratsubaLower algorithm shown in formula (10). The inputs of circuit 13 include the modulus M and q output by circuit 12. In this circuit 13, with the aforementioned low-order evaluation matrix as the circuit parameter, addition operations and corresponding encoding operations are respectively performed on the multiplier segments of the two inputs. After compression and combination of the generated vectors, with the low-order interpolation matrix N+1 as the parameter for combination and compression, the modular multiplication result qM mod 2
[0104] can be obtained. L (which can be achieved by truncation or shifting), through basic circuit operations such as shifting and addition, the final modular multiplication result Z can be obtained.
[0105] In other embodiments, circuits 12 and 13 may also be independent circuits dedicated to performing rounding multiplication and modulo multiplication.
[0106] It can be understood that the operation of circuit 12 depends on the circuit configuration with the high-order transformation matrix as a circuit parameter, and the operation of circuit 13 depends on the circuit configuration with the low-order transformation matrix as a circuit parameter. For the same order k, the forms of the transformation matrices satisfying the forms of formulas (10) and (14) are not unique. Different forms of transformation matrices can optimize the circuit calculation process to varying degrees. A relatively optimal form of the transformation matrix, such as the form of formula (11) can adopt more half-multiplication circuits to optimize the circuit calculation and reduce the calculation cost.
[0107] Therefore, exploring the available transformation matrices for various numbers of terms k while minimizing the calculation cost as much as possible becomes a direction for optimizing the modulo multiplication / rounding multiplication circuit calculation. When k takes a small value, the dimension of the transformation matrix is limited and the search space is not large; while when k gradually increases, the search space increases exponentially, bringing great difficulties to the determination of the transformation matrix and the configuration of the circuit. Therefore, it is hoped that there can be an improved scheme to quickly determine the low-order transformation matrix / high-order transformation matrix with low calculation cost for a given number of terms k,
[0108] Thereby optimizing the configuration of the modulo multiplication circuit / rounding multiplication circuit.
[0109] As described previously in combination with formulas (11)-(13) and Figure 3 , it can be found that when the last row element of the i-th column in the low-order interpolation matrix is non-zero and the other rows are all zero, then this column will generate the target term e L r L,i r k -1 in the low-order result C L . This target term is equivalent to e k mod r in the subsequent modulo operation C L,i mod r, that is, only the last m bits need to be considered. Therefore, this target term can be implemented by a half-multiplication circuit.
[0110] Therefore, the number of columns in the low-order interpolation matrix that meet the above conditions (only the last row element is non-zero) corresponds to the number of half-multiplication calculations, and the number of the remaining columns corresponds to the number of conventional full-multiplication calculations.
[0111] Correspondingly and similarly, if only the element in the first row of the i-th column in the high-order interpolation matrix is not 0, then this column will generate a target term e U in the high-order result C U,i r k-1 . This target term is equivalent to e U / r in the subsequent rounding and shifting operation C k , that is, only the first m bits need to be considered. Therefore, this target term can be implemented by a semi-multiplication circuit. That is, the number of columns in the high-order interpolation matrix U,i that meet the condition (only the element in the first row is not 0) corresponds to the number of semi-multiplication calculations, and the number of remaining columns corresponds to the number of conventional full-multiplication calculations. In addition, it is found that the low-order transformation matrix
[0112] and the high-order transformation matrix and the high-order transformation matrix can have a certain associated correspondence relationship, and the low-order transformation matrix and the high-order transformation matrix with the associated correspondence relationship consume the same number of semi-multiplication times and full-multiplication times.
[0113] For clarity, the following defines the cost-related concepts in the multiplication circuit calculation implemented based on the Karatsuba algorithm.
[0114] T (k) represents the number of full-multiplication calculations required for performing a conventional multiplication (calculating AB) using the k-term Karatsuba algorithm. For the conventional multiplication AB, all operations between multiplier segments need to be implemented using a full-multiplication circuit.
[0115] F (k) represents the number of full-multiplication calculations required for performing a rounding multiplication operation or a modular multiplication operation AB mod r k using the k-term Karatsuba algorithm.
[0116] H (k) represents the number of semi-multiplication calculations required for performing a rounding multiplication operation or a modular multiplication operation AB mod r k using the k-term Karatsuba algorithm.
[0117] Q (k) represents the number of effective multiplication calculations required for performing a rounding multiplication operation or a modular multiplication operation AB mod r k using the k-term Karatsuba algorithm.
[0118] Since the cost of semi-multiplication can be considered as 1 / 2 of that of full multiplication, the effective number of multiplication operations Q (k) can be expressed as:
[0119] Q (k) = F (k) + H (k) / 2 (15)
[0120] Therefore, the goal that the solution in this specification hopes to achieve is to quickly determine the low-order transformation matrix / high-order transformation matrix Furthermore, make the effective number of multiplication operations Q (k) used by the low-order / high-order transformation matrix as small as possible.
[0121] The following will be described in detail in combination with the low-order evaluation matrix After determining the low-order evaluation matrix , the low-order interpolation matrix can be correspondingly determined. The determination method of the high-order transformation matrix is similar.
[0122] As mentioned above, for modular multiplication AB mod r k it is necessary to calculate [C0, C1,..., C k-1 T , where the cross-term combinations in the Karatsuba algorithm can be optimized by using the two-element enumeration method, that is, the following rewrite is performed:
[0123] A i B j + A j B i = -(A i - A j )(B i - B j ) + A i B i + A j B j (16)
[0124] Therefore, the k-order low-order evaluation matrix can be constructed according to the following idea so that the pairwise multiplication of the two intermediate vectors obtained after it acts on two multiplier vectors can achieve the following two parts of multiplication:
[0125] Part (i): Calculate A0B0,..., A k-1 B k-1 , that is, each multiplier of the multiplication comes from a segment or element of the multipliers A and B;
[0126] Part (ii): Calculate (A i - A j )(B i - B j ), that is, each multiplier of the multiplication comes from the combination of two segments of the multiplicands A and B, and the indices of these two segments need to meet certain conditions, namely i < j, and i + j ≤ k - 1.
[0127] To implement the multiplication in part (i), k rows can be constructed, where the (i + 1)-th element of the i-th row is 1 (where i ranges from 0 to k - 1), and the rest of the elements are 0.
[0128] To implement the multiplication in part (ii), i and j that meet the above conditions can be traversed from 0 to k - 1 respectively to construct the following rows:
[0129]
[0130] Combining the rows constructed in the above two parts together, the k-order lower-order evaluation matrix can be obtained.
[0131] Taking k = 4 as an example, according to the above method, the following 4-order lower-order evaluation matrix can be constructed.
[0132]
[0133] As can be seen from equation (18), the upper 4 rows (from row 0 to row 3) of the matrix are used for the multiplication calculation in part (i), and the lower 4 rows (from row 4 to row 7) are used for the multiplication calculation in part (ii), and then for calculating the cross terms shown in formula (16). For example, row 4 is used to calculate A0B3 + A3B0.
[0134] The above two-element enumeration method can be applied to the search for low-order evaluation matrices with a small k value. When k is large, the embodiments of this specification propose a method for efficiently determining the k-order lower-order evaluation matrix by decomposing k.
[0135] Figure 5 The flowchart of the method for determining the processing parameters of the modulo multiplication circuit according to an embodiment is shown. This method can be executed by any circuit, device, and platform with computing and processing capabilities. As Figure 5 shown, this method includes the following steps.
[0136] In step 52, decompose the target number of items k to be processed, and divide it into the sum of the first number of items p and the second number of items q, where the second number of items q corresponds to the low-order segment part.
[0137] For the k-item Karatsuba algorithm, the multiplicand A will be divided into k segments A0,..., Ak-1 For the convenience of calculating modular multiplication, the k segments can be further divided into a high - order segment part A with p terms H , and a low - order segment part A with q terms L , where the value of q can be That is, the number of terms q of the low - order segment part should be greater than or equal to the number of terms p of the high - order segment part. When q takes the value of , the above A H and A L are as follows:
[0138]
[0139] Another multiplier B can also be correspondingly divided. In this way, the product AB can be expressed as:
[0140]
[0141] Furthermore, the modular multiplication AB mod r k can be expressed as:
[0142]
[0143] Thus, the modular multiplication can be decomposed into two parts. The first part is A L B L mod r k , where both A L and B L contain q multiplier segments. Therefore, A L B L can be optimized by a q - order full - scale transformation matrix (E (q) , I (q) ). The second part includes two sub - parts, that is, calculating A H B L mod r p and A L B H mod r p , both of which involve taking the modulus of r for the cross - product calculation between the high - order segment part and the low - order segment part of A and B p .
[0144] Still taking k = 4 as an example. We can take p = k - q = 1, and thus divide the multiplier as follows:
[0145] A H = A3r 0 , A L = A0r 0 + A1r 1 + A2r 2
[0146] B H = B3r 0 ,B L = B0r 0 + B1r 1 + B2r 2 (22)
[0147] When calculating AB mod r 4 there is:
[0148] A H B L mod r 1 = A3B0 mod r
[0149] A L B H mod r 1 = A0B3 mod r (23)
[0150] The two terms in formula (23) correspond to the calculation of the second part.
[0151] Based on the above decomposition concept, perform the following steps 54 to 58.
[0152] In step 54, obtain the full-order evaluation matrix of order q corresponding to the second number q. It can be understood that this full-order evaluation matrix E (q) is used to optimize the calculation of the first part A L B L . In this step, the full-order evaluation matrix of order q determined in advance can be read, or the full-order evaluation matrix of order q can be determined by various search calculations. Specifically, in practice, since q is only slightly larger than k / 2 and the order is generally low, the full-order evaluation matrix of the corresponding order can be determined by enumeration search. Or, it can also be determined by other methods for optimizing the full-order evaluation matrix, which is not limited here.
[0153] In step 56, determine the additional matrix, which is used to calculate the cross product between the first segment part (i.e., the high-order segment part) corresponding to the first number p of the two multipliers and the second segment part (i.e., the low-order segment part) corresponding to the second number q. It can be understood that this additional matrix is used to optimize the calculation of the second part, that is, to optimize A H B L mod r p and A L B H mod r p .
[0154] Then, in step 58, according to the full-order evaluation matrix E of order q obtained above (q)Based on the additional matrix, determine the k-th order low-order evaluation matrix corresponding to the target number of terms k. As described above, the s-th order low-order / full-scale evaluation matrix corresponding to any number of terms s is used to apply to the multiplier vectors composed of s multiplier segments of the two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the low-order / full-scale segments of the product.
[0155] Observing the calculation of the second part shown in formula (21), it can be seen that this part of the calculation further includes two sub-parts, namely, calculating A H B L mod r p and A L B H mod r p . Accordingly, for the determination of the additional matrix and the k-th order low-order evaluation matrix, it can be achieved through two strategies: The first strategy is to independently process A H B L and A L B H respectively, and determine different evaluation matrices for the two multipliers; The second strategy is to combine the processing and determine a common evaluation matrix for the two multipliers. The specific implementations under the two strategies are described below respectively.
[0156] As a common basis for the two strategies, the q-th order full-scale evaluation matrix E (q) can be augmented after being obtained in step 54, and its augmented matrix is used as the basis matrix E base . Specifically, p all-zero columns are added to the right side of the q-th order full-scale evaluation matrix E (q) , and its number of columns is expanded to k columns to obtain the basis matrix, that is:
[0157] E base =[E (q) ,0] (24)
[0158] where [.,0] means adding all-zero columns to the right side for augmentation.
[0159] According to the first strategy, in step 56, two matrices are respectively determined as the additional matrix. Specifically, determine the first matrix E left , which is used to calculate the p-th order modular multiplication of the first segment part of the first multiplier and the second segment part of the second multiplier, that is, used to calculate A H B L mod r p ; In addition, the second matrix E right is also determined, which is used to calculate the p-th order modular multiplication of the second segment part of the first multiplier and the first segment part of the second multiplier, that is, used to calculate A L B H mod r p .
[0160] In one embodiment, the above first matrix and second matrix can be obtained respectively based on a p-order low-order evaluation matrix. More specifically, in one example, q all-zero columns can be added to the right side of the p-order low-order evaluation matrix as the first matrix E left , that is:
[0161]
[0162] Correspondingly, q all-zero columns can be added to the left side of the p-order low-order evaluation matrix as the second matrix E right , that is:
[0163]
[0164] where [0,.] means adding all-zero columns on the left side for augmentation.
[0165] It can be understood that the purpose of the above augmentation is to expand the matrix dimension to k columns.
[0166] After that, in step 58, the basic matrix E base , the above first matrix E left and the second matrix E right are stacked in sequence to obtain the first evaluation matrix acting on the first multiplier A that is:
[0167]
[0168] where vstack means vertical stacking or splicing.
[0169] Correspondingly, the basic matrix E base , the above second matrix E right and the first matrix E left are stacked in sequence to obtain the second evaluation matrix acting on the second multiplier B that is:
[0170]
[0171] In this way, the first evaluation matrix acting on the first multiplier A and the second evaluation matrix acting on the second multiplier B together serve as the k-order low-order evaluation matrix. In such a case, the k-order modular multiplication of the two multipliers A and B can be implemented as follows: the first evaluation matrix acts on the multiplier A, the second evaluation matrix acts on the multiplier B, and the two obtained vectors are multiplied bit by bit, that is, to obtain Apply the subsequent k - order low - order interpolation matrix again.
[0172] Based on the above process of obtaining the k - order low - order evaluation matrix and the above calculation process, the computational cost of performing modular multiplication using the k - order low - order evaluation matrix obtained under the first strategy can be analyzed as follows:
[0173] F (k) = T (q) + 2F (p)
[0174] H (k) = 2H (p)
[0175] Q (k) = T (q) + 2Q (p) (29)
[0176] Taking k = 4 as an example, describe the determination and computational cost of the low - order evaluation matrix under the first strategy. As mentioned before, in the example of k = 4, the multiplier can be decomposed as shown in Equation (22), where p = 1 and q = 3.
[0177] Correspondingly, based on the 3 - order full - scale evaluation matrix E (3) Add a column of all zeros on the right to obtain the following basic matrix E base :
[0178]
[0179] Since Therefore, we have:
[0180] E left = [1, 0, 0, 0]; E right = [0, 0, 0, 1] (31)
[0181] Thus, the first evaluation matrix and the second evaluation matrix under k = 4 are as follows:
[0182]
[0183] The resulting low - order evaluation matrix has a computational cost of: F (4) = 6, H (4) = 2.
[0184] According to the second strategy, the cross - multiplication terms of the high - order fragments and the low - order fragments of the two multipliers are combined and processed. When dealing with the cross - multiplication terms, refer to formula (16) for A i B j + A j B iExpand, and use the two-element enumeration method to determine the additional matrix. It can be understood that the expansion of the cross terms will involve A i B i , A j B j and other single-index terms. For such terms, only the basic matrix E base does not have the rows. Specifically, since E base already contains the rows corresponding to the low-order segment indexes of each q term, only the rows corresponding to the p high-order segment indexes need to be retained in the additional matrix. When traversing part (ii), in addition to i < j and i + j ≤ k - 1, the conditions for i and j also need to include
[0185] Specifically, according to the second strategy, in step 56, enumerate the segment indexes of the multiplier segments that meet the target conditions, and fill the matrix elements according to the segment indexes to obtain the additional matrix E head , where the target conditions include that if a row has a single segment index, the segment index is greater than or equal to q; if a row has a combination of two multiplier segments, the first segment index i in the combination is less than the second segment index j, the sum of the first segment index i and the second segment index j is less than or equal to k - 1, and the second segment index is greater than or equal to q.
[0186] Based on the additional matrix E head obtained in this way, in step 58, stack the basic matrix E base and the additional matrix E head in sequence to obtain the k-order low-order evaluation matrix That is:
[0187]
[0188] The k-order low-order evaluation matrix obtained in this way can act on two multipliers as shown in formula (10), which is convenient for the calculation of modular multiplication. The calculation cost of modular multiplication using the k-order low-order evaluation matrix obtained under the second strategy can be analyzed as follows:
[0189]
[0190] Still taking k = 4 as an example. As shown in formula (23), the cross terms to be calculated include A3B0 + A0B3, which can be expanded as:
[0191] A3B0 + A0B3 = -(A0 - A3)(B0 - B3) + A0B0 + A3B3 (35)
[0192] According to the aforementioned two-element enumeration process, it can be obtained that:
[0193]
[0194] The first row contains a single index for calculating A3B3 (since the row corresponding to A0B0 is already included in E base The second row contains the index combinations used to calculate (A0-A3) and (B0-B3).
[0195] Based on the above additional matrix, the following k=4-order low-order evaluation matrix can be obtained:
[0196]
[0197] In the above, the k-order low-order evaluation matrix is determined by the first strategy of independent processing and the second strategy of combined processing. It can be seen that the above method of determining the k-order low-order evaluation matrix can be applied to any value of the number of items k.
[0198] Furthermore, it can be seen that the calculation costs of the k-order low-order evaluation matrix obtained by different strategies are not the same. Therefore, in one embodiment, a strategy can be selected based on the calculation cost to obtain a low-order evaluation matrix with a lower calculation cost. Figure 6 Describe the strategy choices.
[0199] Figure 6 A method for determining a k-order low-order evaluation matrix according to an embodiment is shown. Figure 6 As shown, first, in step 62, the target number of items k is divided into the sum of the first number of items p and the second number of items q, wherein the second number of items q corresponds to the low-order fragment part. In step 64, the q-order full-quantity evaluation matrix corresponding to the second number of items q is obtained. Steps 62-64 are similar to Figure 5 Steps 52-54 in correspondence are the same.
[0200] Then, in step 65, the first cost value and the second cost value are calculated, wherein the first cost value Q1 is the calculation cost of the low-order evaluation matrix obtained by adopting the independent processing strategy, and the second cost value Q2 is the calculation cost of the low-order evaluation matrix obtained by adopting the combined processing strategy. According to the above formulas (29) and (34), we have:
[0201] Q1=T (q) +2Q (p)
[0202]
[0203] By comparison, it can be seen that the first term of the first cost value and the second cost value are the same. For the purpose of comparison, only the second term of each can be calculated. That is, the effective number of multiplications Q of the modular multiplication operation using the p-term Karatsuba algorithm can be calculated. (p), determine the first cost value; and, determine the second cost value based on the algebraic operation of p (specifically, perform an operation based on a constant after squaring p + 1).
[0204] Furthermore, in step 66, compare the first cost value and the second cost value. If the first cost value is less than the second cost value, then proceed to step 67, where the steps 56 and 58 Figure 5 are executed using the first strategy of independently processing the two multipliers respectively to obtain the k-order low-order evaluation matrix Otherwise, if the first cost value is not less than the second cost value, then proceed to step 68, and execute the steps 56 and 58 Figure 5 using the second strategy of combining the two multipliers to determine the k-order low-order evaluation matrix
[0205] Optionally, after executing step 67, the total multiplication times F (k) and the half multiplication times H (k) required under the k-order low-order evaluation matrix are determined according to the formula group (29). After executing step 68, the corresponding total multiplication times and half multiplication times are determined according to the formula group (34).
[0206] In the above manner, a k-order low-order evaluation matrix with a relatively low computational cost can be efficiently determined as the configuration parameter of the modular multiplication circuit.
[0207] As mentioned above, the high-order evaluation matrix applicable to the rounding multiplication circuit has a cost and computational logic similar to or corresponding to that of the low-order evaluation matrix . Therefore, the above technical concept can be extended to the determination of the high-order evaluation matrix. Figure 7 The flowchart shows a method for determining the processing parameters of a rounding multiplication circuit according to an embodiment. This method can be executed by any circuit, device, platform with computing and processing capabilities. As Figure 7 shown, the method includes the following steps.
[0208] In step 72, for the target number of items k used to divide the multiplier into multiplier segments, divide it into the sum of the first number of items p and the second number of items q, where the first number of items p corresponds to the high-order segment part;
[0209] In step 74, obtain the p-order full-scale evaluation matrix corresponding to the first number of items p;
[0210] In step 76, determine the additional matrix, which is used to calculate the cross product between the first segment part corresponding to the first number of items p of the two multipliers and the second segment part corresponding to the second number of items q;
[0211] In step 78, according to the p-th order full evaluation matrix and the additional matrix, determine the k-th order high-order evaluation matrix corresponding to the target number of terms k; the s-th order high-order / full evaluation matrix corresponding to any number of terms s is used to apply to the multiplier vectors composed of s multiplier segments of two multipliers respectively to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the high-order / full segment of the product; wherein, the k-th order high-order evaluation matrix is used as the processing parameter of the rounding multiplication circuit for rounding the k-th power of r.
[0212] In one embodiment, the first number of terms p is an integer greater than k / 2.
[0213] According to an implementation manner, a strategy of independently processing two multipliers can be adopted to determine the additional matrix and the k-th order high-order evaluation matrix. In this implementation manner, determining the additional matrix may include: determining a first matrix, which is used to calculate the q-th order rounding of the first segment part of the first multiplier and the second segment part of the second multiplier; determining a second matrix, which is used to calculate the q-th order rounding of the second segment part of the first multiplier and the first segment part of the second multiplier.
[0214] More specifically, in one embodiment, p all-zero columns can be added to the right of the q-th order high-order evaluation matrix corresponding to the second number of terms q as the first matrix; p all-zero columns can be added to the left of the q-th order high-order evaluation matrix as the second matrix.
[0215] Correspondingly, determining the k-th order high-order evaluation matrix corresponding to the target number of terms k may include: sequentially stacking the base matrix, the above-mentioned first matrix and the second matrix to obtain the first evaluation matrix for the first multiplier; wherein the base matrix is the augmented matrix of the p-th order full evaluation matrix; sequentially stacking the base matrix, the second matrix and the first matrix to obtain the second evaluation matrix for the second multiplier. The first evaluation matrix and the second evaluation matrix constitute the k-th order high-order evaluation matrix.
[0216] According to another implementation manner, a strategy of combined processing of two multipliers can be adopted to determine the additional matrix and the k-th order high-order evaluation matrix. In this implementation manner, determining the additional matrix includes: enumerating the segment indexes of the multiplier segments that meet the target conditions, and filling the matrix elements according to the segment indexes to obtain the additional matrix, wherein the target conditions include that if a row has a single segment index, the segment index is less than or equal to p; if a row has a combination of two multiplier segments, the first segment index i in the combination is less than the second segment index j, the sum of the first segment index i and the second segment index j is greater than or equal to k-1, and the first segment index is less than q. That is, for the combination of two multiplier segments, the constraint conditions that the first segment index i and the second segment index j need to meet include: (1) i < j; (2) i + j >= k - 1; (3) i < q.
[0217] Under this embodiment, determining the k-th order high-order evaluation matrix corresponding to the target number of terms k may include: successively stacking the base matrix and the above-mentioned additional matrix to obtain the k-th order high-order evaluation matrix, where the base matrix is an augmented matrix of the p-th order full-scale evaluation matrix.
[0218] Similar to Figure 6 For the k-th order high-order evaluation matrix, it is also possible to first compare the computational costs of the matrices obtained under the two strategies, select the strategy corresponding to the lower computational cost, and thus generate a k-th order high-order evaluation matrix with a lower computational cost as the configuration parameter of the rounding multiplication circuit.
[0219] In one embodiment, the above-mentioned rounding multiplication circuit and modulo multiplication circuit are sub-circuits in the modular multiplication circuit. For example, they can respectively correspond to Figure 4 Circuit 12 and circuit 13 in the modular multiplication circuit shown. In other embodiments, the rounding multiplication circuit and the modulo multiplication circuit can also be independent circuits dedicated to performing rounding multiplication operations and modulo multiplication operations respectively.
[0220] According to an embodiment of another aspect, a device for determining the processing parameters of a modulo multiplication circuit is provided. Figure 8 The structural schematic diagram of a modulo multiplication parameter determination device according to an embodiment is shown. This device can be deployed in any device, platform or device cluster with data storage, computing, and processing capabilities. As Figure 8 shown, the modulo multiplication parameter determination device 800 includes:
[0221] A first partitioning unit 82, configured to partition the target number of terms k for partitioning the multiplier into multiplier segments into the sum of a first number of terms p and a second number of terms q, where the second number of terms q corresponds to the low-order segment part;
[0222] A first obtaining unit 84, configured to obtain the q-th order full-scale evaluation matrix corresponding to the second number of terms q;
[0223] A first determining unit 86, configured to determine an additional matrix for calculating the cross product between the first segment part corresponding to the first number of terms p of the two multipliers and the second segment part corresponding to the second number of terms q;
[0224] A first calculating unit 88, configured to determine the k-th order low-order evaluation matrix corresponding to the target number of terms k according to the q-th order full-scale evaluation matrix and the additional matrix; the s-th order low-order / full-scale evaluation matrix corresponding to any number of terms s is used to apply to the multiplier vectors composed of s multiplier segments of the two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the low-order / full-scale segments of the product; wherein, the k-th order low-order evaluation matrix is used as the processing parameter of the modulo multiplication circuit for taking the modulo of r to the k-th power.
[0225] According to an embodiment of another aspect, a device for determining processing parameters of a rounding multiplication circuit is provided. Figure 9 The structural schematic diagram of a rounding multiplication parameter determination device according to an embodiment is shown. This device can be deployed in any device, platform, or device cluster with data storage, computing, and processing capabilities. As Figure 8 shown, the rounding multiplication parameter determination device 900 includes:
[0226] A second partitioning unit 92, configured to partition a target number of items k for partitioning a multiplicand into multiplicand segments into the sum of a first number of items p and a second number of items q, where the first number of items p corresponds to the high-order segment part;
[0227] A second obtaining unit 94, configured to obtain a p-order full-amount evaluation matrix corresponding to the first number of items p;
[0228] A second determining unit 96, configured to determine an additional matrix, which is used to calculate the cross product between the first segment part corresponding to the first number of items p of the two multiplicands and the second segment part corresponding to the second number of items q;
[0229] A second calculating unit 98, configured to determine a k-order high-order evaluation matrix corresponding to the target number of items k according to the p-order full-amount evaluation matrix and the additional matrix; an s-order high-order / full-amount evaluation matrix corresponding to any number of items s, which is used to be applied to a multiplicand vector composed of s multiplicand segments of the two multiplicands to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the high-order / full-amount segment of the product; wherein, the k-order low-order evaluation matrix is used as the processing parameter of the rounding multiplication circuit for rounding to the kth power of r.
[0230] For the implementation manners of the above units in the device, reference can be made to the description in combination with the method embodiments. Through the above device, the k-order low-order evaluation matrix can be quickly determined as the parameter of the modulo multiplication circuit, and the k-order high-order evaluation matrix can be determined as the processing parameter of the rounding multiplication circuit, saving the cost of parameter configuration for the modulo multiplication circuit and the rounding multiplication circuit.
[0231] According to an embodiment of another aspect, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the combination Figure 5 and / or Figure 7 the described method.
[0232] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the combination Figure 5 and / or Figure 7 the described method is implemented.
[0233] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0234] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for determining processing parameters of a modular multiplication circuit, comprising: For a target number of terms k for dividing a multiplier into multiplier segments, dividing it into the sum of a first number of terms p and a second number of terms q, where the second number of terms q corresponds to the low-order segment part; Obtaining a q-th order full-amount evaluation matrix corresponding to the second number of terms q; Determining an additional matrix for calculating the cross product between the first segment part corresponding to the first number of terms p of two multipliers and the second segment part corresponding to the second number of terms q; According to the q-th order full-amount evaluation matrix and the additional matrix, determining a k-th order low-order evaluation matrix corresponding to the target number of terms k; an s-th order low-order / full-amount evaluation matrix corresponding to any number of terms s is used to be applied to multiplier vectors composed of s multiplier segments of two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the low-order / full-amount segments of the product; wherein, the k-th order low-order evaluation matrix is used as the processing parameter of a modular multiplication circuit for taking modulo of r to the k-th power.
2. The method according to claim 1, wherein, The second number of terms q is an integer greater than k / 2.
3. The method according to claim 1, wherein, Determining the additional matrix includes: Determining a first matrix for calculating the p-th order modular multiplication of the first segment part of the first multiplier and the second segment part of the second multiplier; Determining a second matrix for calculating the p-th order modular multiplication of the second segment part of the first multiplier and the first segment part of the second multiplier.
4. The method according to claim 3, wherein, Determining the first matrix includes: adding q all-zero columns to the right of the p-th order low-order evaluation matrix corresponding to the first number of terms p as the first matrix; Determining the second matrix includes: adding q all-zero columns to the left of the p-th order low-order evaluation matrix as the second matrix.
5. The method according to claim 3, wherein, The k-th order low-order evaluation matrix includes a first evaluation matrix acting on the first multiplier and a second evaluation matrix acting on the second multiplier; According to the q-th order full-amount evaluation matrix and the additional matrix, determining the k-th order low-order evaluation matrix corresponding to the target number of terms k includes: Stacking the base matrix, the first matrix and the second matrix in sequence to obtain the first evaluation matrix; the base matrix is the augmented matrix of the q-th order full-amount evaluation matrix; Stacking the base matrix, the second matrix and the first matrix in sequence to obtain the second evaluation matrix.
6. The method according to claim 1, wherein Determining the additional matrix includes: Enumerating the segment indices of multiplier segments that meet the target conditions, and filling the matrix elements according to the segment indices to obtain the additional matrix, where the target conditions include that if a row has a single segment index, the segment index is greater than or equal to q; if a row has a combination of two multiplier segments, the first segment index i in the combination is less than the second segment index j, the sum of the first segment index i and the second segment index j is less than or equal to k - 1, and the second segment index is greater than or equal to q.
7. The method according to claim 6, wherein, According to the q-th order full-amount evaluation matrix and the additional matrix, determining the k-th order low-order evaluation matrix corresponding to the target number of terms k includes: Stacking the base matrix and the additional matrix in sequence to obtain the k-th order low-order evaluation matrix, and the base matrix is the augmented matrix of the q-th order full-amount evaluation matrix.
8. The method according to claim 3, further comprising: Determine a first cost value according to the number of effective multiplications for modular multiplication using the p-term Karatsuba algorithm; and, determine a second cost value based on algebraic operations of p; Determine an additional matrix, including, in response to the first cost value being less than the second cost value, determining the first matrix and the second matrix.
9. The method according to claim 6, further comprising: Determine a first cost value according to the number of effective multiplications for modular multiplication using the p-term Karatsuba algorithm; and, determine a second cost value based on algebraic operations of p; Determine an additional matrix, including, in response to the first cost value not being less than the second cost value, performing fragment indexing for enumerating multiplier fragments that meet the target conditions.
10. The method according to claim 1, further comprising configuring the modular multiplication circuit with the k-th order low-order evaluation matrix as a processing parameter corresponding to the radix r and the number of terms k.
11. The method according to claim 1, wherein, The modular multiplication circuit is a sub-circuit in a modular multiplication circuit that performs modular multiplication using the Barrett reduction method.
12. A method for determining processing parameters of a rounding multiplication circuit, comprising: For a target number of terms k for dividing a multiplier into multiplier fragments, divide it into the sum of a first number of terms p and a second number of terms q, where the first number of terms p corresponds to the high-order fragment part; Obtain a p-th order full evaluation matrix corresponding to the first number of terms p; Determine an additional matrix for calculating the cross product between the first fragment part corresponding to the first number of terms p of two multipliers and the second fragment part corresponding to the second number of terms q; According to the p-th order full evaluation matrix and the additional matrix, determine a k-th order high-order evaluation matrix corresponding to the target number of terms k; an s-th order high-order / full evaluation matrix corresponding to any number of terms s is used to apply to multiplier vectors composed of s multiplier fragments of two multipliers respectively to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the high-order / full fragment of the product; wherein, the k-th order low-order evaluation matrix is used as a processing parameter for a rounding multiplication circuit for rounding to the k-th power of r.
13. An apparatus for determining processing parameters of a modular multiplication circuit, comprising: A first division unit configured to, for a target number of terms k for dividing a multiplier into multiplier fragments, divide it into the sum of a first number of terms p and a second number of terms q, where the second number of terms q corresponds to the low-order fragment part; A first acquisition unit configured to obtain a q-th order full evaluation matrix corresponding to the second number of terms q; A first determination unit configured to determine an additional matrix for calculating the cross product between the first fragment part corresponding to the first number of terms p of two multipliers and the second fragment part corresponding to the second number of terms q; A first calculation unit configured to, according to the q-th order full evaluation matrix and the additional matrix, determine a k-th order low-order evaluation matrix corresponding to the target number of terms k; an s-th order low-order / full evaluation matrix corresponding to any number of terms s is used to apply to multiplier vectors composed of s multiplier fragments of two multipliers respectively to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the low-order / full fragment of the product; wherein, the k-th order low-order evaluation matrix is used as a processing parameter for a modular multiplication circuit for taking the modulus of the k-th power of r.
14. An apparatus for determining processing parameters of a rounding multiplication circuit, comprising: A second partitioning unit configured to partition a target number of items k for partitioning a multiplicand into multiplicand segments into a sum of a first number of items p and a second number of items q, where the first number of items p corresponds to a high-order segment part; A second obtaining unit configured to obtain a p-th order full evaluation matrix corresponding to the first number of items p; A second determining unit configured to determine an additional matrix for calculating a cross product between a first segment part corresponding to the first number of items p of two multiplicands and a second segment part corresponding to the second number of items q; A second calculating unit configured to determine a k-th order high-order evaluation matrix corresponding to the target number of items k according to the p-th order full evaluation matrix and the additional matrix; an s-th order high-order / full evaluation matrix corresponding to any number of items s, which is used to be applied to a multiplicand vector composed of s multiplicand segments of two multiplicands to obtain two intermediate vectors, and an element combination of the two intermediate vectors is used to obtain a high-order / full segment of the product; wherein, the k-th order low-order evaluation matrix is used as a processing parameter of a rounding multiplication circuit for rounding to the k-th power of r.
15. A computing device, comprising a memory and a processor, characterized in that, The memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-12 is implemented.