Method and device for determining processing parameters for modulo multiplication circuit and rounding multiplication circuit
Through factor decomposition and Karatsuba algorithm optimization of the processing parameters of the modular multiplication and rounding multiplication circuit, the problem of insufficient modular multiplication operation performance is solved and efficient computation of encrypted calculation is realized.
Patent Information
- Application Number
- CN202510443694.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the performance optimization of the modular multiplication operation in encrypted computing is insufficient, especially the high cost of division calculation, which affects the efficiency of privacy calculations.
The first term and the second term are determined by factorization, the low-bit evaluation matrix and the high-bit evaluation matrix are used, and the Kronecker product operation is combined with the processing parameters of the modulus multiplication and rounding multiplication circuit. The Karatsuba algorithm is used for multiplication optimization to reduce the number of multiplications and circuit costs.
It effectively reduces the calculation cost of modulus multiplication and rounding multiplication circuits, improves the performance and efficiency of encrypted computing, and is suitable for the field of privacy computing.
Smart Images

Figure CN120335765A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of hardware design for encrypted computing, and particularly to methods and apparatuses for determining processing parameters for a rounding multiplication circuit and a modulo multiplication circuit. Background Art
[0002] With the continuous improvement of people's awareness of data security and privacy protection, privacy computing, as a new type of secure computing mode, is gradually becoming one of the mainstream methods for data processing and analysis. Privacy computing realizes the protection of data privacy and security by directly performing calculations on encrypted or anonymized data without exposing the original data, and has broad application prospects. At present, privacy computing technologies have been widely applied in fields such as artificial intelligence, finance, and healthcare, and have become an important supporting technology for the future digital society.
[0003] Privacy computing relies on the use of various encryption algorithms, including the RSA encryption algorithm, the elliptic curve ECC encryption algorithm, the homomorphic encryption algorithm, and so on. In particular, the homomorphic encryption algorithm has become one of the mainstream technologies in the field of privacy computing, and can perform various common calculation operations in the encrypted state while ensuring the correctness of the results and the privacy of the data. Among the above various encryption algorithms, the arithmetic operations are often performed in a certain modulo space, and the modulo multiplication operation, as a basic operation and primitive, has an important impact on the performance of encrypted computing. At present, some hardware solutions have been proposed to accelerate the performance of the modulo multiplication operation.
[0004] It is hoped that there will be an improved solution to further optimize the implementation process of the basic arithmetic operations related to encrypted computing. Summary of the Invention
[0005] One or more embodiments of this specification describe a method and an apparatus for determining processing parameters for a modulo multiplication circuit and a rounding multiplication circuit, which can save the cost of parameter design and configuration for the modulo multiplication circuit and the rounding multiplication circuit.
[0006] According to a first aspect, there is provided a method for determining processing parameters of a modulo multiplication circuit, including:
[0007] Performing factorization on a target number of items k to be processed, and determining a first number of items k1 and a second number of items k2 from the decomposed factors;
[0008] Determine the k - order low - order evaluation matrix corresponding to the target number of terms k according to the k1 - order low - order evaluation matrix corresponding to the first number of terms k1, the k2 - order low - order evaluation matrix corresponding to the second number of terms k2, and the k2 - order full - amount evaluation matrix. Wherein, the p - order low - order / full - amount evaluation matrix corresponding to any number of terms p is used to apply to the multiplier vectors respectively composed of p multiplier segments corresponding to the two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the low - order / full - amount segments of the product; the k - order low - order evaluation matrix is used as the processing parameter of the modulo - multiplication circuit for taking the modulo of r to the k - th power.
[0009] In one embodiment, determining the k - order low - order evaluation matrix corresponding to the target number of terms k specifically includes:
[0010] Divide the k1 - order low - order evaluation matrix into a first matrix part composed of target rows and a second matrix part composed of the remaining rows; wherein, the element combination between the elements generated by the target rows in the two intermediate vectors is realized through a semi - multiplication circuit;
[0011] Perform a Kronecker product operation on the second matrix part and the k2 - order full - amount evaluation matrix to obtain a first result part;
[0012] Perform a Kronecker product operation on the first matrix part and the k2 - order low - order evaluation matrix to obtain a second result part; the splicing of the first result part and the second result part forms a result matrix, and the result matrix is used to form the k - order low - order evaluation matrix.
[0013] According to one implementation manner, determining the first number of terms k1 and the second number of terms k2 from the decomposed factors specifically includes:
[0014] Determine the first factor and the second factor from the decomposed factors;
[0015] For any factor ki, determine the ratio of the first number and the second number as the cost coefficient, where the first number is the reduced number of the first multiplication times involved in the modulo - multiplication operation using the ki - term Karatsuba algorithm compared to the second multiplication times involved in the conventional multiplication operation; the second number is the minimum number of times of using the full - multiplication circuit in the modulo - multiplication operation using the ki - term Karatsuba algorithm;
[0016] If the cost coefficient of the first factor is greater than that of the second factor, determine the first factor as the first number of terms k1 and the second factor as the second number of terms k2; otherwise, determine the second factor as the first number of terms k1 and the first factor as the second number of terms k2.
[0017] In one example, any element in the k1 - order and k2 - order low - order transformation matrices is selected from 0, 1, and - 1.
[0018] According to one embodiment, the above method further includes using the k-th order low-order evaluation matrix as the processing parameter corresponding to the radix r and the number of terms k, and configuring the modulo multiplication circuit with the parameter.
[0019] In one example, the aforementioned modulo multiplication circuit is a sub-circuit in the modulo multiplication circuit.
[0020] According to a second aspect, a method for processing parameters of a rounding multiplication circuit is provided, including:
[0021] Factorize the target number of terms k to be processed, and determine the first number of terms k1 and the second number of terms k2 from the decomposed factors;
[0022] Determine the k-th order high-order evaluation matrix corresponding to the target number of terms k according to the k1-th order high-order evaluation matrix corresponding to the first number of terms k1, the k2-th order high-order evaluation matrix corresponding to the second number of terms k2, and the k2-th order full-scale evaluation matrix, where the p-th order high-order / full-scale evaluation matrix corresponding to any number of terms p is used to apply to the multiplier vectors respectively composed of p multiplier segments corresponding to the two multipliers to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the high-order / full-scale segment of the product; the k-th order high-order evaluation matrix is used as the processing parameter for the rounding multiplication circuit for rounding the k-th power of r.
[0023] According to a third aspect, a device for determining the processing parameters of a modulo multiplication circuit is provided, including:
[0024] A first decomposition unit configured to factorize the target number of terms k to be processed, and determine the first number of terms k1 and the second number of terms k2 from the decomposed factors;
[0025] A first determination unit configured to determine the k-th order low-order evaluation matrix corresponding to the target number of terms k according to the k1-th order low-order evaluation matrix corresponding to the first number of terms k1, the k2-th order low-order evaluation matrix corresponding to the second number of terms k2, and the k2-th order full-scale evaluation matrix, where the p-th order low-order / full-scale evaluation matrix corresponding to any number of terms p is used to apply to the multiplier vectors respectively composed of p multiplier segments corresponding to the two multipliers to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the low-order / full-scale segment of the product; the k-th order low-order evaluation matrix is used as the processing parameter for the modulo multiplication circuit for taking the modulo of the k-th power of r.
[0026] According to a fourth aspect, a device for determining the processing parameters of a rounding multiplication circuit is provided, including:
[0027] A second decomposition unit configured to factorize the target number of terms k to be processed, and determine the first number of terms k1 and the second number of terms k2 from the decomposed factors;
[0028] A second determination unit, configured to determine a k-th order high-order evaluation matrix corresponding to a target number of terms k according to a k1-th order high-order evaluation matrix corresponding to a first number of terms k1, a k2-th order high-order evaluation matrix corresponding to a second number of terms k2, and a k2-th order full-amount evaluation matrix, where a p-th order high-order / full-amount evaluation matrix corresponding to any number of terms p is used to respectively apply to multiplier vectors formed by p multiplier segments corresponding to two multipliers to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain a high-order / full-amount segment of the product; the k-th order high-order evaluation matrix is used as a processing parameter of a rounding multiplication circuit for rounding the k-th power of r.
[0029] According to a fifth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the first aspect or the second aspect.
[0030] According to a sixth aspect, there is provided a computing device, including a memory and a processor, characterized in that an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect or the second aspect is implemented.
[0031] In the embodiments of the present specification, for a target order to be processed, it is decomposed into factors of lower orders. Through a recursive manner, a high-order transformation matrix is quickly obtained based on lower-order transformation matrices and used as a processing parameter of a modular multiplication circuit / rounding multiplication circuit. Further, in some implementation manners, cost coefficients corresponding to each factor are defined. Based on the cost coefficients of each factor, an optimal decomposition sorting of the factors can be efficiently determined to obtain a high-order transformation matrix with a lower calculation cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 Illustrates an algorithm process of modular multiplication using Barrett modular reduction;
[0034] Figure 2 Illustrates a schematic diagram of the principle of the 2-term Schoolbook algorithm;
[0035] Figure 3 Illustrates a schematic diagram of the number of bits in the operation;
[0036] Figure 4 Illustrates a schematic diagram of a modular multiplication circuit according to an embodiment;
[0037] Figure 5 Flowchart showing a method for determining processing parameters of a modular multiplication circuit according to an embodiment;
[0038] Figure 6 Flowchart showing a method for determining processing parameters of a rounding multiplication circuit according to an embodiment;
[0039] Figure 7 Schematic structural diagram showing a modular multiplication parameter determination device according to an embodiment;
[0040] Figure 8 Schematic structural diagram showing a rounding multiplication parameter determination device according to an embodiment. Detailed implementation manners
[0041] The solutions provided in this specification will be described below with reference to the accompanying drawings.
[0042] To achieve privacy computing, various encryption algorithms are adopted. In particular, homomorphic encryption has become one of the mainstream technologies in the field of privacy computing because it can perform corresponding mathematical calculations on encrypted data. The operations of homomorphic encryption involve a large number of four arithmetic operations in the sense of modulo. The main difficulty of modulo operation lies in that division calculation is involved in the process of calculating the quotient value, and the implementation cost of division is relatively high. For this reason, the Barrett reduction algorithm is proposed. This algorithm achieves an effect similar to division through multiplication and shift operations, thereby improving the efficiency of modulo operation.
[0043] Modular multiplication is a commonly used operation in encryption systems, that is, after multiplying two integers X and Y, a modulo operation is performed on a modulus M, that is, X * Y mod M. Modular multiplication can be implemented by a hardware circuit using the Barrett reduction idea.
[0044] Figure 1 Shows the algorithm process of modular multiplication using Barrett modular reduction. As Figure 1 shown, the inputs of this algorithm include the multiplier X, Y, the modulus M, and the pre-computed value μ corresponding to the modulus M. Here it is assumed that both the multiplier X and Y are not greater than the modulus M corresponding to the modulus space. And the value range of this modulus M can be expressed as [2 β-1 , 2 N ), where β is a very small constant, for example, 2; N is the maximum bit width supported by the system. Thus, the modulus M can take various values within the bit width N acceptable to the system. The parameters α and t involved in the pre-computed value can be set as needed. Specifically, α and β can be appropriately set so that the Z generated in the 4th line of the algorithm is within the range of [0, 2M], so that in lines 5 - 6, at most one correction operation is required.
[0045] It can be seen that in the algorithm process of modular multiplication using Barrett reduction, the following operations are relatively high-cost core operations: (1) Conventional multiplication calculation, such as the calculation X*Y shown in line 1 of the algorithm; (2) Rounding multiplication calculation, that is, calculating the truncated rounding result of the product of two numbers multiplied relative to the target value in the form of a power of 2, such as the calculation shown in line 3 of the algorithm where denotes rounding down; (3) Modular multiplication calculation, that is, calculating the modulo result of the product of two numbers multiplied for the target value in the form of a power of 2, such as the calculation involved in line 4 of the multiplication qM mod 2 N+1 .
[0046] Without loss of generality, conventional multiplication can be represented as AB, rounding multiplication can be represented as Modular multiplication is represented as AB mod r k , where r = 2 m , and where denotes rounding up. That is, the target value in the form of a power of 2 (when applied to the Figure 1 algorithm, here N may be different from the bit width corresponding to the modulus M) is divided into k segments, each segment has a bit width of m, and the base of each segment is r.
[0047] To accelerate the modular multiplication operation, the operation processes and calculation circuits of the above conventional multiplication, rounding multiplication, and modular multiplication can be optimized and improved respectively.
[0048] For conventional multiplication, the Karatsuba algorithm can be used for acceleration. According to the k-term Karatsuba algorithm, first, the multiplier A (and multiplier B) with a bit width of N is divided into k segments, or called k words. Each segment A i has a bit width of Thus, the multiplier A can be represented as where r = 2 m is the base of each segment, and A i = A[(i + 1)m - 1:im].[[]END]]
[0049] Thus, the product result C of AB can be calculated by the following formula:
[0050]
[0051] Taking k = 2 as an example, the process of calculating the product C using the Schoolbook algorithm is shown in the following formula:
[0052] C = AB = A1B1r 2 +(A1B0 + A0B1)r + A0B0 (2)
[0053] Figure 2 Shows the schematic diagram of the principle of the 2-term Schoolbook algorithm.
[0054] It can be found that if the calculation process of formula (2) is strictly followed and A is calculated item by item i B j and then summed up, 4 multiplication operations are required. However, for the second term in formula (2), what matters is the product sum A1B0 + A0B1, rather than the individual products. Therefore, the product sum can be recovered through the following formula:
[0055] A0B1 + A1B0 = -(A0 - A1)(B0 - B1) + A1B1 + A0B0 (3)
[0056] Since the results of A1B1 and A0B0 can be reused in formula (3), by calculating the second term in formula (2) using formula (3), the number of multiplication operations required in the calculation process of formula (2) can be reduced to 3 times, thereby realizing the optimization of the Karatsuba algorithm.
[0057] Still taking k = 2 as an example below, the optimization process of the Karatsuba algorithm is divided into three stages: the evaluation stage, the multiplication stage, and the interpolation stage. In the preparatory process, the multiplier can be represented as a vector composed of each multiplier segment (or word), called the multiplier vector. For example, in the case of k = 2, the two multipliers can be respectively represented as the multiplier vectors: A = [A0, A1] T , B = [B0, B1] T .
[0058] In the evaluation stage, the evaluation matrix E is applied to the multiplier vector to obtain the corresponding generated vector. The elements in the generated vector are combinations of multiplier segments and are used as the basis for subsequent multiplications. For example, when k = 2, the 2nd-order evaluation matrix E (2) can take the following form:
[0059]
[0060] Applying this evaluation matrix to the multiplier vector corresponding to multiplier A, the generated vector e corresponding to multiplier A can be obtained as follows A :
[0061]
[0062] Similarly, the generated vector corresponding to multiplier B can also be obtained.
[0063] Then, in the multiplication stage, the generated vector e corresponding to multiplier A A and the generated vector e corresponding to multiplier B B are multiplied bit by bit to obtain the intermediate vector, that is: e = eA ⊙e B , where ⊙ is the Hadamard operator, representing element-wise multiplication. In the example where k = 2, the intermediate vector e obtained is as follows:
[0064]
[0065] Next, in the interpolation stage, multiply the intermediate vector by the interpolation matrix I, and the resulting vector shows each segment of the product result. In the example where k = 2, the 2nd-order interpolation matrix I (2) can take the following form:
[0066]
[0067] In this way, the resulting vector is as follows:
[0068]
[0069] The radix vector R = [r 0 , r 1 , r 2 T can act on the above resulting vector, that is, calculate R T ·C, thereby recovering the value of the product result C.
[0070] Combining the above three stages, the product calculation process under the k-term Karatsuba algorithm can be summarized as follows:
[0071] C = Karatsuba(A, B) = R T ·I (k) ·((E (k) ·A) ⊙ (E (k) ·B)) (9)
[0072] According to formula (9), pre-construct the transformation matrix under the k-term Karatsuba algorithm, including the evaluation matrix E (k) and the interpolation matrix I (k) . Use the evaluation matrix E (k) to act on the vector composed of the multiplier segments of the multiplicand A and the vector composed of the multiplier segments of the multiplier B respectively, multiply the two resulting vectors element-wise to obtain an intermediate vector; then multiply the intermediate vector by the interpolation matrix I (k) , and the resulting vector C thus obtained shows each segment of the product.
[0073] In practice, the construction of the evaluation matrix E (k) and the interpolation matrix I (k) is not unique, and different matrix forms may bring different computational costs, such as different numbers of multiplications.
[0074] The above optimization ideas based on the k-term Karatsuba algorithm can also be applied to integer multiplication and modular multiplication AB mod r k .
[0075] It can be understood that according to formula (1), in the k-term Karatsuba algorithm, the product C of AB contains segments from C0 to C 2k-2 The segment subscripts correspond to the orders of r. For modular multiplication AB mod r k , the segments of the product C with orders higher than k have no influence on the result. Therefore, only the product segments of the lower-order part within the k-th order need to be considered. For this purpose, the idea of implementing the k-term Karatsuba algorithm using the transformation matrix (E (k) , I (k) ) can be referred to, and a lower-order transformation matrix is constructed Through similar operations, the lower-order part C L in the product result is obtained. Specifically, the k-term Karatsuba lower-order algorithm for modular multiplication can be implemented as follows:
[0076]
[0077] where R L = [r 0 , r 1 , …, r k-1 T .
[0078] Then take the modulus of C L with respect to r k . It can be seen that the process of obtaining the lower-order product part using the lower-order transformation matrix is similar to the process of implementing conventional multiplication: using the lower-order evaluation matrix to act on the vectors composed of the multiplier segments of the multiplier A and the vectors composed of the multiplier segments of the multiplier B respectively, multiplying the two obtained vectors bit by bit to obtain the intermediate vector e L ; then multiplying the intermediate vector e by the lower-order interpolation matrix L , and the resulting result vector C L shows the lower-order part of the product.
[0079] Still taking k = 2 as an example to describe an example process. In this example, the lower-order evaluation matrix and the lower-order interpolation matrix can take the following forms respectively:
[0080]
[0081] In this way, the obtained result is:
[0082]
[0083] in the form of formula (11) can further simplify circuit calculations because when calculating C next L for r 2 when taking the modulus, in e L,1 and e L,2 only the last m bits in the terms need to be considered, that is, only e L,1 mod r and e L,2 mod r need to be considered. Its principle can be illustrated by Figure 3 the digit diagram of
[0084] It can be understood that r = 2 m is the base or size of a segment, and m is the bit width of a segment. Any e L,i is obtained by multiplying the additive combination of multiplier segment A i and the additive combination of multiplier segment B j . The maximum bit width of each additive combination is m + 1 (carry may occur during addition). Therefore, the maximum bit width of e L,i is 2m + 2, slightly exceeding the size of two segments. As Figure 3 shown, e L,i r is obtained by shifting e L,i to the left by one segment. When calculating C L mod r 2 , the part exceeding the size of r 2 has no effect on the result, as shown by the shaded part in the figure. Only the white part falling within the range of r 2 has an effect on the result. It can be seen that only the last m bits of e L,i r fall within the range of r 2 .
[0085] Therefore, the modulus of C L can be calculated as:
[0086] C L mod r 2 = e L,0 (r + 1) + (e L,1 mod r)r - (e L,2 mod r)r mod r 2 (13)
[0087] Since only the last m bits of e L,1 and e L,2 need to be considered, their calculations can be implemented by a half - multiplication circuit. The half - multiplication circuit includes a high - order half - multiplication circuit that only calculates the high - order part and a low - order half - multiplication circuit that only calculates the low - order part. In formula (13), the low - order half - multiplication circuit can be used to calculate eL,1 and e L,2 , obtain e L,1 mod r and e L,2 mod r. Compared with the full multiplication circuit that calculates all digits, the half multiplication circuit can save circuit area and calculation cost. The calculation cost of the half multiplication circuit can be considered as 0.5 times of full multiplication calculation. Correspondingly, by calculating the modular multiplication through formula (13), only the calculation cost of 1 + 2 * 0.5 = 2 times of full multiplication operations is required. Compared with the 3 multiplications of the 2-term Karatsuba algorithm for implementing conventional multiplication, the cost is further reduced.
[0088] In practice, in order to reduce the search range and implement the operation simply and efficiently, preferably, any element in the low-order evaluation matrix is constrained to be selected from 0, 1, and -1. The interpolation matrix is generated according to the corresponding evaluation matrix and may contain other values.
[0089] Relative to the modular multiplication, for the integer multiplication , the fragments in the product C with orders lower than k have no influence on the result, so only the product fragments of the high-order part with orders above k need to be considered. For this purpose, similarly referring to the idea of using the transformation matrix (E (k) , I (k) ) to implement the k-term Karatsuba algorithm, construct the high-order transformation matrix Through similar operations, obtain the high-order part in the product result, that is, the high-order product result C U . Specifically, the k-term Karatsuba high-order algorithm for modular multiplication can be implemented as follows:
[0090]
[0091] where, R U = [r k-1 , r k , …, r 2k-2 T .
[0092] Then, calculate C U / r k to obtain. It can be seen that the process of obtaining the high-order product result using the high-order transformation matrix is similar to the process of implementing conventional multiplication: using the high-order evaluation matrix to act on the vector composed of the multiplier segments of the multiplier A and the vector composed of the multiplier segments of the multiplier B respectively, multiply the two obtained vectors bit by bit to obtain the intermediate vector e U ; then multiply the intermediate vector e by the high-order interpolation matrix U , and the resulting result vector C U shows the high-order part of the product.
[0093] Similar to optimizing modular multiplication using an appropriate form, the process of integer multiplication can also be optimized by adopting an appropriate form of high-order transformation matrix to reduce the computational cost, including using a high-order semi-multiplication circuit to calculate partial product terms. And, similar to the low-order evaluation matrix, preferably, any element in the high-order evaluation matrix is selected from 0, 1, and -1. integer multiplication is optimized, including using a high-order semi-multiplication circuit to calculate partial product terms, thereby reducing the computational cost. And, similar to the low-order evaluation matrix, preferably, any element in the high-order evaluation matrix is selected from 0, 1, and -1.
[0094] Figure 4 shows a schematic diagram of a modular multiplication circuit according to an embodiment. This modular multiplication circuit utilizes the idea of the k-term Karatsuba algorithm to implement Figure 1 the calculation process of modular multiplication through Barrett reduction as in Figure 4 . As shown, this modular multiplication circuit includes a circuit 11 for conventional multiplication, a circuit 12 for integer multiplication, and a circuit 13 for modular multiplication. Specifically, the inputs of circuit 11 include the multiplicands X and Y. In this circuit, using the evaluation matrix E (k) as the circuit parameter, addition operations are respectively performed on the multiplier segments corresponding to the multiplicands X and Y to obtain the generated vectors e A and e B for each element in. After compression and combination of the generated vectors, using the interpolation matrix I (k) as the parameter, the result T of XY multiplication as shown in the Figure 1 algorithm can be obtained.
[0095] For the sake of distinction in the following text, the transformation matrices (E (k) , I (k) ) in the k-term Karatsuba algorithm for implementing conventional multiplication are also referred to as the full-scale evaluation matrix and the full-scale interpolation matrix.
[0096] Circuit 12 implements integer multiplication using the KaratsubaUpper algorithm shown in formula (14). The inputs of circuit 12 include the pre-computed value μ and T H shifted from T (as shown in the second line of the Figure 1 algorithm). In this circuit 12, using the aforementioned high-order evaluation matrix as the circuit parameter, addition operations and corresponding encoding operations are respectively performed on the multiplier segments of the two inputs. After compression and combination of the generated vectors, using the high-order interpolation matrix as the parameter for combination and compression, the result of integer multiplication shown in the third line of the algorithm
[0097] Circuit 13 implements modular multiplication using the KaratsubaLower algorithm shown in formula (10). The inputs of circuit 13 include the modulus M and q output by circuit 12. In this circuit 13, using the aforementioned low-order evaluation matrix For circuit parameters, addition operations and corresponding encoding operations are respectively performed on the multiplier segments of the two input paths. After the generated vectors are compressed and combined, a low-order interpolation matrix is used as a parameter for combination and compression, and the modular multiplication result qM mod 2 can be obtained N+1 .
[0098] Finally, the modular multiplication result is combined with T obtained by processing T L (which can be achieved by truncation or shifting), and through basic circuit operations such as shifting and addition, the final modular multiplication result Z can be obtained.
[0099] In other embodiments, circuit 12 and circuit 13 can also be used as independent circuits dedicated to performing rounding multiplication and modular multiplication.
[0100] It can be understood that the operation of circuit 12 depends on the circuit configuration with the high-order transformation matrix as a circuit parameter, and the operation of circuit 13 depends on the circuit configuration with the low-order transformation matrix as a circuit parameter. For the same order k, the forms of the transformation matrices satisfying the forms of formulas (10) and (14) are not unique. Different forms of transformation matrices can optimize the circuit calculation process to different degrees. A better form of the transformation matrix, such as the form of formula (11) can use more half-multiplication circuits to optimize the circuit calculation and reduce the calculation cost.
[0101] Therefore, exploring the available transformation matrices for various numbers of terms k while minimizing the calculation cost as much as possible becomes a direction for optimizing the circuit calculation of modular multiplication / rounding multiplication. When k takes a small value, the dimension of the transformation matrix is limited and the search space is small; while when k gradually increases, the search space increases exponentially, bringing great difficulties to the determination of the transformation matrix and the configuration of the circuit. Therefore, it is hoped that there can be an improved solution to quickly determine the low-order transformation matrix / high-order transformation matrix with low calculation cost for a given number of terms k, so as to optimize the configuration of the modular multiplication circuit / rounding multiplication circuit.
[0102] First, the calculation cost corresponding to the transformation matrix is analyzed and defined below.
[0103] As described previously in combination with formulas (11)-(13) and Figure 3 , it can be found that when the i-th column of the low-order interpolation matrix has only the last row element non-zero (other rows are all zero), then this column will generate the target term e L in the low-order result C L,i r k-1 , this target item in the subsequent modulo operation C L mod r k is equivalent to e L,i mod r, that is, only the last m bits need to be considered. Therefore, this target item can be implemented by a half - multiplication circuit.
[0104] Therefore, the number of columns in the low - order interpolation matrix that meet the above conditions (only the elements in the last row are not zero), that is, the number corresponding to the half - multiplication calculation, and the number of the remaining columns, that is, the number corresponding to the regular full - multiplication calculation.
[0105] Similarly, if in the high - order interpolation matrix only the element in the first row of the i - th column is not zero, then this column will generate a target item e U in the high - order result C U,i r k-1 , this target item in the subsequent rounding - shift operation C U / r k is equivalent to e U,i / r, that is, only the first m bits need to be considered. Therefore, this target item can be implemented by a half - multiplication circuit. That is, the number of columns in the high - order interpolation matrix that meet the conditions (only the element in the first row is not zero), that is, the number corresponding to the half - multiplication calculation, and the number of the remaining columns, that is, the number corresponding to the regular full - multiplication calculation.
[0106] In addition, it is found that the low - order transformation matrix and the high - order transformation matrix can have a certain associated correspondence relationship, and the low - order transformation matrix and the high - order transformation matrix with an associated correspondence relationship consume the same number of half - multiplication times and full - multiplication times.
[0107] For clarity, the following defines the cost - related concepts in the multiplication circuit calculation based on the Karatsuba algorithm.
[0108] T (k) represents the number of full - multiplication calculations required for a regular multiplication (calculating AB) using the k - term Karatsuba algorithm. For the regular multiplication AB, all operations between multiplier segments need to be implemented by a full - multiplication circuit.
[0109] F (k) represents the number of full - multiplication calculations required for a rounding multiplication operation or a modulo multiplication operation AB mod r k using the k - term Karatsuba algorithm.
[0110] H(k) Denotes the number of half - multiplication calculations required for integer multiplication or modular multiplication \(AB\bmod r\) using the \(k\) - term Karatsuba algorithm. or modular multiplication \(AB\bmod r\) k when performing the operation.
[0111] Q (k) Denotes the number of effective multiplication calculations required for integer multiplication or modular multiplication \(AB\bmod r\) using the \(k\) - term Karatsuba algorithm. or modular multiplication \(AB\bmod r\) k when performing the operation.
[0112] R (k) Denotes the number of multiplication calculations required for integer multiplication or modular multiplication \(AB\bmod r\) using the \(k\) - term Karatsuba algorithm, which is the number of times reduced compared to ordinary multiplication \(AB\). That is, \(R\) or modular multiplication \(AB\bmod r\) k when performing the operation. That is, \(R\) (k) is \(F\) (k) +\(H\) (k) compared to \(T\) (k) the number of times saved / reduced, that is:
[0113] T (k) =\(F\) (k) +\(H\) (k) +\(R\) (k) (15)
[0114] In addition, since the cost of half - multiplication calculation can be considered as 1 / 2 of the full - multiplication calculation, the number of effective multiplication calculations \(Q\) (k) can be expressed as:
[0115] Q (k) =\(F\) (k) +\(H\) (k) / 2 (16)
[0116] Therefore, the goal that the solution in this specification hopes to achieve is to quickly determine the low - order transformation matrix / high - order transformation matrix / high - order transformation matrix under the number of terms \(k\). Further, make the number of effective multiplication calculations \(Q\) (k) used by the low - order / high - order transformation matrix as small as possible.
[0117] The following will be described in detail in combination with the low - order evaluation matrix After determining the low - order evaluation matrix the low - order interpolation matrix can be correspondingly determined The determination method of the high - order transformation matrix is similar.
[0118] It can be understood that the low - order evaluation matrix The order k, which corresponds to the number of terms k in the Karatsuba algorithm, represents splitting the multiplicand into k multiplier segments of equal bit-width, or k words. If the number of terms k is a composite number, it can be further expressed as the product of lower-order factors. For example, if k = k1k2, then each segment can be further split. Correspondingly, it is possible to attempt to use the lower-order evaluation matrices corresponding to the lower number of terms and to recursively determine the lower-order evaluation matrix of order k
[0119] Specifically, the multiplier vector A corresponding to the k-term Karatsuba algorithm has k = k1k2 segments, and each segment A i has a bit-width of The multiplier vector A can be rewritten as a vector A' of dimension k1, where the i-th segment A i ' has a bit-width of k2m and contains k2 sub-segments, that is:
[0120]
[0121] Thus, the k1-order lower-order transformation matrix can be used to apply the k1-term Karatsuba lower-order algorithm to the multiplier vectors A', B'. The i-th element e in the generated vector obtained by applying to the multiplier vector A' LA,i is the inner product of the i-th row of LA and the multiplier vector A', so each element of e LA has k2 words (i.e., sub-segments). The above calculation process is as follows:
[0122]
[0123] where is the element in the i-th row and j-th column of
[0124] In the multiplication stage, it is necessary to traverse i to calculate all e LA,i e LB,i . Note that since each element of e LA and e LB has k2 sub-segments, the Karatsuba algorithm can be further used for acceleration. More specifically, if e LA,i e LB,i needs to perform a full multiplication, the full transformation matrix can be further applied to calculate e LA,i e LB,i . If e LA,i e LB,i only needs a half multiplication, the lower-order transformation matrix can be further applied to calculate eLA,i e LB,i 。
[0125] Based on such analysis, a method for iteratively and recursively determining a k-th order low-order evaluation matrix (as a processing parameter of the modulo multiplication circuit) is proposed. Figure 5 The flowchart of the method for determining the processing parameter of the modulo multiplication circuit according to an embodiment is shown. This method can be executed by any circuit, device, or platform with computing and processing capabilities. As Figure 5 shown, the method includes the following steps.
[0126] In step 52, factorize the target number of terms k to be processed, and determine the first number of terms k1 and the second number of terms k2 from the decomposed factors. For simplicity, it may be assumed first that k = k1k2.
[0127] In step 54, according to the k1-th order low-order evaluation matrix corresponding to the first number of terms k1 the k2-th order low-order evaluation matrix corresponding to the second number of terms k2 and the k2-th order full-scale evaluation matrix determine the k-th order low-order evaluation matrix corresponding to the target number of terms k wherein, the p-th order low-order / full-scale evaluation matrix corresponding to any number of terms p is used to respectively apply to the multiplier vectors composed of p multiplier segments corresponding to the two multipliers to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain the low-order / full-scale segment of the product; the k-th order low-order evaluation matrix is used as the processing parameter of the modulo multiplication circuit for taking the modulo of r to the k-th power.
[0128] Specifically, the k1-th order low-order evaluation matrix can be divided into a first matrix part composed of target rows and a second matrix part composed of the remaining rows; wherein, the target rows are the rows corresponding to the semi-multiplication operation, that is, when the k1-th order low-order evaluation matrix is applied to the two multiplier vectors to obtain two intermediate vectors, the element combination between the elements generated by the target rows in the two intermediate vectors can be realized by a semi-multiplication circuit. Combining the previous analysis of the low-order interpolation matrix, it can be known that if some columns in the k1-th order low-order interpolation matrix have only the last row element not equal to 0, these columns can be called target columns, and the corresponding operations can be realized by semi-multiplication. The rows in the k1-th order low-order evaluation matrix corresponding to these target columns are the target rows. The order of rows / columns can be adjusted to group the target rows / target columns together. Correspondingly, the k1-th order low-order evaluation matrix can be divided into a first matrix part corresponding to the semi-multiplication operation and a second matrix part corresponding to the full-multiplication operation as follows:
[0129]
[0130] Based on this, the second matrix part of the full multiplication operation and the full evaluation matrix of order k2 are subjected to a Kronecker product operation to obtain a first result part; and the first matrix part of the semi-multiplication operation and the low-order evaluation matrix of order k2 are subjected to a Kronecker product operation to obtain a second result part. The concatenation of the first result part and the second result part forms a result matrix. When k = k1k2, the result matrix forms the low-order evaluation matrix of order k This operation process is as shown in formula (20):
[0131]
[0132] In the above formula, represents the Kronecker product operation symbol, and represents the custom operation process on the right side of formula (20).
[0133] Through the above operation process, based on the low-order evaluation matrix of order k1 the low-order evaluation matrix of order k2 and the full evaluation matrix of order k2 the low-order evaluation matrix of order k corresponding to the higher order k = k1k2 is recursively obtained
[0134] The above calculation idea can be extended to more factors. For example, assume k = k1k2k3. At this time, according to formulas (19) and (20), for further partition and expand the calculation, and the following can be obtained:
[0135]
[0136] It is easy to verify that the above custom operator satisfies the associative law:
[0137]
[0138] From this, it can be obtained that when the target number of terms k contains more factors, the operation process can be decomposed into further superimposing other factors on the basis of the operation results of two factors and calculating by analogy; in other words, the operation process can include repeatedly executing the calculation process for two factors. Step 54 describes the calculation process for two factors with k1 and k2 as examples, and this calculation process can be analogized to other factors. Thus, the low-order evaluation matrix of the target number of terms k with more factors is obtained.
[0139] In the above manner, for the number of terms \(k\) that can be factorized, it is possible to quickly and recursively obtain the low-order evaluation matrix of the higher-order number of terms \(k\) based on the low-order evaluation matrix and the full evaluation matrix of the low-order factors.
[0140] Furthermore, by observing formula (20), it can be found that the calculation of this formula is not symmetric with respect to the number of terms \(k1\) and the number of terms \(k2\). That is to say, the calculation cost of formula (20) is related to the order of \(k1\) and \(k2\). Exchanging the order of \(k1\) and \(k2\) may result in different calculation costs. For example, if \(k = 8\), we can set \(k1 = 2\) and \(k2 = 4\); or \(k = 4\) and \(k2 = 2\). The two low-order evaluation matrices obtained from these two factorization orders have different calculation costs.
[0141] In one implementation, after decomposing the number of terms \(k\) into multiple factors, the calculation costs corresponding to various orderings of the multiple factors can be determined respectively (measured by the number of effective multiplication operations \(Q\) (k) ), and the calculation cost can be determined from them.
[0142] In order to more quickly search and determine the low-order evaluation matrix with fewer effective multiplication operations \(Q\) (k) , the relationship between the calculation cost and the order of the number of terms will be explored below.
[0143] Still assuming the decomposition method of \(k = k1k2\). According to the aforementioned formulas (17) and (18), first use the \(k1\)-order low-order transformation matrix to apply the \(k1\)-term Karatsuba low-order algorithm to the multiplier vectors \(A'\) and \(B'\) to obtain \(e\) LA and \(e\) LB . Then in the multiplication stage, it is necessary to traverse \(i\) to calculate all \(e\) LA,i \(e\) LB,i . Among them, each element of \(e\) LA and \(e\) LB has \(k2\) sub-fragments, and the Karatsuba algorithm can be further used for acceleration.
[0144] If \(e\) LA,i \(e\) LB,i needs to perform \(k2m\)-bit full multiplication, the full transformation matrix can be applied for calculation, which requires times of full multiplication with a bit width of \(m\). For simplicity, it can be assumed that all \(F\) (k1) times of full multiplication are implemented using the same transformation matrix. In this way, times of \(m\)-bit full multiplication operations are required.
[0145] If \(e\) LA,i \(e\) LB,i only requires the semi-multiplication result, then use the low-order transformation matrix Performing the Karatsuba low-order algorithm for k2 terms can achieve better performance. Since the k2m-bit semi-multiplication inherently requires times of full multiplications, and times of semi-multiplications, all times of semi-multiplications altogether require times of m-bit full multiplications, as well as times of m-bit semi-multiplications.
[0146] Generally speaking, if k is first decomposed into k1 words, and each word is further decomposed into k2 sub-fragments, and the transformation matrices obtained by successively applying the Karatsuba algorithm The computational cost is:
[0147]
[0148] In order to explore a lower-order evaluation matrix with fewer effective multiplication computations Q (k) the decomposition order of the two factors can be exchanged to determine the change in the computational cost Q (k) before and after the exchange order, and based on this, study the relationship between the factor order and the computational cost.
[0149] Without loss of generality, assume that the target number of terms k can be decomposed into k α , k i , k i+1 , k β , where k α and k β can be 1 or can be further decomposed. For example:
[0150] k α = k1k2…k i-1
[0151] k β = k i+2 k i+3 …k q (24)
[0152] According to the above decomposition order, the low-order evaluation matrix corresponding to the number of terms k can be calculated by the following formula:
[0153]
[0154] Expanding term by term according to formula (23), the computational cost corresponding to formula (25) can be obtained as:
[0155]
[0156] Assume that the order of k i and k i+1 is exchanged, that is, k i+1 is in the front and k iAfter that, similarly expanding according to formula (23), the computational cost after swapping the order can be obtained as follows:
[0157]
[0158] Mathematically, it can be easily deduced that the equivalent condition for the cost after swapping shown in formula (27) to be less than the cost before swapping shown in (26):
[0159]
[0160] The reasoning of formula (28) means that if is greater than then the order of k i and k i+1 should be swapped to obtain a smaller computational cost; otherwise, retaining the current order can obtain a smaller computational cost.
[0161] Therefore, for any factor k i a cost coefficient can be defined which is the ratio of the first number and the second number where the first number is the reduction number of the first multiplication times involved in modular multiplication using the k i -term Karatsuba algorithm compared to the second multiplication times involved in performing a conventional multiplication, and the second number is the number of times of performing operations using a full multiplication circuit in modular multiplication using the k i -term Karatsuba algorithm.
[0162] Each factor can be sorted in reverse order according to the magnitude of the cost coefficient, that is, the factor with a larger cost coefficient is sorted in the front. In other words, when the factorization of the number of terms k contains two factors, the first factor and the second factor, if the cost coefficient corresponding to the first factor is greater than that of the second factor, then the first factor is ranked in the front, for example, as the first number of terms k1 in formula (20), and the second factor is ranked in the back, as the second number of terms k2. Conversely, if the cost coefficient of the first factor is less than that of the second factor, then the second factor is used as the first number of terms k1 in the front, and the first factor is used as the second number of terms k2 in the back.
[0163] For example, k = 8 can be factored into 4 * 2, or it can be factored into 2 * 4. Since R (2) / F (2) = 0, while R (4) / F (4)= 0.2, so 4*2 with factor 4 placed in front is a better decomposition scheme. Similarly, for the case of k = 2*3*4 = 24, through the sorting of the cost coefficients of each factor, it can be quickly determined that the has a lower computational cost.
[0164]
[0165] The costs (T (k) , F (k) , H (k) , R (k) ) of each term with a low order number (e.g., 1 - 4 orders) can be easily determined and preset. Based on this, the cost coefficients of each factor of the term number k can be quickly obtained, so as to quickly determine the sorting of the factors, so as to obtain a low-order evaluation matrix with a lower computational cost in a recursive manner Compared with the method of traversing various sorting combinations to determine the corresponding computational cost, such a method is obviously more efficient. In this way, the low-order evaluation matrix can be efficiently determined as the configuration parameter of the modulo multiplication circuit.
[0166] As mentioned above, the high-order evaluation matrix applicable to the rounding multiplication circuit has a cost and computational logic similar to or corresponding to that of the low-order evaluation matrix. Figure 6 FIG. shows a flowchart of a method for determining the processing parameters of a rounding multiplication circuit according to an embodiment. This method can be executed by any circuit, device, platform with computational processing capabilities. As Figure 6 shown, this method includes the following steps.
[0167] In step 62, factorize the target term number k to be processed, and determine the first term number k1 and the second term number k2 from the decomposed factors.
[0168] In step 64, according to the k1-order high-order evaluation matrix corresponding to the first term number k1 the k2-order high-order evaluation matrix corresponding to the second term number k2 and the k2-order full-scale evaluation matrix determine the k-order high-order evaluation matrix corresponding to the target term number k wherein, the p-order high-order / full-scale evaluation matrix corresponding to any term number p is used to respectively apply to the multiplier vectors composed of p multiplier segments corresponding to each of the two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the high-order / full-scale segments of the product; the k-order high-order evaluation matrix is used as the processing parameter of the rounding multiplication circuit for rounding the k-th power of r.
[0169] More specifically, referring to formula (20), the k-order high-order evaluation matrix It can be calculated accordingly through the following formula:
[0170]
[0171] That is to say, first divide the k1-order high-order evaluation matrix into the first matrix part composed of target rows and the second matrix part composed of the remaining rows Among them, the element combination between the elements generated by the target rows in the two intermediate vectors is realized through a semi-multiplication circuit. Then, the second matrix part is subjected to a Kronecker product operation with the k2-order full evaluation matrix to obtain the first result part; the first matrix part is subjected to a Kronecker product operation with the k2-order high-order evaluation matrix to obtain the second result part; the concatenation of the first result part and the second result part forms a result matrix to form a k-order high-order evaluation matrix
[0172] Determine the factorization order on which the high-order evaluation matrix depends in a recursive manner, which is determined by the above sorting based on the cost coefficient. In this way, the high-order evaluation matrix can be efficiently determined as the configuration parameter of the rounding multiplication circuit.
[0173] In one embodiment, the above rounding multiplication circuit and modulo multiplication circuit are sub-circuits in the modular multiplication circuit. For example, they can respectively correspond to Figure 4 Circuit 12 and Circuit 13 in the modular multiplication circuit shown. In other embodiments, the rounding multiplication circuit and the modulo multiplication circuit can also be independent circuits dedicated to performing rounding multiplication operations and modulo multiplication operations respectively.
[0174] According to an embodiment of another aspect, a device for determining the processing parameters of a modulo multiplication circuit is provided. Figure 7 The structural schematic diagram of a modulo multiplication parameter determination device according to an embodiment is shown. This device can be deployed in any device, platform or device cluster with data storage, computing, and processing capabilities. As Figure 7 shown, the modulo multiplication parameter determination device 700 includes:
[0175] The first decomposition unit 72 is configured to perform factorization on the target number of items k to be processed, and determine the first number of items k1 and the second number of items k2 from the decomposed factors;
[0176] The first determination unit 74 is configured to determine a k - order low - order evaluation matrix corresponding to the target number of terms k according to a k1 - order low - order evaluation matrix corresponding to the first number of terms k1, a k2 - order low - order evaluation matrix corresponding to the second number of terms k2, and a k2 - order full - amount evaluation matrix. Wherein, a p - order low - order / full - amount evaluation matrix corresponding to any number of terms p is used to be respectively applied to multiplier vectors formed by p multiplier segments corresponding to two multipliers to obtain two intermediate vectors, and the combination of the elements of the two intermediate vectors is used to obtain a low - order / full - amount segment of the product; the k - order low - order evaluation matrix is used as a processing parameter of a modulo - multiplication circuit for taking the modulo of r to the k - th power.
[0177] According to an embodiment of yet another aspect, a device for determining processing parameters of a rounding multiplication circuit is provided. Figure 8 The structural schematic diagram of a rounding multiplication parameter determination device according to an embodiment is shown. This device can be deployed in any device, platform or device cluster with data storage, computing, and processing capabilities. As Figure 8 shown, the rounding multiplication parameter determination device 800 includes:
[0178] The second decomposition unit 82 is configured to perform factorization on the target number of terms k to be processed, and determine the first number of terms k1 and the second number of terms k2 from the decomposed factors;
[0179] The second determination unit 84 is configured to determine a k - order high - order evaluation matrix corresponding to the target number of terms k according to a k1 - order high - order evaluation matrix corresponding to the first number of terms k1, a k2 - order high - order evaluation matrix corresponding to the second number of terms k2, and a k2 - order full - amount evaluation matrix. Wherein, a p - order high - order / full - amount evaluation matrix corresponding to any number of terms p is used to be respectively applied to multiplier vectors formed by p multiplier segments corresponding to two multipliers to obtain two intermediate vectors, and the combination of the elements of the two intermediate vectors is used to obtain a high - order / full - amount segment of the product; the k - order high - order evaluation matrix is used as a processing parameter of a rounding multiplication circuit for rounding r to the k - th power.
[0180] For the implementation manners of each unit in the above device, reference can be made to the description in combination with the method embodiments. Through the above device, the k - order low - order evaluation matrix can be quickly determined as the parameter of the modulo - multiplication circuit, and the k - order high - order evaluation matrix can be determined as the processing parameter of the rounding multiplication circuit, saving the cost of parameter configuration for the modulo - multiplication circuit and the rounding multiplication circuit.
[0181] According to an embodiment of another aspect, a computer - readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in combination with Figure 5 or / or Figure 6 described.
[0182] According to an embodiment of another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, it implements the method in combination with Figure 5 and / or Figure 6 described above.
[0183] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0184] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for determining processing parameters of a modular multiplication circuit, comprising: Factorizing a target number of items k to be processed, and determining a first number of items k1 and a second number of items k2 from the decomposed factors; Determining a k - order low - order evaluation matrix corresponding to the target number of items k according to the k1 - order low - order evaluation matrix corresponding to the first number of items k1, the k2 - order low - order evaluation matrix corresponding to the second number of items k2, and the k2 - order full - volume evaluation matrix, wherein for any number of items p, the p - order low - order / full - volume evaluation matrix is used to apply to multiplier vectors respectively formed by p multiplier segments corresponding to two multipliers to obtain two intermediate vectors, and the combination of elements of the two intermediate vectors is used to obtain the low - order / full - volume segment of the product; the k - order low - order evaluation matrix is used as the processing parameter of the modular multiplication circuit for taking the modulus of r to the k - th power.
2. The method according to claim 1, wherein, Determining the k - order low - order evaluation matrix corresponding to the target number of items k includes: Dividing the k1 - order low - order evaluation matrix into a first matrix part composed of target rows and a second matrix part composed of the remaining rows; wherein, the combination of elements between the elements generated by the target rows in the two intermediate vectors is implemented by a semi - multiplication circuit; Performing a Kronecker product operation on the second matrix part and the k2 - order full - volume evaluation matrix to obtain a first result part; Performing a Kronecker product operation on the first matrix part and the k2 - order low - order evaluation matrix to obtain a second result part; the splicing of the first result part and the second result part forms a result matrix, and the result matrix is used to form the k - order low - order evaluation matrix.
3. The method according to claim 1, wherein, Determining the first number of items k1 and the second number of items k2 from the decomposed factors includes: Determining a first factor and a second factor from the decomposed factors; For any factor ki, determining the ratio of a first number to a second number as a cost coefficient, where the first number is the reduced number of the first multiplication times involved in the modular multiplication operation using the ki - term Karatsuba algorithm relative to the second multiplication times involved in the conventional multiplication operation; the second number is the number of times of performing operations using a full - multiplication circuit in the modular multiplication operation using the ki - term Karatsuba algorithm; If the cost coefficient of the first factor is greater than that of the second factor, determining the first factor as the first number of items k1 and the second factor as the second number of items k2; otherwise, determining the second factor as the first number of items k1 and the first factor as the second number of items k2.
4. The method according to claim 1, wherein, Any element in the k1 - order and k2 - order low - order evaluation matrices is selected from 0, 1, and - 1.
5. According to the method described in claim 1, it further includes using the k - order low - order evaluation matrix as the processing parameter corresponding to the base r and the number of items k, and configuring the parameters of the modular multiplication circuit.
6. The method according to claim 1, wherein The modular multiplication circuit is a sub - circuit in a modular multiplication circuit that uses the Barrett modular reduction method for modular multiplication.
7. A method for determining processing parameters of a rounding multiplication circuit, comprising: Factorizing a target number of items k to be processed, and determining a first number of items k1 and a second number of items k2 from the decomposed factors; Based on the k1-order high-order evaluation matrix corresponding to the first number k1, the k2-order high-order evaluation matrix corresponding to the second number k2, and the k2-order full-scale evaluation matrix, determine the k-order high-order evaluation matrix corresponding to the target number k. Among them, the p-order high-order / full-scale evaluation matrix corresponding to any number p is used to respectively apply to the multiplier vectors composed of p multiplier segments corresponding to the two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the high-order / full-scale segments of the product; the k-order high-order evaluation matrix is used as the processing parameter of the rounding multiplication circuit for rounding the k-th power of r.
8. The method according to claim 1, wherein Determining the k-order high-order evaluation matrix corresponding to the target number k includes: Divide the k1-order high-order evaluation matrix into a first matrix part composed of target rows and a second matrix part composed of the remaining rows; among them, the element combination between the elements generated by the target rows in the two intermediate vectors is realized by a semi-multiplication circuit; Perform a Kronecker product operation on the second matrix part and the k2-order full-scale evaluation matrix to obtain a first result part; Perform a Kronecker product operation on the first matrix part and the k2-order high-order evaluation matrix to obtain a second result part; the splicing of the first result part and the second result part forms a result matrix, and the result matrix is used to form the k-order high-order evaluation matrix.
9. The method according to claim 1, wherein, Determining the first number k1 and the second number k2 from the decomposed factors includes: Determine the first factor and the second factor from the decomposed factors; For any factor ki, determine the ratio of the first number to the second number as the cost coefficient, where the first number is the reduction number of the first multiplication times involved in the rounding multiplication operation using the ki-term Karatsuba algorithm compared to the second multiplication times involved in the conventional multiplication operation; the second number is the number of times of using the full-multiplication circuit in the rounding multiplication operation using the ki-term Karatsuba algorithm; If the cost coefficient of the first factor is greater than that of the second factor, determine the first factor as the first number k1 and the second factor as the second number k2; otherwise, determine the second factor as the first number k1 and the first factor as the second number k2.
10. A device for determining the processing parameters of a modulo multiplication circuit, including: A first decomposition unit configured to perform factor decomposition on the target number k to be processed, and determine the first number k1 and the second number k2 from the decomposed factors; A first determination unit configured to determine the k-order low-order evaluation matrix corresponding to the target number k according to the k1-order low-order evaluation matrix corresponding to the first number k1, the k2-order low-order evaluation matrix corresponding to the second number k2, and the k2-order full-scale evaluation matrix. Among them, the p-order low-order / full-scale evaluation matrix corresponding to any number p is used to respectively apply to the multiplier vectors composed of p multiplier segments corresponding to the two multipliers to obtain two intermediate vectors, and the element combinations of the two intermediate vectors are used to obtain the low-order / full-scale segments of the product; the k-order low-order evaluation matrix is used as the processing parameter of the modulo multiplication circuit for taking the modulo of the k-th power of r.
11. A device for determining the processing parameters of a rounding multiplication circuit, including: A second decomposition unit, configured to perform factorization on a target number of items k to be processed, and determine a first number of items k1 and a second number of items k2 from the decomposed factors; A second determination unit, configured to determine a k-th order high-order evaluation matrix corresponding to the target number of items k according to a k1-th order high-order evaluation matrix corresponding to the first number of items k1, a k2-th order high-order evaluation matrix corresponding to the second number of items k2, and a k2-th order full-amount evaluation matrix, wherein a p-th order high-order / full-amount evaluation matrix corresponding to any number of items p is used to respectively apply to a multiplier vector composed of p multiplier segments corresponding to each of the two multipliers to obtain two intermediate vectors, and the element combination of the two intermediate vectors is used to obtain a high-order / full-amount segment of the product; the k-th order high-order evaluation matrix is used as a processing parameter of a rounding multiplication circuit for rounding the k-th power of r.
12. A computing device, comprising a memory and a processor, characterized in that, An executable code is stored in the memory, and when the processor executes the executable code, the method described in any one of claims 1-9 is implemented.