Sparse data feature-based number-theory transformation acceleration method
Through the number theory transformation acceleration method based on sparse data features, the number theory transformation calculation in homomorphic convolution is optimized, and the problem of high computing overhead in the existing technology is solved, and more efficient privacy neural network reasoning is achieved.
Patent Information
- Application Number
- CN202510210465.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the inference calculation overhead of privacy neural network based on hybrid HE/2PC schemes is high, especially the number-theoretical transformation and inverse transformation calculation in homomorphic convolution, resulting in memory consumption and calculation delay problems.
A number theory transformation acceleration method based on sparse data features is proposed. Through bit inversion, pre-calculation, performing optimization calculations and subsequent butterfly network calculations, the number of multiplication and additions in the butterfly network is reduced by using sparseness.
By optimizing the pre-level calculation data flow of the butterfly network, skipping unnecessary calculations and merging multi-level calculations, the number of multiplication and additions required in number theory transformation is significantly reduced, the calculation delay is reduced, and the overall energy efficiency of privacy reasoning is improved.
Smart Images

Figure CN120145443A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of accelerator system design in privacy computing, and particularly relates to a number-theoretic transform acceleration method for accelerating polynomial multiplication in privacy computing. Background Art
[0002] Privacy protection has become one of the main concerns when deploying deep neural networks (DNNs) in the cloud. In recent years, homomorphic encryption (HE) has been proposed and attracted extensive attention. By encrypting data into ciphertext polynomials, HE allows computations to be performed on encrypted data without revealing any information about the data itself. To apply HE to private DNN inference, there are mainly two methods: fully homomorphic encryption (FHE) schemes and hybrid HE / two-party computation (2PC) schemes. The main difference between the two lies in the implementation of non-linear activation functions. The hybrid scheme uses a two-party secure computation protocol, which helps to avoid activation function approximation and high-cost bootstrapping operations in the FHE scheme. Therefore, the present invention will focus on optimizing the hybrid scheme.
[0003] Privacy neural network inference based on the hybrid HE / 2PC scheme is still much slower than plaintext inference, mainly because of the high computational overhead. Compared with the computation of non-linear layers, homomorphic convolution (HConv) in linear layers becomes the main computational bottleneck, and its computational process is as Figure 1 shown. The client has the input activation vector X of the convolutional layer, encodes it into a polynomial and encrypts it, and then sends it to the server for computation. The server has the weight vector W, encodes it into a polynomial and computes it with the received ciphertext, and then sends it back to the client for decryption and post-processing. In this way, the computed data can be obtained at the client without leaking the input data, and the client will not know the neural network-related information of the server. In particular, the polynomial computation of the weight W and X is accelerated by the number-theoretic transform (NTT). The reason for the large computational overhead of homomorphic convolution lies in the repeated computation of a large number of number-theoretic transforms and their inverse transforms (inverse number-theoretic transform, INTT), especially in the computation of weight polynomials. The basic idea of the number-theoretic transform is to reduce the complexity of multiplying two N-degree polynomials from O(N 2 ) to O(N log N) through a divide-and-conquer method. The specific computation is implemented through an N-point butterfly network, and a large number of multiply-add computations are involved in the butterfly network, and all are modular operations. Figure 2Shows the butterfly network schematic diagram of the 8-point number-theoretic transform when N = 8, which includes two parts: bit-reversal and butterfly network calculation. Each unit of the butterfly network contains one multiplication calculation and two addition / subtraction calculations. For a more general N-point number-theoretic transform, its input is a sequence of length N composed of N integers modulo q, and bit-reversal (bit-reverse) needs to be performed first, followed by log 2 N-level calculations. Each level contains N / 2 butterfly units. Inside each butterfly unit, pairwise modular multiplication and modular addition (subtraction) calculations are performed on the data. As Figure 2 shown, the intermediate data m[i] and m[j] and the rotation factor (W N is a primitive root) perform butterfly calculations to obtain M[i] and M[j]:
[0004]
[0005] Although the weight polynomials in the NTT domain can be pre-computed and stored, this will result in significant memory overhead. For example, storing all the weights of a 4-bit quantized ResNet-50 in the NTT domain requires 23 GB of memory, which results in a memory consumption more than 1000 times higher than the normal case.
[0006] In recent years, researchers have proposed different homomorphic encryption (HE) accelerators to accelerate the computationally expensive number-theoretic transform (NTT). Some studies focus on optimizing the data flow and avoiding stalls between stages to improve parallelism [1][2]. Others achieve acceleration by decomposing large-scale NTTs into multiple small NTTs [3]. Although these accelerators have made promising progress in acceleration, they usually face high area and power consumption costs. For example, the area of the accelerator proposed in [3] exceeds 150 mm 2 , and the power consumption is about 100 W. It is found that the input of the number-theoretic transform of neural network weights has a very high sparsity, up to 90%, but the existing work does not take into account the characteristics of the input data of the number-theoretic transform. Therefore, it is of great significance to study how to utilize the input characteristics of the number-theoretic transform of network model weights to improve the computational efficiency of polynomial modular multiplication.
[0007] [1] M.S. Riazi, K. Laine, B. Pelton, and W. Dai, “HEAX: An architecture for computing on encrypted data,” in Proceedings of the twenty-fifth international conference on architectural support for programming languages and operating systems, 2020, pp. 1295–1309.
[0008] [2] X. Ren, Z. Chen, Z. Gu, Y. Lu, R. Zhong, W.-J. Lu, J. Zhang, Y. Zhang, H. Wu, X. Zheng et al., “CHAM: A customized homomorphic encryption accelerator for fast matrix-vector product,” in 2023 60th ACM / IEEE Design Automation Conference (DAC). IEEE, 2023, pp. 1–6.
[0009] [3] N. Samardzic, A. Feldmann, A. Krastev, S. Devadas, R. Dreslinski, C. Peikert, and D. Sanchez, “F1: A fast and programmable accelerator for fully homomorphic encryption,” in MICRO-54: 54th Annual IEEE / ACM International Symposium on Microarchitecture, 2021, pp. 238–252. Summary of the Invention
[0010] In view of the problems existing in the above prior art, the present invention proposes a method for accelerating number-theoretic transform based on sparsified data features, also known as sparse butterfly network computing acceleration. The present invention utilizes the feature of sparsifying the input data of number-theoretic transform in homomorphic convolution, and through data flow design, avoids unnecessary calculations, thereby improving computational energy efficiency. The present invention designs corresponding optimization methods according to the data distribution type after bit-reversal in number-theoretic transform, so as to achieve efficient acceleration of homomorphic convolution. The specific process of realizing sparse number-theoretic transform acceleration of the present invention will be introduced below.
[0011] The technical solution of the present invention is as follows:
[0012] A method for accelerating number-theoretic transform based on sparsified data features. In homomorphic convolution, the height and width of the input activation vector X of the current convolution layer are denoted as H, and the pre-processed sequence of the weight vector W is accelerated by this method for number-theoretic transform. The input sequence of the number-theoretic transform NTT is denoted as m, and its length is N, N = 2 k , k is the bit width of the index, and the index of the sequence ranges from 0 to N - 1; it is characterized in that, by utilizing the sparsity of the input data in number-theoretic transform, the number of multiply-accumulate calculations in the butterfly network is reduced, and the overall process includes four steps: bit-reversal, pre-computation, execution of optimized calculation, and execution of subsequent butterfly network calculation; the specific steps are as follows:
[0013] The first step: Bit-reversal, rearrange the indices of the input sequence m in the way of binary bit-reversal to obtain the sequence m br , and the index list IndicesLst of the non-zero values in m br ;
[0014] The second step: Pre-computation, pre-compute and store the optimization parameters according to H and IndicesLst;
[0015] When H is a power of 2, the non-zero values in the bit-reversed sequence m br are distributed in a concentrated data distribution: there are only continuous N con non-zero values within each interval of length N skip , and the rest of the values are all 0. Every N con numbers are recorded as a group, and there are N group = N / N con groups. Calculate the number of replication times N dup = N con / N skip , and at this time, the corresponding optimization method is skipped, and no additional pre-computation is performed, but directly enter the third step to execute the optimized calculation;
[0016] When H is not a power of 2, the non-zero values in the bit-reversed sequence m br are distributed in a scattered data distribution: there are N conThere is only one non - zero value within the interval, and N within the current interval con The N - point NTT butterfly operations are combined into one level for execution. At this time, the corresponding merging optimization method is used; pre - calculate the merging intervals and the rotation factors after merging to obtain the list of interval lengths dxIntervalLen, and the interval length N mrg = min(IdxIntervalLen), the number of merged rotation factors N comp
[0017] , the number of copying times N con = N mrg , the number of groups that need to be calculated is N group = N / N con , the list of exponents ExpSet of the rotation factors after merging relative to the primitive root and the list of sign bits SignSet after merging
[0018] Step 3: Execute the optimized calculation
[0019] 1) Select the optimization method for calculation
[0020] When H is a power of 2, skip the optimization method and calculate the butterfly network with the normal input number N skip to obtain the N skip - point butterfly network calculation result m'[k], k = 0,..., N skip - 1, and the pre - calculated N dup are passed to the next step together, and record N opt = N skip ;
[0021] When H is not a power of 2, adopt the merging optimization method. For the i - th value idx in IndicesLst, the m br [idx] in the corresponding sequence, combined with the N comp obtained in the pre - calculation Calculate the first N comp results after multi - level merging k = 0,..., N comp - 1, take its negation to obtain the calculation results of consecutive 2N comp , that is, m'[k + N comp = - m'[k], k = 0,..., N comp - 1, and the N dup in the pre - calculation are passed to the next step together, and record N opt = 2N comp ;
[0022] 2) Copy
[0023] Duplicate m'[k], k = 0, …, N obtained in the previous step opt N copies, where N dup = N opt or 2N skip . After duplication, we get M'[k + l×N comp = m'[k], k = 0, …, N opt -1, l = 0, …, N opt -1; At this point, we obtain the calculation results of M'[j], j = 0, …, N dup ; con The calculation results of
[0024] 3) Repeat the above two steps for each optimization interval of the N group groups to obtain all N calculation results M′[k], k = 0, …, N - 1 in the M' stage;
[0025] Fourth step: Perform subsequent butterfly network calculations
[0026] For the calculations from the M' stage to the final M stage, just perform the subsequent calculations of the N-point butterfly network to obtain the final output of M[k], k = 0, …, N - 1, which is the final output of the NTT.
[0027] Furthermore, for the number-theoretic transform acceleration method based on sparse data characteristics, the first step of bit reversal is specifically: represent each index of the input sequence m as a k-bit binary number, reverse the binary representation of each index, that is, invert the bit order of the binary number, and then rearrange the input sequence according to the reversed index order to generate a new sequence m br , and obtain the index list IndicesLst of its non-zero values. When H is a power of 2, the non-zero values in the sequence m br are distributed in a concentrated data distribution, and when H is not a power of 2, the non-zero values in the sequence m br are distributed in a dispersed data distribution.
[0028] Furthermore, for the number-theoretic transform acceleration method based on sparse data characteristics, in the second step of pre-computation, when H is not a power of 2, pre-compute the merging intervals and the merged rotation factors, specifically:
[0029] a) Determine the merging intervals
[0030] For the i-th value in IndicesLst, that is, the index of the i-th non-zero value after bit reversal, perform the following operations to determine the mergeable interval corresponding to this index value;
[0031] First, denote idx = IndicesLst[i], idx -1= IndicsLst[i - 1], idx +1 = IndicesLst[i + 1], and idx and idx -1 are represented in binary form. The binary form of idx has k bits. Find the highest non - identical bit between the binary forms of idx and idx -1 , denoted as the x - th bit. That is, the [k - 1:x + 1] bits of the binary representations of idx and idx -1 are the same, while idx[x] ≠ idx -1 [x]; Extract the high - order bits [k - 1:x] of idx and fill the low x bits with 0 to get IdxStart = {idx[k - 1:x], x'b0}, which is used as the starting value of the mergeable interval; Similarly, for idx +1 Performing the above calculations gives Idx +1 Start; The ending value of the current interval IdxEnd = Idx +1 Start - 1; Calculate the interval length and store it in a list, that is, IdxIntervalLen[i] = IdxEnd - IdxStart + 1;
[0032] Traverse each element in IndicesLst. After performing the above operations, a list IdxIntervalLen of the same length as IndicesLst is obtained, where each element corresponds to the length of the mergeable interval for each index value; Finally, determine that the length of the mergeable interval for the entire butterfly network is N mrg = min(IdxIntervalLen); Denote N con = N mrg , N group = N / N con ;
[0033] b) Calculate the merged twiddle factors, and calculate the exponent and sign bit of the merged twiddle factor relative to the primitive root;
[0034] For the i - th value in IndicesLst, denote the relative position of this index within the mergeable interval as i rel = IndicesLst[i] - IdxStart, and perform the following operations respectively to determine the merged twiddle factor corresponding to each index value;
[0035] First, the number of merged twiddle factors to be calculated is At the same time, obtain the number of times N to be copied in subsequent calculations dup = N mrg / 2N comp ; The exponent of the twiddle factor relative to the primitive root is the sum of all exponents TwiddleExp[j] on the merge path, that is Denoted as k = 0, 1, ... N comp -1; Meanwhile, the sign bit of the combined rotation factor is the product of the sign bits sign[j] of all the rotation factors on the combination path, denoted as k = 0, 1, ... N comp -1; Therefore, the N combined rotation factors corresponding to each index value IndicesLst[i] are recorded as a list comp whose sign bit is a list of the same length After traversing each element in IndicesLst and performing the above operations, lists ExpSet and SignSet of the same length as IndicesLst are obtained, where each element ExpSet
[0036] and SignSet i are both a list of length N i comp
[0037] The technical effects of the present invention are as follows:
[0038] A number-theoretic transform acceleration method based on sparse data features provided by the present invention utilizes the input sparsity of the number-theoretic transform in privacy inference. According to the data distribution type after bit reversal in the number-theoretic transform, corresponding optimization methods are designed to optimize the data flow of the pre-stage calculation of the butterfly network, skip unnecessary calculations, and combine multi-stage calculations, greatly reducing the number of multiplications and additions required in the number-theoretic transform, reducing the calculation latency, and improving the overall energy efficiency of privacy inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of a homomorphic convolution implemented based on a number-theoretic transform in a hybrid HE / 2PC encryption scheme;
[0040] Figure 2 is a schematic diagram of the algorithm implementation of an 8-point number-theoretic transform, which is the algorithm implementation of the butterfly network of the number-theoretic transform;
[0041] Figure 3 is a flowchart of the number-theoretic transform acceleration method based on sparse data features proposed by the present invention;
[0042] Figure 4 is a schematic diagram of the centralized data distribution and its corresponding skip optimization method in the embodiment of the present invention;
[0043] Figure 5 is a schematic diagram of the dispersed data distribution and its corresponding combination optimization method in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The present invention will be further clearly and completely described below in conjunction with the accompanying drawings and through specific embodiments.
[0045] The present invention proposes a number-theoretic transform acceleration method based on sparse data features, which utilizes the sparsity of the input data in the number-theoretic transform to reduce the number of multiplications and additions in the butterfly network. The overall process is as Figure 3 shown, mainly including four steps: bit reversal, pre-computation, execution of optimized calculation, and execution of subsequent butterfly network calculation; in the homomorphic convolution as Figure 1 shown, the height and width of the input activation vector X of the current convolution layer are denoted as H. After the weight vector W undergoes pre-processing and other operations, the input to the number-theoretic transform NTT is a sequence m of length N, where N = 2 k , k is the bit width of the index, and the index of the sequence ranges from 0 to N - 1; the acceleration method of the number-theoretic transform includes the following steps:
[0046] The first step: Bit reversal, re-arrange the indices of the input sequence m in the way of binary bit reversal to obtain the sequence m br ;
[0047] The goal of bit reversal is to re-arrange the indices of the input sequence in the way of binary bit reversal for the convenience of subsequent butterfly network calculation. Represent each index of the input sequence m as a k-bit binary number, reverse the binary representation of each index, that is, invert the bit order of the binary number, and then re-arrange the input sequence according to the reversed index order to generate a new sequence m br , and obtain the index list IndicesLst of the non-zero values therein. When H is a power of 2, the non-zero values in the sequence m br are distributed in a concentrated data distribution, as Figure 4 shown. When H is not a power of 2, the non-zero values in the sequence m br are distributed in a scattered data distribution, as Figure 5 shown.
[0048] As Figure 2 shown, the input sequence m contains N = 8 values, corresponding to 0 to 7 index values, and the bit width of each index value is k = log 2 N = 3 bits. Reverse the binary representation of each index value. For example, the index of m[4] is 4 = (100) 2 , and after reversal, it is (001) 2 = 1, that is, m br [1] = m[4]. Based on this, the sequence m br after bit reversal can be obtained. Further, the index list IndicesLst of the non-zero values in m br can be obtained.
[0049] Step 2: Pre - calculation. According to H and IndicesLst, pre - calculate the optimization parameters and store them for subsequent reuse;
[0050] When H is a power of 2, the non - zero value distribution in the bit - reversed sequence m br is a concentrated data distribution: within each interval of length N com there are only consecutive N skip non - zero values, and the rest of the values are 0. Every N con numbers are recorded as a group, and there are N group = N / N con groups. For example, Figure 4 as shown in con N = 16, N skip = 4. At this time, the corresponding optimization method is skipped. Calculate the number of replication times N dup = N con / N skip = 4. Instead of performing additional pre - calculation, directly enter Step 3 to perform the optimization calculation;
[0051] When H is not a power of 2, the non - zero value distribution in the bit - reversed sequence m br is a dispersed data distribution: within each interval of length N con there is only one non - zero value. The N con point NTT butterfly operations in the current interval can be merged into one level, corresponding to the merging optimization method, and it is necessary to pre - calculate the merging interval and the rotation factor after merging; specifically:
[0052] a) Determine the merging interval
[0053] For the i - th value in IndicesLst, that is, the index of the i - th non - zero value after bit - reversal, perform the following operations to determine the mergeable interval corresponding to this index value;
[0054] First, denote idx = IndicesLst[i], idx -1 = IndicesLst[i - 1], idx +1 = IndicesLst[i + 1]. Represent idx and idx -1 in binary form. As mentioned before, the binary form of idx has k bits. Find the highest non - identical bit in the binary forms of idx and idx -1 , denoted as the x - th bit, that is, in the binary representation, the [k - 1:x + 1] bits of idx and idx -1 are the same, and idx[x] ≠ idx -1[x]. Extract the high-order part [k - 1:x] of idx and fill the low x bits with 0 to obtain IdxStart = {idx[k - 1:x], x'b0}, which serves as the starting value of the mergeable interval. Similarly, for idx +1 Performing the above calculations yields Idx +1 Start. The ending value of the current interval IdxEnd = Idx +1 Start - 1. Calculate the interval length and store it in the list, i.e., IdxIntervalLen[i] = IdxEnd - IdxStart + 1.
[0055] After traversing each element in IndicesLst and performing the above operations, a list IdxIntervalLen of the same length as IndicesLst can be obtained, where each element corresponds to the length of the mergeable interval for each index value. Finally, determine that the length of the mergeable interval for the entire butterfly network is N mrg = min(IdxIntervalLen). Denote N con = N mrg , N group = N / N con .
[0056] For example, as Figure 5 shown, taking the sequence length N = 2 6 = 64 as an example, k = 6, IndicesLst[0] = 6, IndicesLst[1] = 22, Indices[2] = 49. For the i = 1st value, calculate its mergeable interval. First, denote idx = Indices[1] = 22 = (010110) 2 , idx -1 = IndicesLst[1 - 1] = 6 = (000110) 2 , idx +1 = Indiceslst[1 + 1] = 36 = (100100) 2 , the 5th bit of idx is the same as that of idx -1 , and the 4th bit is different. Extract the idx[5:4] bits and fill the low 4 bits with 0 to obtain IdxStart = (010000) 2 = 16, which is the starting position of the mergeable interval. Similarly, Idx +1 Start = 32. Therefore, the ending value of the current interval IdxEnd = Idx +1Start-1 = 31, and the length of the mergeable interval is denoted as IdxIntervalLen[1] = 31 - 16 + 1 = 16. After traversing each element in IndicesLst and performing the above operations, a list IdxIntervalLen of the same length as IndicesLst can be obtained, where IdxIntervalLen[0] = 16, IdxIntervalLen[1] = 16, IdxIntervalLen[2] = 32. Finally, it is determined that the length of the mergeable interval of the entire butterfly network is N mrg = min(IdxIntervalLen) = 16. Denote N con = N mrg = 16, N group = N / N con = 4.
[0057] b) Calculate the combined twiddle factors
[0058] As Figure 2 shown, the twiddle factors in the butterfly network of NTT are all powers of the primitive root W N Therefore, the key to this step is to calculate the exponent of the combined twiddle factor relative to the primitive root. For the i-th value in IndicesLst, denote the relative position of this index within the merge interval as i rel = indicesLst[i] - IdxStart, and perform the following operations respectively to determine the combined twiddle factors corresponding to each index value. First, the number of combined twiddle factors to be calculated is At the same time, obtain the number of times N dup = N mrg / 2N comp . The exponent of the twiddle factor relative to the primitive root is the sum of all exponents TwiddleExp[j] on the merge path, that is Denoted as k = 0, 1,... N comp - 1; At the same time, because there are addition and subtraction operations in the butterfly calculation unit, there is a difference in the positive and negative signs of the combined twiddle factors. The sign bit of the combined twiddle factor is the product of the sign bits sign[j] of all twiddle factors on the merge path, denoted as k = 0, 1,... N comp - 1; Therefore, the N comp combined twiddle factors corresponding to each index value IndicesLst[i] are denoted as a list whose sign bit is a list of the same length
[0059] Traverse each element in IndicesLst. After performing the above operations, two lists ExpSet and SignSet with the same length as IndicesLst can be obtained. Each element in ExpSet i and SignSet i is a list of length N comp .
[0060] For example Figure 5 as shown, taking i = 0, Indices[0] = 6 as an example, i rel = 6. First, the number of combined rotation factors to be calculated is N dup = N mrg / 2N comp = 16 / (2 × 4) = 2. The exponent of each combined rotation factor corresponds to the sum of all exponents on the combined path, that is k = 0, 1, 2, 3, and the corresponding sign bits are k = 0, 1, 2, 3. For N group = 4 groups, perform the above calculations respectively to obtain the optimized parameter list ExpSet = {ExpSet 0 , ExpSet 1 , ExpSet 2}. Since all values from index 32 to 47 are 0 and no calculation is required, the length of the optimized parameter list is the same as the length of IndicesLst.
[0061] Step 3: Perform optimized calculations
[0062] 1) Select an optimization method for calculation
[0063] When H is a power of 2, skip the corresponding optimization method and normally calculate the butterfly network with N skip inputs to obtain the calculation results m'[k], k = 0,..., N skip - 1 of the N skip -point butterfly network, and pass it together with the pre-calculated N dup to the next step. Taking Figure 4 as an example, it is necessary to normally calculate the butterfly network with N skip = 4 inputs to obtain the calculation results m'[k], k = 0,..., 3, of the N skip -point butterfly network, and pass it together with the pre-calculated N dup to the next step, and record N opt = 4. At this time, as shown in Figure 4 , the 12 butterfly units (including 12 modular multiplications and 24 modular additions) in the dotted arrow part of the "4-point NTT" stage are effectively reduced.
[0064] When H is not a power of 2, the merging optimization method is adopted. For the i-th value idx in IndicesLst, the corresponding m br [idx] in the sequence, combined with the N comp , calculate the first N comp results after multi-level merging k = 0, …, N comp -1, take its negation to obtain the calculation results of consecutive 2N comp , that is, m′[k + N comp = -m′[k], k = 0, …, N comp -1, and the N dup in the pre-computation are passed to the next step together. Taking Figure 5 as an example, for the 0-th value idx = IndicesLst[0] = 6 in IndicesLst, the corresponding m br [6] in the sequence, combined with the N comp = 4, calculate the first N comp results after multi-level merging k = 0, …, 3, take its negation to obtain the calculation results of consecutive 2N comp , that is, m′[k + 4] = -m′[k], k = 0, …, 3. So far, m′[k], k = 0, …, 7 has been calculated, and the N dup = 2 in the pre-computation are passed to the next step, and denote N opt = 2N comp = 8. At this time, the 24 butterfly units (including 24 modular multiplications and 48 modular additions) in the dashed arrow part of the "merging calculation" stage shown in Figure 5 are replaced by 4 complex multiplications, greatly reducing the overall number of multiplication and addition calculations.
[0065] 2) Duplication
[0066] Duplicate the m′[k], k = 0, …, N opt -1 calculated in the previous step for N dup copies, where N opt = N skip or 2N comp . After duplication, we get M'[k + l×N opt = m′[k], k = 0, …, N opt -1, l = 0, …, N dup -1. So far, M'[j], j = 0, …, N conThe calculation result. In this step, like Figure 4 and Figure 5 the butterfly unit calculations in the "copy" stage shown are eliminated, so the number of multiplication and addition calculations is also greatly reduced.
[0067] 3) Repeat the above two steps for each optimization interval of the N group groups to obtain all N calculation results M′[k], k = 0, …, N−1 in the M' stage, as Figure 4 and 5 shown.
[0068] Step 4: Perform subsequent butterfly network calculations. For the calculations from the M' stage to the final M stage, just perform the subsequent calculations of the N-point butterfly network to obtain the final output of M[k], k = 0, …, N−1, that is, the final output of the NTT.
[0069] Through the above method, the number-theoretic transform of the weight polynomial in homomorphic convolution can be optimized. Since each layer of convolution only corresponds to one input image size, that is, the same H, then the IndicesLst of the weight polynomial is consistent, Figure 3 the "second-step pre-calculation" required in only needs to be calculated once and can be reused in the number-theoretic transform process of all weight polynomials in the current convolution layer. Therefore, the additional calculation overhead introduced is very small compared to the overall improvement in calculation efficiency brought by this method, but it can greatly reduce the number of multiplications and additions required in the number-theoretic transform, reduce the calculation latency, and improve the overall energy efficiency of homomorphic convolution.
[0070] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention is subject to the scope defined by the claims.
Claims
1. A method for accelerating number-theoretic transformation based on sparse data features. In homomorphic convolution, the height and width of the current convolution layer input activation vector X are recorded as H, and the sequence preprocessed by the weight vector W is accelerated by number-theoretic transformation through this method. The input sequence of number-theoretic transformation NTT is recorded as m, and its length is N, N = 2 k , k is the bit width of the index, and the index of the sequence is from 0 to N-1; it is characterized by, The sparsity of input data in number theory transformation is used to reduce the number of multiplication and addition calculations in the butterfly network. The overall process includes four steps: bit reversal, pre-calculation, optimization calculation, and subsequent butterfly network calculation. The specific steps are: Step 1: Bit reversal: rearrange the index of the input sequence m in a binary bit reversal manner to obtain sequence m br , and m br The index list IndicesLst of non-zero values in; Step 2: Pre-calculation: According to H and IndicesLst, pre-calculate the optimization parameters and store them; When H is a power of 2, the bit-reversed sequence m br The non-zero value distribution is the concentrated data distribution: the length of each segment is N con There are only N consecutive skip non-zero values, and the rest are 0. con The number is recorded as a group, with a total of N group =N / N con Group, calculate the number of replications N dup =N con / N skip , in this case, the optimization method is skipped, no additional pre-calculation is performed, and the optimization calculation is directly performed in the third step; When H is not a power of 2, the bit-reversed sequence m br The non-zero value distribution is a scattered data distribution: the length of each segment is N con There is only one non-zero value in the interval, and N in the current interval con The point NTT butterfly operation is merged into one level, which corresponds to the merge optimization method; pre-calculate the merge interval and the merged rotation factor to obtain the interval length list dxIntervalLen, the interval length N mrg = min(IdxIntervalLen), the number of combined rotation factors N comp , the number of copies N con =N mrg , the number of groups to be calculated is N group =N / N con , the index list ExpSet of the merged rotation factors relative to the original root and the merged sign bit list SignSet, Step 3: Perform optimization calculations 1) Select the optimization method for calculation When H is a power of 2, the optimization method is skipped and the number of normal calculation inputs is N. skip The butterfly network is N skip The calculation result of the butterfly network is m′[k], k=0,…,N skip -1, and the precomputed N dup Pass them to the next step and record N opt =N skip ; When H is not a power of 2, the merge optimization method is used. For the i-th value idx in IndicesLst, the corresponding m in the sequence br [idx], combined with N obtained in pre-calculation comp , Calculate the top N after merging multiple levels comp Results After inverting it, we get continuous 2N comp The calculation result is m′[k+Nc omp ]=-m′[k], k=0,…,N comp -1, and N in precalculation dup Pass them to the next step and record N opt =2N comp ; 2) Copy The m′[k] calculated in the previous step, k=0,…,N opt -1 Copy N dup , where N opt =N skip or 2N comp , after copying, we get M'[k+l×N opt ]=m′[k],k=0,…,N opt -1,l=0,…,N dup -1; So far, we get M′[j], j=0,…,N con The calculation results of 3) For N group Repeat the above two steps for each optimization interval of the group to obtain all N calculation results M′[k], k=0,…,N-1 in the M′ stage; Step 4: Perform subsequent butterfly network calculations For the calculation from the M' stage to the final M stage, the post-stage calculation of the N-point butterfly network is performed normally to obtain the final output M[k], k=0, ..., N-1, that is, the final output of NTT.
2. The number theory transformation acceleration method based on sparse data features as claimed in claim 1, characterized in that: The first step of bit reversal is to represent each index of the input sequence m as a k-bit binary number, reverse the binary representation of each index, that is, invert the bit order of the binary number, and then rearrange the input sequence according to the reversed index order to generate a new sequence m. br , and get the index list IndicesLst of the non-zero values. When H is a power of 2, the sequence m br The non-zero value distribution in the sequence m is the concentrated data distribution. When H is not a power of 2, the sequence m br The non-zero value distribution is a dispersed data distribution.
3. The number theory transformation acceleration method based on sparse data features as claimed in claim 1, characterized in that: In the second step of pre-calculation, when H is not a power of 2, the merged interval and the merged rotation factor are pre-calculated, specifically: a) Determine the merge interval For the i-th value in IndicesLst, that is, the index of the i-th non-zero value after bit reversal, perform the following operations to determine the mergeable interval corresponding to the index value; First, remember idx = IndicesLst[i], idx -1 =IndicesLst[i-1], idx +1 =IndicesLst[i+1], replace idx and idx -1 Represented in binary form, the binary form of idx has a total of k bits. Find idx and idx -1 The highest different bit in the binary form is recorded as the xth bit, that is, idx and idx in the binary form -1 The [k-1:x+1] bits of the -1 [x]; extract the high bit [k-1:x] of idx and fill the low bit x with 0, and get IdxStart = {idx[k-1:x], x′b0} as the starting value of the mergeable interval; similarly, for idx +1 Performing the above calculations yields Idx +1 Start; the end value of the current interval IdxEnd=Idx +1 Start-1; calculate the interval length and store it in the list, that is, IdxIntervalLen[i] = IdxEnd-IdxStart+1; Traverse each element in IndicesLst, and after performing the above operations, a list IdxIntervalLen of the same length as IndicesLst is obtained, in which each element corresponds to the length of the mergeable interval of each index value; finally, the length of the mergeable interval of the entire butterfly network is determined to be N mrg =min(IdxIntervalLen); record N con =N mrg , N group =N / N con ; b) Calculate the combined rotation factor, and calculate the exponent and sign bit of the combined rotation factor relative to the original root; For the i-th value in IndicesLst, the relative position of the index in the merged interval is recorded as i rel =IndicesLst[i]-IdxStart, perform the following operations respectively to determine the merge rotation factor corresponding to each index value; First, the number of merge rotation factors that need to be calculated is At the same time, we get the number of times N that needs to be copied in subsequent calculations dup =N mrg / 2N comp ; The index of the rotation factor relative to the original root is the sum of all the indices TwiddleExp[j] on the merge path, that is, Recorded as k=0,1,...N comp -1; at the same time, the sign bit of the combined rotation factor is the product of the sign bits sign[j] of all the rotation factors on the combined path, recorded as Therefore, each index value IndicesLst[i] corresponds to N comp The combined rotation factors are recorded as a list Its sign bit is a list of the same length Traverse each element in IndicesLst and perform the above operations to obtain lists ExpSet and SignSet of the same length as IndicesLst, where each element ExpSet i With SignSet i They are all a length of N comp .
Citation Information
Cited By
Polynomial multiplication accelerator
CN120762625A
Data processing method, device and equipment for sparse calculation of neural network
CN122088577A
Data processing method, device and equipment for neural network sparse computation
CN122088577B