Homomorphic convolution acceleration method based on approximate fast Fourier transform
By replacing the number theory transformation into the fast Fourier transform in homomorphic convolution calculation and using different bit widths in multiple stages, the problem of large overhead of homomorphic convolution calculation is solved, and efficient calculation and significant power consumption reduction is achieved.
Patent Information
- Application Number
- CN202510210462.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
There are a large number of repeated calculations of number theory transformations and inverse transformations in the calculation of weight polynomials, resulting in large calculation overhead and high memory consumption.
By replacing the number theory transform with a fast Fourier transform (FFT), and using different bit widths in multiple stages of the FFT, the optimal balance between power consumption and accuracy is found through multi-objective design space exploration methods.
The efficiency improvement of homomorphic convolutional calculation is achieved, the hardware cost is reduced, and the power consumption reduction is achieved by maintaining acceptable error growth.
Smart Images

Figure CN120145442A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of algorithm optimization for privacy computing, and particularly relates to a method for accelerating homomorphic convolution by using approximate fast Fourier transform. Background Art
[0002] Privacy protection has become one of the main concerns when deploying deep neural networks (DNNs) in the cloud. Homomorphic encryption (HE) has recently been proposed and has attracted wide attention. By encrypting data into ciphertext polynomials, HE allows computations to be performed on encrypted data without revealing any information about the data itself. To apply HE to private DNN inference, there are mainly two methods: fully homomorphic encryption (FHE) schemes and hybrid HE / two-party computation (2PC) schemes. The main difference between the two lies in the implementation of non-linear activation functions. The hybrid scheme uses a two-party secure computation protocol, which helps to avoid activation function approximation and high-cost bootstrapping operations in the FHE scheme. Therefore, the present invention will focus on optimizing the hybrid scheme.
[0003] Private neural network inference is still much slower than plaintext inference, mainly due to the high computational overhead. Compared with the computations of non-linear layers, homomorphic convolution (HConv) has become the main computational bottleneck, and its computational process is as Figure 1As shown in the figure. The client has the input activation vector X of the convolutional layer, encodes it into a polynomial and encrypts it, and then sends it to the server for preprocessing and N-point NTT transform calculation. The server has the weight vector W, encodes it into a polynomial, performs preprocessing and NTT transform calculation, then multiplies it point by point with the N-point NTT transform result of the ciphertext, and then sends the calculation result back to the client for decryption and postprocessing. In this way, the calculated data can be obtained at the client without leaking the input data, and the client will not know the information related to the neural network on the server side. In particular, the polynomial calculation of the weights W and X is accelerated by the number theoretic transform (NTT). The reason for the large computational overhead of homomorphic convolution lies in the repeated calculations of a large number of number theoretic transforms and their inverse transforms (INTT), especially in the calculation of the weight polynomial. Although the weight polynomial in the NTT domain can be precomputed and stored, this will result in significant memory overhead. For example, storing all the weights of a 4-bit quantized ResNet-50 in the NTT domain requires 23GB of memory, which results in a memory consumption more than 1000 times higher than the normal case.
[0004] In recent years, different homomorphic encryption accelerators have been proposed to accelerate the costly number theoretic transforms. Some studies focus on optimizing the data flow and avoiding stalls between stages to improve parallelism [1][2]. Others achieve acceleration by decomposing large-scale NTTs into multiple small NTTs [3]. Although these accelerators have made promising progress in acceleration, they usually face high area and power consumption costs. For example, the area of the accelerator proposed in [3] exceeds 150mm 2 , and the power consumption is about 100W. It is found that homomorphic privacy inference has three levels of fault tolerance, which come from homomorphic decryption, neural network quantization, and the fault tolerance of the network itself. However, the number theoretic transform is based on modular arithmetic, and the modulo operation makes it have no fault tolerance and even amplifies the proportion of the error in the original data. The fast Fourier transform (FFT) can not only accelerate polynomial multiplication but also has the potential for approximation. Therefore, it is of great significance to study how to use the fast Fourier transform to replace the number theoretic transform and further reduce the computational overhead through approximation to improve the computational efficiency of homomorphic convolution.
[0005] The fast Fourier transform and the number theoretic transform can generally be used to accelerate polynomial multiplication. FFT and NTT share the same data flow structure and adopt the Cooley-Turkey (CT) butterfly algorithm. As Figure 2 shown, the complexity of polynomial multiplication is reduced from O(N 2) Reduced to O(NlogN). The difference between the two lies in the data type and the arithmetic units used. FFT processes data with complex floating-point coefficients and requires complex adders and multipliers, while NTT processes polynomials over a ring and requires modular adders and multipliers. After encoding and encryption, the calculations are performed over a polynomial ring. Therefore, in traditional schemes, NTT is used to calculate the exact multiplication results of the weight and activation polynomials. In contrast, due to data type issues, directly using FFT to replace NTT will introduce calculation errors unless the data width is large enough. Mainstream homomorphic convolution accelerators are based on NTT and the inverse convolution theorem, as Figure 1 shown. However, NTT itself is more sensitive to errors, and the twiddle factors vary with different moduli, while FFT does not have these problems. Therefore, based on the fault tolerance in homomorphic privacy inference applications, it is of great significance to study a homomorphic convolution acceleration method that replaces number-theoretic transform with fast Fourier transform and introduces approximation.
[0006] [1] M.S.Riazi, K.Laine, B.Pelton, and W.Dai, “HEAX: An architecture for computing on encrypted data,” in Proceedings of the twenty-fifth international conference on architectural support for programming languages and operating systems, 2020, pp.1295–1309.
[0007] [2] X.Ren, Z.Chen, Z.Gu, Y.Lu, R.Zhong, W.-J.Lu, J.Zhang, Y.Zhang, H.Wu, X.Zheng et al., “CHAM: A customized homomorphic encryption accelerator for fast matrix-vector product,” in 2023 60th ACM / IEEE Design Automation Conference (DAC). IEEE, 2023, pp.1–6.
[0008] [3] N. Samardzic, A. Feldmann, A. Krastev, S. Devadas, R. Dreslinski, C. Peikert, and D. Sanchez, “F1: A fast and programmable accelerator for fully homomorphic encryption,” in MICRO-54: 54th Annual IEEE / ACM International Symposium on Microarchitecture, 2021, pp. 238–252. Summary of the Invention
[0009] To solve the problems existing in the prior art, the present invention proposes a homomorphic convolution acceleration method based on approximate fast Fourier transform. By utilizing the fault tolerance feature of homomorphic convolution, the number theoretic transform for privacy inference is replaced with the fast Fourier transform, and an approximation method is introduced to further reduce the bit width, so as to reduce the hardware cost of each operation.
[0010] The technical solution of the present invention is as follows:
[0011] A homomorphic convolution acceleration method based on approximate fast Fourier transform, characterized in that by utilizing the fault tolerance feature of homomorphic convolution, the number theoretic transform for privacy inference is replaced with the fast Fourier transform, and an approximation method is introduced to further reduce the bit width, so as to reduce the hardware cost of each operation; in the process of introducing approximation, by constructing a parameter space, different bit widths are used in multiple stages of the fast Fourier transform and explored to find the best balance result between power consumption and accuracy. The specific steps are as follows:
[0012] The first step: Client encryption
[0013] The client encodes the input activation vector X of the convolutional neural network into a polynomial of length N and encrypts it according to the hybrid HE / 2PC encryption protocol; the obtained ciphertext polynomial X EN is sent to the server for homomorphic convolution operation;
[0014] The second step: Server-side calculation
[0015] The server-side calculation includes the following steps:
[0016] S1) Input the ciphertext polynomial X EN and the preprocessing of the weight polynomial W; the preprocessing includes two operations of folding and rotation. After preprocessing, the input polynomial X rot and the weight polynomial W rot in complex form are obtained;
[0017] S2) Input polynomial X rot and weight polynomial W rot to perform fast Fourier transform FFT to obtain the input polynomial and weight polynomial in the FFT domain;
[0018] S3) Multiply the input polynomial and weight polynomial in the FFT domain point - by - point to obtain the processing result Y EN ;
[0019] Step 3: Client - side decryption
[0020] The client receives the processing result Y EN from the server, and performs decryption calculation according to the encryption protocol to obtain Y fft , then performs the corresponding IFFT calculation to obtain Y rot , and post - processes the decrypted polynomial. Perform rounding and modulo - q operations on the coefficients of the polynomial Y appr obtained after post - processing, and finally obtain the output polynomial Y of the homomorphic convolution.
[0021] Furthermore, for the above - mentioned homomorphic convolution acceleration method based on approximate fast Fourier transform, the pre - processing in step S1) of the server - side calculation in the second step includes two operations: folding and rotation. Specifically:
[0022] S1 - 1) Folding operation: Transform the polynomials X N and W belonging to Z[X] / (X EN + 1)modq into complex - number polynomials X and W fold in the form of fold ;
[0023] S1 - 2) Rotation operation: Multiply the folded N / 2 - level complex - number polynomials X fold and W fold point - by - point with a specific rotation factor to obtain X rot and W rot , and the rotation factor is
[0024] Furthermore, for the above - mentioned homomorphic convolution acceleration method based on approximate fast Fourier transform, the fast Fourier transform FFT of the input polynomial X rot and weight polynomial W rot in step S2) of the server - side calculation in the second step is implemented by a butterfly network. For N / 2 - point FFT calculation, the input is N / 2 complex numbers, and perform log 2N / 2-level calculation, each level of calculation contains N / 4 butterfly units. Inside each butterfly unit, complex multiplication and addition calculations are performed pairwise on the data. In the process of determining the bit width of the FFT calculation, multi-objective design space exploration is used to approximate the intermediate data and rotation factors at each level of the butterfly network, quantize the original floating-point numbers into fixed-point numbers and reduce the bit width, which specifically includes the following steps:
[0025] S2-1) Construct the parameter space: The N / 2-point FFT butterfly network has a total of log 2 N / 2 levels. Complex multiplication is only necessary in the last log 2 N / 2 - 2 levels. The power consumption of complex multiplication is related to the bit width dw i of the intermediate data and the bit width tw i of the rotation factor, where the selected parameter set Ω in the dimension is expressed as:
[0026] Each parameter in each dimension represents the bit width of the intermediate data or rotation factor input to the complex multiplier, and the value range is from b L to b H , the upper limit b H represents the highest supported bit width, and the lower limit b L represents the lowest bit width allowed by the FFT calculation error tolerance;
[0027] S2-2) Design space evaluation: For the samples obtained in the constructed parameter space, construct the corresponding approximate FFT, perform error estimation and hardware overhead calculation; for error evaluation, use the method of multiple simulations; for hardware overhead evaluation, obtain the estimated power consumption size through the look-up table for the approximate FFT parameters obtained by sampling;
[0028] S2-3) Multi-objective design space exploration: Use the multi-objective Bayesian optimization method to search for the Pareto optimal design;
[0029] Through the steps S2-1) to S2-3), a set of Pareto fronts regarding the power consumption and error of the approximate FFT is obtained. According to the preset maximum error tolerance in the actual application, select a set of approximate FFT parameters with the corresponding lowest power consumption; perform approximate fast Fourier transform on the input polynomial and weight polynomial respectively according to this parameter to obtain the input polynomial and weight polynomial in the FFT domain.
[0030] Furthermore, for the above homomorphic convolution acceleration method based on approximate fast Fourier transform, in the design space evaluation of step S2-2) in the second step of server-side calculation, specifically:
[0031] For error evaluation, multiple simulations are adopted. First, a set of corresponding activation tensors and weight tensors are randomly obtained from the real network input. In Matlab, encoding is implemented for this set of input values according to the encryption protocol to obtain the plaintext polynomial. For the plaintext polynomial of the activation values, it is input into the SEAL library for encryption to obtain the ciphertext polynomial. Then, homomorphic convolution calculation based on the current approximate FFT is implemented in Matlab, and the obtained result is sent back to the SEAL library for decryption. The decrypted plaintext polynomial continues to be post - processed in Matlab, including coefficient extraction, activation function, and quantization, to obtain the final calculated value, which is compared with the accurate calculated value to obtain the calculation error.
[0032] For hardware overhead evaluation, the estimated power consumption size is obtained by looking up the table with the sampled approximate FFT parameters. The method based on the lookup table LUT is adopted, and the hardware cost estimation of different data - width configurations is obtained from the LUT by aggregating the cost of pre - synthesized butterfly units.
[0033] Furthermore, in the above - mentioned homomorphic convolution acceleration method based on the approximate fast Fourier transform, the post - processing of the decrypted polynomial in the third - step client decryption includes two steps: rotation and expansion. Specifically:
[0034] 1) Rotation operation: Multiply the N / 2 - level complex polynomial Y rot by a specific rotation factor. The corresponding rotation factors for multiplication result in Y fold ;
[0035] 2) Expansion operation: Transform the complex polynomial Y in the form of flod back to Y N in the form of R[X] / (X appr +1).
[0036] The technical effects of the present invention are as follows:
[0037] A homomorphic convolution acceleration method based on approximate fast Fourier transform optimizes the calculation of homomorphic convolution, replaces the number-theoretic transform NTT with the fast Fourier transform FFT, and introduces an approximation method to further reduce the bit width, so as to reduce the hardware cost of each operation. In the process of determining the approximate FFT bit width, a multi-objective design space exploration method is adopted, different bit widths are used in multiple stages of the fast Fourier transform, a space evaluation method based on a lookup table is designed, and multi-objective space exploration is carried out on it to achieve a design balance between calculation accuracy and power consumption. The present invention is applicable to any convolutional layer, reduces the overall overhead of homomorphic convolution, and improves the calculation efficiency. In practical applications, different approximation degrees are determined according to the acceptable different accuracies to achieve the optimal hardware efficiency. Compared with the common NTT-based homomorphic convolution scheme, the approximate FFT-based homomorphic convolution of the present invention can achieve more than 10 times power consumption reduction while maintaining an acceptable error increase. Description of the Drawings
[0038] Figure 1 It is a flow chart of homomorphic convolution implemented based on number-theoretic transform in the hybrid HE / 2PC encryption scheme;
[0039] Figure 2 It is a schematic diagram of the implementation of the butterfly network algorithm for 8-point number-theoretic transform in the algorithm implementation of the fast Fourier transform;
[0040] Figure 3 It is a flow chart of homomorphic convolution implemented based on the fast Fourier transform proposed by the present invention;
[0041] Figure 4 It is a work flow chart of design space evaluation in the multi-objective design space exploration proposed by the present invention. Detailed Embodiment
[0042] The present invention will be further clearly and completely elaborated below with reference to the drawings through specific embodiments.
[0043] A homomorphic convolution acceleration method based on approximate fast Fourier transform proposed by the present invention has an overall process as Figure 3 shown, which is mainly divided into three steps: the client encrypts the data and sends it to the server, the server performs homomorphic calculation on the encrypted data, and the client receives the calculated data and decrypts it to obtain the final calculation result; the following will be elaborated in detail with reference to the drawings.
[0044] The First Step: Client Encryption
[0045] The client encodes the input activation vector X of the convolutional neural network into a polynomial of length N and encrypts it according to the hybrid HE / 2PC encryption protocol; the obtained ciphertext polynomial X EN is sent to the server for homomorphic convolution operation;
[0046] Step 2: Server-side calculation
[0047] The server-side calculation includes the following steps:
[0048] S1) Input the ciphertext polynomial X EN and the preprocessing of the weight polynomial W; X EN and W both belong to the polynomial ring Z[X] / (X N +1)modq, the degree of the polynomial is N, and the coefficients are integers modulo q; the preprocessing is divided into two operations: folding and rotation:
[0049] S1-1) Folding operation, transform the polynomials X N and W that belong to Z[X] / (X EN +1)modq into complex polynomial forms of X and W fold respectively; specifically, map the N coefficients {X fold [j]}, j = 0,..., N-1, of the input polynomial to the N / 2 complex coefficients {X EN [k]} of the N / 2-degree complex polynomial, where X fold [k] = X fold [2j] + X EN [2j + 1]i, where i is the imaginary unit, X EN [2j] is the real part, and X EN [2j + 1] is the imaginary part; for example, X EN = x EN + 2x 0 + 3x 1 + 4x 2 + 5x 3 + 6x 4 + 7x 5 + 8x 6 + 8x 7 , and the folded complex polynomial is X fold = (1 + 2i)x 0 + (3 + 4i)x 1 + (5 + 6i)x 2 + (7 + 8i)x 3 . Similarly, W can be folded and transformed into W fold .
[0050] S1-2) Rotation operation, multiply the folded N / 2-degree complex polynomial X fold by a specific rotation factor to get X rot , multiply the folded N / 2-degree complex polynomial W fold by a specific rotation factor to get Wrot , the corresponding multiplication rotation factor is That is For example:
[0051] S2) Input polynomial X rot , weight polynomial W rot Fast Fourier Transform (FFT). The FFT is implemented by a butterfly network, where each butterfly unit contains one complex multiplication calculation and two complex addition / subtraction calculations. Figure 2 shows the calculation process of an 8-point FFT butterfly network. For a more general N / 2-point FFT calculation, the input is N / 2 complex numbers, and log 2 N / 2 levels of calculations are performed. Each level of calculation contains N / 4 butterfly units. Inside each butterfly unit, complex multiplication and addition calculations are performed pairwise between data, as Figure 2 shown. The intermediate data m[i] and m[j] and the rotation factor perform a butterfly calculation to obtain M[i] and M[j]:
[0052]
[0053] During the FFT calculation process, the present invention proposes to approximate the intermediate data and the rotation factor, quantize the original floating-point numbers into fixed-point numbers and reduce the bit width. The reduction of the approximate bit width will significantly reduce the hardware overhead of the multiplier, but will introduce certain errors in the calculation. As mentioned above, the neural network privacy inference itself has a certain degree of fault tolerance. Therefore, determining the approximation degree to balance the hardware cost and the calculation accuracy is a key issue. The present invention proposes to use multi-objective design space exploration (DSE) to determine the fixed-point bit widths of the intermediate data and the rotation factor at each level during the FFT calculation process, mainly including the following steps:
[0054] S2-1) Construct the parameter space: As mentioned above, the N / 2-point FFT butterfly network has a total of log 2 N / 2 levels. Importantly, complex multiplication is only necessary in the last log 2 N / 2 - 2 levels. In the first stage, since all rotation factors are 1, no multiplication operation is required; in the second stage, since the rotation factors are 1 and -j, simple operations such as data exchange and sign flipping are involved. The power consumption of complex multiplication is related to the bit width dw i of the intermediate data and the bit width tw i of the rotation factor, where Therefore, the selected dimensional parameter set Ω is expressed as: Each parameter in each dimension represents the bit-width of the intermediate data or rotation factor input to the complex multiplier, with a value range from b L to b H 。The upper limit b H represents the highest supported bit-width, while the lower limit b L represents the lowest bit-width allowed for the tolerance of FFT calculation errors. For example, when N = 4096 and taking the bit-width of the traditional NTT-based homomorphic convolution as 39 bits, the upper limit b H of the bit-width of the replaced FFT intermediate data and rotation factor is generally set to 25 bits, and the lower limit b L represents the lowest bit-width allowed for the tolerance of FFT calculation errors. From the coarse-grained experimental results, it can be obtained that when the bit-width is reduced to 15 bits, the calculation errors introduced will cause the network accuracy to be very low. Therefore, b L can be taken as 15 bits.
[0055] S2-2) Design space evaluation: For the samples obtained in the constructed parameter space, construct the corresponding approximate FFT, perform error estimation and hardware overhead calculation, and the flow chart is as Figure 4 shown. For the evaluation of errors, the present invention adopts the method of multiple simulations. First, a set of corresponding activation tensors and weight tensors are randomly obtained in the real network input, and the encoding of this set of input values according to the encryption protocol is implemented in Matlab to obtain the plaintext polynomial. The plaintext polynomial of the activation value is input into the SEAL library for encryption to obtain the ciphertext polynomial. Next, the homomorphic convolution calculation based on the current approximate FFT is implemented in Matlab, and the obtained result is sent back to the SEAL library for decryption. The decrypted plaintext polynomial continues to be post-processed in Matlab, including coefficient extraction (the inverse process of encoding), activation function, and weight quantization, to obtain the final calculated value and compare it with the accurate calculated value to obtain the calculation error. At the same time, for the evaluation of hardware overhead, the estimated power consumption size is obtained through the look-up table for the sampled approximate FFT parameters. Performing power consumption evaluation for each configuration through register transfer level (RTL) simulation may be very time-consuming. To quickly estimate the hardware cost, a look-up table (LUT)-based method is adopted. This method obtains the hardware cost estimation of different data width configurations from the LUT by aggregating the costs of pre-synthesized butterfly units. For example, the present invention constructs a two-dimensional look-up table containing the power consumption of multipliers with two input bit-widths ranging from 15 to 25 bits respectively. Then, according to the bit-width parameters in the sample, the sum of the power consumption of the approximate FFT is calculated by looking up the table. Although the LUT-based method provides inaccurate hardware cost estimation for designers, it can achieve fast evaluation and provide support for subsequent multi-objective exploration.
[0056] S2-3) Multi-objective design space exploration: The present invention is designed with two objectives, hardware cost and computing accuracy. Based on the above evaluation method, a multi-objective Bayesian optimization method is used to search for the Pareto optimal design.
[0057] Through the above method of design space exploration, a set of Pareto fronts regarding the power consumption and error of the approximate FFT can be obtained. According to the maximum tolerable error in actual applications, a set of approximate FFT parameters with the lowest power consumption can be selected. Then, according to these parameters, the input polynomial and the weight polynomial are respectively subjected to approximate fast Fourier transform to obtain two polynomials in the FFT domain.
[0058] S3) Perform a dot product on the input polynomial and the weight polynomial in the FFT domain obtained in the previous step to obtain Y EN 。
[0059] Step 3: Client decryption. As Figure 3 shown, after the client receives the processing result Y EN from the server, it performs decryption calculation according to the encryption protocol to obtain Y fft , and then performs the corresponding IFFT calculation to obtain Y rot . Corresponding to the previous preprocessing, in order to obtain the final result, post-processing of the decrypted polynomial is also required, including two operations: rotation and expansion:
[0060] 1) The rotation operation is to perform a dot product on the N / 2-level complex polynomial Y rot with a specific rotation factor. The corresponding rotation factor for multiplication results in Y fold .
[0061] 2) The expansion operation transforms the complex polynomial Y in the form of flod back to the form of R[X] / (X N +1). The specific implementation method is to map the N / 2 complex coefficients {Y fold [k]} of the N / 2-level complex polynomial, where Y fold [k] = Y appr [2j] + Y appr [2j + 1]i, where i is the imaginary unit, Y appr [2j] is the real part, and Y appr [2j + 1] is the imaginary part, to the N coefficients {Y appr}[j] of the N-level polynomial Y appr , j = 0,..., N - 1.
[0062] The polynomial Y N in the form of R[X] / (X apprIts coefficients need to be rounded and taken modulo q, and transformed into the form of Z[X] / (X N +1) mod q to obtain the output polynomial Y of the homomorphic convolution, as Figure 3 shown.
[0063] Through the above method, the homomorphic convolution can be optimized for calculation. Compared with the common NTT-based homomorphic convolution scheme, the homomorphic convolution based on approximate FFT of the present invention can achieve more than 10 times power consumption reduction while maintaining an acceptable error growth.
[0064] Finally, it should be noted that the purpose of publishing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention is subject to the scope defined by the claims.
Claims
1. A homomorphic convolution acceleration method based on approximate fast Fourier transform, characterized in that: By taking advantage of the fault-tolerance characteristics of homomorphic convolution, we replace the number-theoretic transformation of privacy reasoning with the fast Fourier transform, and introduce an approximation method to further reduce the bit width to reduce the hardware cost of each operation. In the process of introducing the approximation, we construct a parameter space, use different bit widths in multiple stages of the fast Fourier transform, and explore them to find the best balance between power consumption and accuracy. The specific steps are as follows: Step 1: Client-side encryption The client encodes the input activation vector X of the convolutional neural network into a polynomial of length N and encrypts it according to the hybrid HE / 2PC encryption protocol; the obtained ciphertext polynomial X EN Send to the server for homomorphic convolution operation; Step 2: Server-side calculation The server-side calculation includes the following steps: S1) Input ciphertext polynomial X EN and weight polynomial W; the preprocessing includes two steps of folding and rotation, and the complex input polynomial X is obtained after preprocessing rot , weight polynomial W rot ; S2) Input polynomial X rot , weight polynomial W rot Fast Fourier transform FFT, obtain the input polynomial and weight polynomial in the FFT domain; S3) Perform point multiplication on the input polynomial and weight polynomial in the FFT domain to obtain the processing result Y EN ; Step 3: Client Decryption The client receives the processing result Y from the server EN After that, decryption calculation is performed according to the encryption protocol to obtain Y fft , and then perform the corresponding IFFT calculation to obtain Y rot , and post-process the decrypted polynomial to obtain the polynomial Y appr The coefficients of are rounded and modulo q, and finally the output polynomial Y of the homomorphic convolution is obtained.
2. The homomorphic convolution acceleration method based on approximate fast Fourier transform according to claim 1, characterized in that: The preprocessing in step S1) of the second step of server-side calculation includes two steps of folding and rotating, specifically: S1-1) folding operation, which belongs to Z[X] / (X N +1) modq of the polynomial X EN and W, respectively, are transformed into A complex polynomial of the form X fold and W fold ; S1-2) Rotation operation, the folded N / 2 level complex polynomial X fold and W fold , multiply by a specific rotation factor to get X rot and W rot , the rotation factor is 3. The homomorphic convolution acceleration method based on approximate fast Fourier transform according to claim 1, characterized in that: In the second step of server-side calculation, step S2) inputs the polynomial X rot , weight polynomial W rot Fast Fourier transform FFT, FFT is implemented by butterfly network, for N / 2 point FFT calculation, the input is N / 2 complex numbers, log2N / 2 level calculation is performed, each level of calculation contains N / 4 butterfly units, and inside each butterfly unit, complex multiplication and addition calculation is performed between two data; in the process of determining the bit width of FFT calculation, multi-objective design space exploration is used to approximate the intermediate data and rotation factors of each level of butterfly network, quantize the original floating point number into fixed point number and reduce the bit width, which specifically includes the following steps: S2-1) Construct parameter space: The N / 2-point FFT butterfly network has a total of log2N / 2 stages. Complex multiplication is only necessary in the last log2N / 2-2 stages. The power consumption of complex multiplication is related to the intermediate data bit width dw i and the bit width tw of the rotation factor i about, among which Selected The dimension parameter set Ω is expressed as: Each parameter in each dimension represents the bit width of the intermediate data or rotation factor input to the complex multiplier, ranging from b L to b H , upper limit b H Indicates the maximum bit width supported, and the lower limit b L It indicates the minimum bit width allowed by the FFT calculation error tolerance; S2-2) Design space evaluation: For the samples collected in the constructed parameter space, construct the corresponding approximate FFT, perform error estimation and hardware cost calculation; for error evaluation, use multiple simulations; for hardware cost evaluation, use the approximate FFT parameters obtained by sampling to obtain the estimated power consumption through a lookup table; S2-3) Multi-objective design space exploration: using multi-objective Bayesian optimization methods to search for Pareto optimal designs; Through the steps S2-1) to S2-3), a set of Pareto fronts about power consumption and error of approximate FFT is obtained, and a set of approximate FFT parameters with corresponding minimum power consumption is selected according to the maximum error tolerance preset in the actual application; according to the parameters, the input polynomial and the weight polynomial are respectively subjected to approximate fast Fourier transform to obtain the input polynomial and the weight polynomial in the FFT domain.
4. The homomorphic convolution acceleration method based on approximate fast Fourier transform according to claim 3, characterized in that: The step S2-2) design space evaluation is specifically as follows: For error evaluation, multiple simulations are used. First, a set of corresponding activation tensors and weight tensors are randomly obtained in the real network input. Then, the input values are encoded according to the encryption protocol in Matlab to obtain the plaintext polynomial. The activation value plaintext polynomial is input into the SEAL library for encryption to obtain the ciphertext polynomial. Then, the homomorphic convolution calculation based on the current approximate FFT is implemented in Matlab, and the obtained result is sent back to the SEAL library for decryption. The decrypted plaintext polynomial is further post-processed in Matlab, including coefficient extraction, activation function and re-quantization, to obtain the final calculated value, which is compared with the accurate calculated value to obtain the calculation error; For hardware cost evaluation, the approximate FFT parameters obtained by sampling are used to obtain the estimated power consumption through a lookup table. A lookup table (LUT)-based method is used to obtain hardware cost estimates for different data width configurations from the LUT by aggregating the costs of the pre-synthesized butterfly units.
5. The homomorphic convolution acceleration method based on approximate fast Fourier transform according to claim 1, characterized in that: The post-processing of the decrypted polynomial in the third step of client decryption includes two steps: rotation and expansion, which are specifically: 1) Rotation operation, transform the N / 2-order complex polynomial Y rot Multiply the point with a specific rotation factor, corresponding to the multiplied rotation factor Get Y fold ; 2) Expand the operation and A complex polynomial Y of the form flod Change back to R[X] / (X N +1) form of Y appr .
Citation Information
Patent Citations
Privacy amplification algorithm for quantum secret key distribution
CN104426655A
Fully homomorphic encryption deep learning reasoning method and system based on FPGA
CN112699384A
Privacy computing heterogeneous acceleration method and device based on fully homomorphic encryption
CN115622684A
Method and apparatus for hardware-based accelerated arithmetic operation on homomorphically encrypted message
US20230163945A1