Homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-thread mapping
By optimizing the homomorphic encryption algorithm on the GPU, using technical means such as nested loop optimization and BFV root pre-calculation, the problem of underutilizing the GPU characteristics in the existing technology is solved, and efficient polynomial multiplication calculation and homomorphic encryption performance improvement is achieved.
Patent Information
- Application Number
- CN202411811579.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-06
AI Technical Summary
When the prior art optimizes homomorphic encryption algorithms on GPUs, the memory access, thread scheduling characteristics and thread hierarchical structure of the GPU are not fully utilized, resulting in serious data dependence, affecting parallel efficiency, and inefficient calculation efficiency of fully homomorphic encryption.
Based on the CUDA-GPU structure model, a nested loop optimization method is adopted to perform loop optimization on the NTT transformation algorithm and the INTT transformation algorithm. Combined with the BFV original root pre-calculation and pre-store joint optimization algorithm, homomorphic multiplication in the GPU parallel computing environment is improved and polynomial multiplication calculation is optimized.
By optimizing algorithms and thread mapping, the execution efficiency of polynomial multiplication in GPU parallel computing environment is improved, data dependence is reduced, computing resource utilization is improved, and the performance of homomorphic encryption computing is significantly improved.
Smart Images

Figure CN119939618A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and in particular to a homomorphic encryption acceleration method and device based on polynomial multiplication optimization and GPU multi-threaded mapping. Background Art
[0002] Microsoft's SEAL library is a fully homomorphic encryption scheme based on the CPU. However, due to the limited parallel processing capability of the CPU, for homomorphic encryption that requires a large number of polynomial calculations, the CPU cannot efficiently perform calculations in parallel, resulting in slow calculation speed. With its excellent hardware performance, GPUs have been fully utilized in recent years to accelerate homomorphic encryption schemes, including: (1) arithmetic operations are optimized through algebraic tools such as CRT and DGT, reducing the data transmission overhead between the host and the device. However, this scheme does not solve the polynomial multiplication performance bottleneck caused by the growth of polynomial coefficients; (2) the first GPU hardware implementation of the CKKS scheme is proposed, which improves performance by designing and implementing memory-centric optimizations. However, this scheme does not fully consider the GPU computing scheduling process. Therefore, fully homomorphic encryption algorithms still face the problems of high algorithm complexity and low efficiency in practical applications, especially in polynomial multiplication calculations, where ciphertext data becomes polynomials with large coefficients and high degrees, with high computational complexity and consumption of a large amount of computing resources.
[0003] The computational complexity of the traditional term-by-term multiplication and convolution operation method is , which is not very efficient. However, the complexity can be reduced to , which can improve the performance of polynomial multiplication, but there are still efficiency issues in practical applications. The implementation schemes based on GPU acceleration include: (1) using GPU constant memory to solve thread divergence problems, solving shared memory conflicts through reasonable memory splitting and padding, and improving the cuHE NTT calculation speed by 20% - 50% under different problem sizes. However, this scheme does not fully utilize the GPU's hierarchical thread structure for optimization; (2) proposing three GPU microarchitecture extension methods to reduce memory bottlenecks and improve the throughput of the arithmetic pipeline in fully homomorphic operations, and improving the performance by 14.6 times compared with NVIDIAV100 GPU. However, this scheme does not consider the complex loop calculation overhead of fully homomorphic operations; (3) making full use of TCUs in Nvidia GPU to accelerate NTT transformation, changing the butterfly operation method of calculating NTT to directly using matrix multiplication to calculate NTT transformation, and using the TCU unit in GPU to accelerate by dividing matrix multiplication into blocks, the overall performance is improved by 32.3%. However, this scheme does not consider the complex loop calculation overhead of large matrix multiplication.
[0004] In summary, the optimization algorithm of the above method on GPU is not sufficient, and does not fully utilize the GPU's memory access, thread scheduling characteristics and thread hierarchy structure; when using the NTT transformation algorithm to accelerate polynomial multiplication under the CUDA framework, the common multi-layer loops will lead to serious data dependence, affecting the parallel efficiency of single instruction multiple data operations, resulting in low efficiency of fully homomorphic encryption calculations. Summary of the invention
[0005] In order to solve the technical problems that the existing technology does not fully utilize the memory access, thread scheduling characteristics and thread hierarchical structure of the GPU; when using the NTT transformation algorithm to accelerate polynomial multiplication under the CUDA framework, the common multi-layer loop will lead to serious data dependence, affecting the parallel efficiency of single instruction multiple data operations, and resulting in low efficiency of fully homomorphic encryption calculations, the embodiment of the present invention provides a homomorphic encryption acceleration method and device based on polynomial multiplication optimization and GPU multi-thread mapping. The technical solution is as follows:
[0006] On the one hand, a homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping is provided, and the method is implemented by a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping, and the method includes:
[0007] S1. Based on the CUDA-GPU structural model and according to the parallel computing characteristics of GPU, a nested loop optimization method is used to perform loop optimization on the structure of the NTT transformation algorithm to obtain a loop optimization algorithm based on the NTT transformation; a nested loop optimization method is used to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation;
[0008] S2. According to the loop optimization algorithm based on NTT transformation, the loop optimization algorithm based on INTT transformation, and the BFV primitive root pre-computation and pre-storage joint optimization algorithm, the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment is improved to obtain an improved homomorphic multiplication;
[0009] S3. Input two ciphertext data to be processed; use the improved homomorphic multiplication to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain a processed ciphertext result.
[0010] On the other hand, a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping is provided, and the device is applied to a homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping, and the device includes:
[0011] The first acquisition unit is used to optimize the loop of the NTT transformation algorithm structure based on the CUDA-GPU structure model and the parallel computing characteristics of the GPU by using a nested loop optimization method to obtain a loop optimization algorithm based on the NTT transformation; and optimize the loop of the INTT transformation algorithm structure based on the nested loop optimization method to obtain a loop optimization algorithm based on the INTT transformation;
[0012] A second acquisition unit is used to improve the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment according to the loop optimization algorithm based on the NTT transformation, the loop optimization algorithm based on the INTT transformation, and the BFV primitive root pre-computation and pre-storage joint optimization algorithm to obtain an improved homomorphic multiplication;
[0013] The third acquisition unit is used to input two ciphertext data to be processed; use the improved homomorphic multiplication to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain a processed ciphertext result.
[0014] On the other hand, a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping is provided, and the homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping comprises: a processor; a memory, wherein computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned homomorphic encryption acceleration methods based on polynomial multiplication optimization and GPU multi-threaded mapping is implemented.
[0015] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned homomorphic encryption acceleration methods based on polynomial multiplication optimization and GPU multi-threaded mapping.
[0016] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0017] The embodiment of the present invention firstly adopts a nested loop optimization method based on the CUDA-GPU structural model and the parallel computing characteristics of the GPU to perform loop optimization on the structure of the NTT transformation algorithm to obtain a loop optimization algorithm based on the NTT transformation; adopts a nested loop optimization method to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation; secondly, according to the loop optimization algorithm based on the NTT transformation, the loop optimization algorithm based on the INTT transformation and the BFV primitive root pre-computation and pre-storage joint optimization algorithm, the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment is improved to obtain the improved homomorphic multiplication; finally, two ciphertext data to be processed are input; the improved homomorphic multiplication is used to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain the processed ciphertext result.
[0018] Embodiments of the present invention The present invention performs GPU acceleration on SEAL fully homomorphic encryption through a CUDA environment, optimizes an NTT transformation algorithm and an INTT transformation algorithm according to the parallel computing characteristics of a GPU, expands a multi-layer loop nesting structure, improves parallelism and computing resource utilization, reduces data dependence, and performs pre-processing and post-processing on required data, thereby greatly improving the GPU's execution efficiency of polynomial multiplication operations; according to the GPU thread hierarchical structure, the present invention combines a loop optimization algorithm based on an NTT transformation and a loop optimization algorithm based on an INTT transformation with GPU thread mapping optimization, and combines the two layers of NTT transformation and INTT transformation in the optimized polynomial multiplication with a GPU scheduling process, thereby improving the GPU parallel computing environment polynomial multiplication execution efficiency by dividing an optimal number of threads and thread blocks. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 It is a flow chart of a homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-thread mapping provided by an embodiment of the present invention;
[0021] Figure 2 is a homomorphic multiplication flow chart provided by an embodiment of the present invention;
[0022] Figure 3 It is a framework diagram of mapping a fully homomorphic encryption algorithm and a GPU computing architecture provided by an embodiment of the present invention;
[0023] Figure 4 It is a block diagram of a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-thread mapping provided by an embodiment of the present invention;
[0024] Figure 5 It is a structural schematic diagram of a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0026] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0027] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0028] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0029] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0030] The embodiment of the present invention provides a homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping, which can be implemented by a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping, and the homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping can be a terminal or a server. Figure 1 The flowchart of the homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping is shown. The processing flow of the method may include the following steps:
[0031] S1. Based on the CUDA-GPU structural model and according to the parallel computing characteristics of GPU, a nested loop optimization method is used to perform loop optimization on the structure based on the NTT transformation algorithm to obtain a loop optimization algorithm based on the NTT transformation; a nested loop optimization method is used to perform loop optimization on the structure based on the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation.
[0032] Among them, the NTT transformation algorithm can accelerate the calculation of polynomial multiplication, but the NTT transformation algorithm is suitable for integer fields, especially integers on finite fields; in homomorphic encryption, the NTT transformation algorithm greatly reduces the computational complexity by converting polynomial multiplication into point multiplication operations, making the calculation of encrypted data efficient and feasible.
[0033] Optionally, S1 adopts a nested loop optimization method to perform loop optimization on the NTT transformation algorithm structure to obtain a loop optimization algorithm based on NTT transformation, including:
[0034] The three-layer loop based on the NTT transformation algorithm is nested by removing one layer of loop, and the three-layer loop is reduced to two layers, and only an outer loop that keeps loop iteration and an inner loop that keeps internal high parallelism is retained, so as to obtain a loop optimization algorithm based on NTT transformation.
[0035] Optionally, the implementation process of the loop optimization algorithm based on NTT transformation includes:
[0036] The input pointer points to the array storing the polynomial coefficients, the number of polynomials, the number of coefficient moduli, the power of the modulus of the polynomial, the data structure based on the NTT transformation, the signal for selecting the root power or the inverse root power, the thread index, and the value of the polynomial modulus;
[0037] Calculate the current layer interval according to the power of the modulus of the polynomial;
[0038] Calculate the root index and coefficient index according to the current layer interval and thread index;
[0039] Use the polynomial number for loop processing, judge by the number of coefficient moduli, select the corresponding index root power and modulus, and end the loop when the polynomial number is satisfied, and output the result based on NTT transformation.
[0040] Among them, Table 1 is the code of the loop optimization algorithm based on NTT transformation.
[0041] Table 1
[0042]
[0043] Optionally, S1 uses a nested loop optimization method to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation, including:
[0044] The three-layer loop based on the INTT transformation algorithm is nested by removing one layer of loop, and the three-layer loop is reduced to two layers, and only an outer loop that keeps loop iteration and an inner loop that keeps internal high parallelism is retained, so as to obtain a loop optimization algorithm based on the INTT transformation.
[0045] Optionally, the implementation process of the loop optimization algorithm based on INTT transformation includes:
[0046] The input pointer points to the array storing the polynomial coefficients, the number of polynomials, the number of coefficient moduli, the power of the modulus of the polynomial, the data structure of the INTT transformation, the signal for selecting the root power or the inverse root power, the thread index, and the value of the polynomial modulus;
[0047] Calculate the current layer interval according to the power of the modulus of the polynomial;
[0048] Calculate the root index and coefficient index based on the current layer interval, the power of the modulus of the polynomial, and the thread index;
[0049] Use the polynomial number for loop processing, judge by the number of coefficient moduli, select the corresponding index root power and modulus, and end the loop when the polynomial number is satisfied, and output the result of INTT transformation.
[0050] Among them, Table 2 is the code of the loop optimization algorithm based on INTT transformation.
[0051] Table 2
[0052]
[0053] S2. According to the loop optimization algorithm based on NTT transformation, the loop optimization algorithm based on INTT transformation and the BFV primitive root pre-computation and pre-storage joint optimization algorithm, the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment is improved to obtain the improved homomorphic multiplication.
[0054] Among them, since the NTT transformation algorithm requires a larger modulus in fully homomorphic encryption to accommodate the noise growth caused by homomorphic operations and the demand for complex parameters, preprocessing and post-processing based on NTT transformation are crucial.
[0055] Among them, for homomorphic multiplication under large modulus, preprocessing and postprocessing operations can be performed before and after the NTT transformation, and the Chinese remainder theorem can be used to decompose the large integer polynomial ring into multiple small polynomial rings; the modulus basis expansion can be performed before the NTT transformation to increase the operation range and accuracy, and the modulus basis conversion after the NTT transformation can reduce noise and ensure the accuracy and reversibility of the operation. By performing preprocessing and postprocessing operations before and after the NTT transformation, multiplication operations can be performed efficiently while maintaining the homomorphic properties of the encrypted data.
[0056] Among them, the Chinese remainder theorem is a basic theorem in number theory, which can solve integer problems that satisfy multiple congruence equations at the same time. It is a method mastered by technical personnel in this field and will not be further elaborated in the present invention.
[0057] Among them, the calculation of complex parameters, including primitive root calculation and inverse element calculation, because the calculation of the above complex parameters only needs to be calculated once, the complex parameters can be pre-calculated and stored in the global memory for data preparation. By directly calling the method instead of direct calculation, the calculation steps in the specific execution are reduced. For GPU acceleration, data association can be reduced to facilitate parallel processing.
[0058] In a feasible implementation, for the calculation of primitive roots in the integer domain of the BFV scheme and the BGV scheme, the minimum primitive root can be calculated through the BFV primitive root pre-computation and pre-storage joint optimization algorithm, and the product can be cyclically multiplied to obtain the primitive roots of all items, and finally stored in the device global memory at the beginning of the NTT data table initialization stage.
[0059] Among them, the BFV primitive root pre-computation and pre-storage joint optimization algorithm is a fully homomorphic encryption algorithm and an important means to improve the performance of the fully homomorphic encryption algorithm. It improves the speed of encryption and decryption operations by reducing online calculations and optimizing data access.
[0060] In a feasible implementation, the specific implementation steps of the BFV primitive root pre-computation and pre-storage joint optimization algorithm may include:
[0061] Enter the number of coefficients, modulus, and array to store roots;
[0062] Precompute primitive root parameters;
[0063] According to the number of coefficients, the minimum primitive root is calculated using the estimated primitive root parameter, modulus and primitive root, and the primitive roots of all items are obtained through loop processing; the primitive roots of all items are stored in the GPU global memory.
[0064] Among them, Table 3 is the code of the BFV primitive root pre-computation and pre-storage joint optimization algorithm.
[0065] Table 3
[0066]
[0067] In a feasible implementation, the homomorphic multiplication process is a multiplication process between ciphertext polynomials, such as Figure 2 The flowchart of homomorphic multiplication is shown; in one embodiment of the method, the specific implementation process of homomorphic multiplication includes: inputting two ciphertext polynomials, before multiplying the two ciphertext polynomials, placing the original polynomial ring in the polynomial ring The ciphertext polynomial is converted to a polynomial ring The modulus Q is set much larger than the modulus q to improve the calculation accuracy. The first ciphertext polynomial and the second ciphertext polynomial Perform polynomial multiplication based on NTT transformation to convert the ciphertext polynomial from Convert to polynomial , where the above conversion operation can reduce noise; the three elements generated in the obtained result can be expressed by the following formula (1):
[0068] (1)
[0069] Where c represents the result of the first multiplication before the ciphertext polynomial is relinearized; t represents the modulus of the plaintext coefficient; q represents the modulus of the ciphertext coefficient; Represents the first polynomial contained in the ciphertext A; Represents the second polynomial contained in the ciphertext A; Represents the first polynomial contained in the ciphertext B; Represents the second polynomial contained in the ciphertext B.
[0070] Among them, the obtained result generates three elements that cannot be used for the next multiplication. The result is reduced to two terms by performing a relinearization operation, and the result ciphertext polynomial is further obtained. The relinearization process can be expressed by the following formula (2):
[0071]
[0072] :
[0073] (2)
[0074] in, represents the conversion key; where is the noise distribution on R, , , ; C represents the resulting ciphertext polynomial obtained after relinearization; Indicates the first polynomial contained in the multiplication result c of the first step; Indicates the second polynomial contained in the multiplication result c of the first step; Indicates the third polynomial contained in the multiplication result c of the first step; represents a random sampling item on the ring; Represents a portion of the transformed key obtained by encryption operation; represents the noise random sampling item; s represents the private key; The logarithm base is used to represent Control scale.
[0075] Among them, under GPU acceleration, the three-layer nested loop algorithms based on NTT transformation and INTT transformation are limited; multiple layers of loops will form dependencies between data, so each iteration of the outer loop needs to wait for the inner loop to complete. In the three-layer nested loop, each iteration of the outer loop needs to be synchronized, which will increase the control flow overhead; the three-layer nested loop will form a parallel interference for the execution of single instruction multiple data. Among them, the two-layer loop structure is suitable for the GPU's SIMT single instruction multiple thread architecture, which can enable more threads to execute the same instructions at the same time, so that the GPU's computing resources can be fully utilized.
[0076] Among them, Table 4 is the code of the improved homomorphic multiplication.
[0077] Table 4
[0078]
[0079] Among them, by removing one layer of loop nesting from the three-layer loop based on NTT transformation, the three-layer loop is reduced to two layers, retaining only an outer loop that maintains loop iterations and an inner loop that maintains internal high parallelism. This solves the problem of data dependence caused by multi-layer loops in GPU parallel computing, and can achieve high parallelization, cooperate with optimized thread configuration, improve computing efficiency and reduce resource overhead.
[0080] S3. Input two ciphertext data to be processed; use the improved homomorphic multiplication to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain the processed ciphertext result.
[0081] Among them, polynomial multiplication is involved in the key generation process, encryption process, decryption process, homomorphic addition process, homomorphic multiplication process and bootstrapping process.
[0082] Optionally, the implementation process of the improved homomorphic multiplication includes:
[0083] Input two ciphertext data; wherein the ciphertext data includes: first ciphertext data and second ciphertext data;
[0084] Preprocess the two ciphertext data using the preset modulus base and NTTTables iterator to obtain the preprocessing result;
[0085] The first ciphertext data and the second ciphertext data are processed respectively using a cyclic optimization algorithm based on NTT transformation to obtain an NTT transformation result of the first ciphertext data and an NTT transformation result of the second ciphertext data;
[0086] Using polynomial multiplication to calculate the preprocessing result, the first ciphertext data and the second ciphertext data to obtain a calculation result;
[0087] Using a cyclic optimization algorithm based on INTT transformation to process the NTT transformation result of the first ciphertext data and the NTT transformation result of the second ciphertext data respectively, to obtain the INTT transformation result of the first ciphertext data and the INTT transformation result of the second ciphertext data;
[0088] According to the calculation result, the INTT transformation result of the first ciphertext data and the INTT transformation result of the second ciphertext data are relinearized by using the modulus-to-digital basis conversion to obtain the storage result of the first ciphertext data.
[0089] Optionally, after the step of inputting two ciphertext data to be processed and processing the two ciphertext data to be processed using the improved homomorphic multiplication to obtain the processed ciphertext result, S3 further includes:
[0090] According to the hierarchical structure characteristics of GPU multi-threaded parallel computing, each thread block executes the outer loop iterative calculation of the loop optimization algorithm based on NTT transformation and the loop optimization algorithm based on INTT transformation, and the thread bundles inside the thread block execute the inner loop iterative calculation of the loop optimization algorithm based on NTT transformation and the loop optimization algorithm based on INTT transformation through shared memory. By dividing the thread blocks and thread bundles with the optimal number and proportion, the hardware and software collaboration of homomorphic encrypted polynomial multiplication operations in the GPU multi-threaded parallel computing environment is optimized.
[0091] In a feasible implementation, Figure 3 The diagram shows the mapping framework between the improved homomorphic multiplication and the GPU computing framework. Table 5 shows the experimental environment parameters.
[0092] Table 5
[0093]
[0094] In order to complete the system functions and evaluate the optimization methods and the performance of the homomorphic library, the scheme proposed in this application is used for experiments and tests; the hardware environment includes GPU and CPU, the GPU is a single-core 4GB, and the high-performance computing core GPU model is: GeForce RTX 2080, and its video memory is 8GB. The software environment includes the operating system Ubuntu18.04 version, based on the CUDA framework version 12.2.91.
[0095] In a feasible implementation mode, the present application optimizes GPU threads, combines a loop optimization algorithm based on NTT transformation and a loop optimization algorithm based on INTT transformation with GPU thread mapping optimization, allocates homomorphic encryption polynomial multiplication loop calculation tasks to different thread levels through GPU resource mapping, optimizes the allocation of thread blocks and threads to improve the computational efficiency of polynomial multiplication in homomorphic encryption in the GPU, reduces parallel interference, and improves computing performance; combines a loop optimization algorithm based on NTT transformation and a loop optimization algorithm based on INTT transformation with GPU scheduling, and reasonably divides thread blocks according to the number of iterations of the outer loop in the two-layer nested loops of NTT transformation and INTT transformation, each thread block corresponds to one iteration of the outer loop of NTT transformation and INTT transformation; wherein, threads scheduled by thread bundles within the thread block can collaborate to complete the calculation of the inner loop of NTT transformation and INTT transformation through shared memory, reduce access to global memory, and improve the execution efficiency of polynomial multiplication by dividing the optimal number of threads and thread blocks, further accelerating the homomorphic encryption operation process in the GPU parallel computing environment.
[0096] Among them, Table 6 shows the thread and thread block allocation. Under the hardware and software experimental environment shown in Table 5 and the condition that the polynomial modulus is 8192, after testing the time efficiency of different thread and thread block allocation schemes, the optimal thread and thread block allocation scheme is obtained. Finally, the number of thread blocks is set to 64 and the number of threads in a block is set to 128.
[0097] Table 6
[0098]
[0099] Among them, by comparing the homomorphic encryption acceleration solution based on polynomial multiplication optimization and GPU multi-threaded mapping in the CUDA environment with the SEAL library solution implemented by the CPU, it can be found that the SEAL library implemented through software and hardware acceleration has greatly improved the operation time of each solution step of CKKS, BGV and BFV, and the total time has been increased by more than 30 times.
[0100] The loop optimization algorithm based on NTT transformation, the loop optimization algorithm based on INTT transformation and the improved homomorphic multiplication of the present invention enable the SEAL library to improve the computing efficiency and processing power when processing homomorphic encryption, especially when processing large amounts of data and complex polynomial operations, and can realize the GPU acceleration function of ciphertext homomorphic operations, and realize GPU acceleration for the Microsoft SEAL library; this application uses a single GPU card, and compared with the Intel i7 CPU single core, the acceleration ratios for the CKKS, BFV and BGV methods are 31.97, 53.45 and 50.79 respectively. As shown in Table 7, the time comparison results of different schemes can greatly reduce the execution time and greatly improve the performance.
[0101] Table 7
[0102]
[0103] The embodiment of the present invention firstly adopts a nested loop optimization method based on the CUDA-GPU structural model and the parallel computing characteristics of the GPU to perform loop optimization on the structure of the NTT transformation algorithm to obtain a loop optimization algorithm based on the NTT transformation; adopts a nested loop optimization method to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation; secondly, according to the loop optimization algorithm based on the NTT transformation, the loop optimization algorithm based on the INTT transformation and the BFV primitive root pre-computation and pre-storage joint optimization algorithm, the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment is improved to obtain the improved homomorphic multiplication; finally, two ciphertext data to be processed are input; the improved homomorphic multiplication is used to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain the processed ciphertext result.
[0104] Embodiments of the present invention The present invention performs GPU acceleration on SEAL fully homomorphic encryption through a CUDA environment, optimizes an NTT transformation algorithm and an INTT transformation algorithm according to the parallel computing characteristics of a GPU, expands a multi-layer loop nesting structure, improves parallelism and computing resource utilization, reduces data dependence, and performs pre-processing and post-processing on required data, thereby greatly improving the GPU's execution efficiency of polynomial multiplication operations; according to the GPU thread hierarchical structure, the present invention combines a loop optimization algorithm based on an NTT transformation and a loop optimization algorithm based on an INTT transformation with GPU thread mapping optimization, and combines the two layers of NTT transformation and INTT transformation in the optimized polynomial multiplication with a GPU scheduling process, thereby improving the GPU parallel computing environment polynomial multiplication execution efficiency by dividing an optimal number of threads and thread blocks.
[0105] Figure 4The present invention is a block diagram of a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-thread mapping according to an exemplary embodiment. The device is used for a homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-thread mapping. Figure 4 , the device includes a first acquisition unit 410, a second acquisition unit 420 and a third acquisition unit 430. Wherein:
[0106] The first acquisition unit 410 is used to optimize the loop of the NTT transformation algorithm structure based on the CUDA-GPU structure model and the parallel computing characteristics of the GPU by using a nested loop optimization method to obtain a loop optimization algorithm based on the NTT transformation; and optimize the loop of the INTT transformation algorithm structure based on the nested loop optimization method to obtain a loop optimization algorithm based on the INTT transformation;
[0107] A second acquisition unit 420 is used to improve the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment according to the loop optimization algorithm based on the NTT transformation, the loop optimization algorithm based on the INTT transformation, and the BFV primitive root pre-computation and pre-storage joint optimization algorithm to obtain an improved homomorphic multiplication;
[0108] The third acquisition unit 430 is used to input two ciphertext data to be processed; use the improved homomorphic multiplication to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain a processed ciphertext result.
[0109] Optionally, the nested loop optimization method is used to perform loop optimization on the NTT transformation algorithm structure to obtain a loop optimization algorithm based on NTT transformation, including:
[0110] The three-layer loop based on the NTT transformation algorithm is nested by removing one layer of loop, and the three-layer loop is reduced to two layers, and only an outer loop that keeps loop iteration and an inner loop that keeps internal high parallelism is retained, so as to obtain a loop optimization algorithm based on NTT transformation.
[0111] Optionally, the implementation process of the NTT transformation-based loop optimization algorithm includes:
[0112] The input pointer points to the array storing the polynomial coefficients, the number of polynomials, the number of coefficient moduli, the power of the modulus of the polynomial, the data structure based on the NTT transformation, the signal for selecting the root power or the inverse root power, the thread index, and the value of the polynomial modulus;
[0113] Calculate the current layer interval according to the power of the modulus of the polynomial;
[0114] Calculate the root index and coefficient index according to the current layer interval and thread index;
[0115] Use the polynomial number for loop processing, judge by the number of coefficient moduli, select the corresponding index root power and modulus, and end the loop when the polynomial number is satisfied, and output the result based on NTT transformation.
[0116] Optionally, the nested loop optimization method is used to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation, including:
[0117] The three-layer loop based on the INTT transformation algorithm is nested by removing one layer of loop, and the three-layer loop is reduced to two layers, and only an outer loop that keeps loop iteration and an inner loop that keeps internal high parallelism is retained, so as to obtain a loop optimization algorithm based on the INTT transformation.
[0118] Optionally, the implementation process of the loop optimization algorithm based on INTT transformation includes:
[0119] The input pointer points to the array storing the polynomial coefficients, the number of polynomials, the number of coefficient moduli, the power of the modulus of the polynomial, the data structure of the INTT transformation, the signal for selecting the root power or the inverse root power, the thread index, and the value of the polynomial modulus;
[0120] Calculate the current layer interval according to the power of the modulus of the polynomial;
[0121] Calculate the root index and coefficient index based on the current layer interval, the power of the modulus of the polynomial, and the thread index;
[0122] Use the polynomial number for loop processing, judge by the number of coefficient moduli, select the corresponding index root power and modulus, and end the loop when the polynomial number is satisfied, and output the result of INTT transformation.
[0123] Optionally, the implementation process of the improved homomorphic multiplication includes:
[0124] Input two ciphertext data; wherein the ciphertext data includes: first ciphertext data and second ciphertext data;
[0125] Preprocess the two ciphertext data using the preset modulus base and NTTTables iterator to obtain the preprocessing result;
[0126] The first ciphertext data and the second ciphertext data are processed respectively using a cyclic optimization algorithm based on NTT transformation to obtain an NTT transformation result of the first ciphertext data and an NTT transformation result of the second ciphertext data;
[0127] Using polynomial multiplication to calculate the preprocessing result, the first ciphertext data and the second ciphertext data to obtain a calculation result;
[0128] Using a cyclic optimization algorithm based on INTT transformation to process the NTT transformation result of the first ciphertext data and the NTT transformation result of the second ciphertext data respectively, to obtain the INTT transformation result of the first ciphertext data and the INTT transformation result of the second ciphertext data;
[0129] According to the calculation result, the INTT transformation result of the first ciphertext data and the INTT transformation result of the second ciphertext data are relinearized by using the modulus-to-digital basis conversion to obtain the storage result of the first ciphertext data.
[0130] Optionally, after the step of inputting two ciphertext data to be processed and using the improved homomorphic multiplication to process the two ciphertext data to be processed to obtain the processed ciphertext result, the step further includes:
[0131] According to the hierarchical structure characteristics of GPU multi-threaded parallel computing, each thread block executes the outer loop iterative calculation of the loop optimization algorithm based on NTT transformation and the loop optimization algorithm based on INTT transformation, and the thread bundles inside the thread block execute the inner loop iterative calculation of the loop optimization algorithm based on NTT transformation and the loop optimization algorithm based on INTT transformation through shared memory. By dividing the thread blocks and thread bundles with the optimal number and proportion, the hardware and software collaboration of homomorphic encrypted polynomial multiplication operations in the GPU multi-threaded parallel computing environment is optimized.
[0132] The embodiment of the present invention firstly adopts a nested loop optimization method based on the CUDA-GPU structural model and the parallel computing characteristics of the GPU to perform loop optimization on the structure of the NTT transformation algorithm to obtain a loop optimization algorithm based on the NTT transformation; adopts a nested loop optimization method to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation; secondly, according to the loop optimization algorithm based on the NTT transformation, the loop optimization algorithm based on the INTT transformation and the BFV primitive root pre-computation and pre-storage joint optimization algorithm, the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment is improved to obtain the improved homomorphic multiplication; finally, two ciphertext data to be processed are input; the improved homomorphic multiplication is used to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain the processed ciphertext result.
[0133] Embodiments of the present invention The present invention performs GPU acceleration on SEAL fully homomorphic encryption through a CUDA environment, optimizes an NTT transformation algorithm and an INTT transformation algorithm according to the parallel computing characteristics of a GPU, expands a multi-layer loop nesting structure, improves parallelism and computing resource utilization, reduces data dependence, and performs pre-processing and post-processing on required data, thereby greatly improving the GPU's execution efficiency of polynomial multiplication operations; according to the GPU thread hierarchical structure, the present invention combines a loop optimization algorithm based on an NTT transformation and a loop optimization algorithm based on an INTT transformation with GPU thread mapping optimization, and combines the two layers of NTT transformation and INTT transformation in the optimized polynomial multiplication with a GPU scheduling process, thereby improving the GPU parallel computing environment polynomial multiplication execution efficiency by dividing an optimal number of threads and thread blocks.
[0134] Figure 5 is a structural diagram of a homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multithreading mapping provided by an embodiment of the present invention, such as Figure 5 As shown, the homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multithreaded mapping may include the above Figure 4 The homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-thread mapping is shown. Optionally, the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multi-thread mapping may include a first processor 2001.
[0135] Optionally, the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multi-threaded mapping may also include a memory 2002 and a transceiver 2003 .
[0136] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0137] Combine the following Figure 5 The components of the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multi-threaded mapping are specifically introduced:
[0138] The first processor 2001 is the control center of the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multithreaded mapping, and can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs).
[0139] Optionally, the first processor 2001 can perform various functions of the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multi-threaded mapping by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.
[0140] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in FIG.
[0141] In a specific implementation, as an embodiment, the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multithreading mapping may also include multiple processors, such as Figure 5 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0142] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.
[0143] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be accessed through the interface circuit ( Figure 5 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0144] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0145] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 5 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0146] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and may be implemented through an interface circuit ( Figure 5 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0147] It should be noted that Figure 5 The structure of the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multi-threaded mapping shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0148] In addition, the technical effects of the homomorphic encryption acceleration device 510 based on polynomial multiplication optimization and GPU multi-threaded mapping can refer to the technical effects of the homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping described in the above method embodiment, which will not be repeated here.
[0149] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0150] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0151] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0152] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0153] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0154] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0155] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0157] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0158] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0159] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0160] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0161] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping, characterized in that: The method comprises: S1. Based on the CUDA-GPU structural model and according to the parallel computing characteristics of GPU, a nested loop optimization method is used to perform loop optimization on the structure of the NTT transformation algorithm to obtain a loop optimization algorithm based on the NTT transformation; a nested loop optimization method is used to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation; S2. According to the loop optimization algorithm based on NTT transformation, the loop optimization algorithm based on INTT transformation, and the BFV primitive root pre-computation and pre-storage joint optimization algorithm, the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment is improved to obtain an improved homomorphic multiplication; S3. Input two ciphertext data to be processed; use the improved homomorphic multiplication to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain a processed ciphertext result.
2. According to claim 1, the homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping is characterized in that: The S1 adopts a nested loop optimization method to perform loop optimization on the NTT transformation algorithm structure to obtain a loop optimization algorithm based on NTT transformation, including: The three-layer loop based on the NTT transformation algorithm is nested by removing one layer of loop, and the three-layer loop is reduced to two layers, and only an outer loop that keeps loop iteration and an inner loop that keeps internal high parallelism is retained, so as to obtain a loop optimization algorithm based on NTT transformation.
3. The homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multithreaded mapping according to claim 2 is characterized in that: The implementation process of the loop optimization algorithm based on NTT transformation includes: The input pointer points to the array storing the polynomial coefficients, the number of polynomials, the number of coefficient moduli, the power of the modulus of the polynomial, the data structure based on the NTT transformation, the signal for selecting the root power or the inverse root power, the thread index, and the value of the polynomial modulus; Calculate the current layer interval according to the power of the modulus of the polynomial; Calculate the root index and coefficient index according to the current layer interval and thread index; Use the polynomial number for loop processing, judge by the number of coefficient moduli, select the corresponding index root power and modulus, and end the loop when the polynomial number is satisfied, and output the result based on NTT transformation.
4. The homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multithreaded mapping according to claim 1 is characterized in that: The S1 adopts a nested loop optimization method to perform loop optimization on the structure of the INTT transformation algorithm to obtain a loop optimization algorithm based on the INTT transformation, including: The three-layer loop based on the INTT transformation algorithm is nested by removing one layer of loop, and the three-layer loop is reduced to two layers, and only an outer loop that keeps loop iteration and an inner loop that keeps internal high parallelism is retained, so as to obtain a loop optimization algorithm based on the INTT transformation.
5. The homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multithreaded mapping according to claim 4 is characterized in that: The implementation process of the loop optimization algorithm based on INTT transformation includes: The input pointer points to the array storing the polynomial coefficients, the number of polynomials, the number of coefficient moduli, the power of the modulus of the polynomial, the data structure of the INTT transformation, the signal for selecting the root power or the inverse root power, the thread index, and the value of the polynomial modulus; Calculate the current layer interval according to the power of the modulus of the polynomial; Calculate the root index and coefficient index based on the current layer interval, the power of the modulus of the polynomial, and the thread index; Use the polynomial number for loop processing, judge by the number of coefficient moduli, select the corresponding index root power and modulus, and end the loop when the polynomial number is satisfied, and output the result of INTT transformation.
6. The homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multithreaded mapping according to claim 1 is characterized in that: The implementation process of the improved homomorphic multiplication includes: Input two ciphertext data; wherein the ciphertext data includes: first ciphertext data and second ciphertext data; Preprocess the two ciphertext data using the preset modulus base and NTTTables iterator to obtain the preprocessing result; The first ciphertext data and the second ciphertext data are processed respectively using a cyclic optimization algorithm based on NTT transformation to obtain an NTT transformation result of the first ciphertext data and an NTT transformation result of the second ciphertext data; Using polynomial multiplication to calculate the preprocessing result, the first ciphertext data and the second ciphertext data to obtain a calculation result; Using a cyclic optimization algorithm based on INTT transformation to process the NTT transformation result of the first ciphertext data and the NTT transformation result of the second ciphertext data respectively, to obtain the INTT transformation result of the first ciphertext data and the INTT transformation result of the second ciphertext data; According to the calculation result, the INTT transformation result of the first ciphertext data and the INTT transformation result of the second ciphertext data are relinearized by using the modulus-to-digital basis conversion to obtain the storage result of the first ciphertext data.
7. The homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multithreaded mapping according to claim 1 is characterized in that: The input of S3 is two ciphertext data to be processed; After the step of using the improved homomorphic multiplication to process two ciphertext data to be processed to obtain the processed ciphertext result, the method further includes: According to the hierarchical structure characteristics of GPU multi-threaded parallel computing, each thread block executes the outer loop iterative calculation of the loop optimization algorithm based on NTT transformation and the loop optimization algorithm based on INTT transformation, and the thread bundles inside the thread block execute the inner loop iterative calculation of the loop optimization algorithm based on NTT transformation and the loop optimization algorithm based on INTT transformation through shared memory. By dividing the thread blocks and thread bundles with the optimal number and proportion, the hardware and software collaboration of homomorphic encrypted polynomial multiplication operations in the GPU multi-threaded parallel computing environment is optimized.
8. A homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping, the homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping is used to implement the homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-threaded mapping as described in any one of claims 1-7, characterized in that: The device comprises: The first acquisition unit is used to optimize the loop of the NTT transformation algorithm structure based on the CUDA-GPU structure model and the parallel computing characteristics of the GPU by using a nested loop optimization method to obtain a loop optimization algorithm based on the NTT transformation; and optimize the loop of the INTT transformation algorithm structure based on the nested loop optimization method to obtain a loop optimization algorithm based on the INTT transformation; A second acquisition unit is used to improve the homomorphic multiplication in the homomorphic encryption operation in the GPU parallel computing environment according to the loop optimization algorithm based on the NTT transformation, the loop optimization algorithm based on the INTT transformation, and the BFV primitive root pre-computation and pre-storage joint optimization algorithm to obtain an improved homomorphic multiplication; The third acquisition unit is used to input two ciphertext data to be processed; use the improved homomorphic multiplication to perform polynomial multiplication calculation on the two ciphertext data to be processed to obtain a processed ciphertext result.
9. A homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multithreaded mapping, characterized in that: The homomorphic encryption acceleration device based on polynomial multiplication optimization and GPU multi-threaded mapping includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
GPU multi-thread parallel Hawk algorithm acceleration method
CN122044804A
A GPU multi-thread parallel-based hawk algorithm acceleration method
CN122044804B