Key switching method and system based on graphics processor
By designing high-parallel, high-throughput number theory transformation, precise basis transformation and internal product cores in the graphics processor, the problem of low key switching calculation efficiency in the existing technology is solved, efficient key switching is achieved, and the application of fully homomorphic encryption computing is promoted.
Patent Information
- Application Number
- CN202510297418.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing GPU-based key switching algorithm is difficult to meet the current computing efficiency requirements, and the FPGA-based key switching accelerator has the problem of insufficient resources on-chip, which cannot achieve large-scale deployment and application.
Design high-parallel, high-throughput number-theoretical transformation, precise basis transformation and internal product kernels, and implement high-throughput key switching by executing these kernel functions in the graphics processor.
By implementing large-scale parallel computing on the GPU, the computing efficiency of key switching is improved, flexible parameter selection is supported, and the delay of single calculation is reduced, which promotes the application promotion of fully homomorphic encryption computing.
Smart Images

Figure CN120150918A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security, and more specifically, relates to a key switching method and system based on a graphics processing unit. Background Art
[0002] Data has been listed as one of the most important production factors, basic resources, and strategic resources. However, privacy leakage problems will occur during the data circulation process, and data privacy protection has become a key issue of common concern in the academic and industrial fields. Fully Homomorphic Encryption (FHE) is a reliable, general, and efficient user data privacy protection solution that allows a server to perform calculations on encrypted data, thereby maintaining the confidentiality of the data while processing it.
[0003] In FHE, key switching can be used to implement many homomorphic operation kernels such as homomorphic multiplication, homomorphic rotation, homomorphic conjugation, and automorphism kernel functions of FHE. It is the most core and time-consuming core kernel function among many FHE kernel functions. However, the existing FHE key switching methods and devices mainly have two deficiencies: First, the existing methods used for the key switching scheme mainly adopt the Hybrid key switching algorithm introduced in the article "Better Bootstrapping for Approximate Homomorphic Encryption", which contains a large number of high-width number theory transformation kernel functions, resulting in a performance bottleneck of the key switching kernel function on the GPU. Second, the existing number theory transformation algorithms used in the GPU-based key switching algorithm all adopt a naive number theory transformation calculation scheme, and there are a large number of pipeline stalls during the calculation process, bringing serious performance challenges to the key switching kernel function on the GPU. Third, the existing FPGA-based key switching accelerator has the problem of insufficient on-chip resources, and it is impossible to realize the large-scale deployment and application of the key switching algorithm on the FPGA.
[0004] That is, the existing GPU-based key switching algorithms are difficult to meet the current computational efficiency requirements. Summary of the Invention
[0005] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a key switching method and system based on a graphics processing unit, aiming to achieve high-throughput key switching by designing high-parallel, high-throughput number theory transformation, exact basis conversion, and inner product kernels, and solve the technical problem that the existing GPU-based key switching algorithms are difficult to meet the current computational efficiency requirements.
[0006] To achieve the above object, according to one aspect of the present invention, there is provided a key switching method based on a graphics processing unit, including: performing the following steps in the graphics processing unit:
[0007] S1: Perform a number-theoretic inverse transform kernel function on the point-value polynomial of the input residue number system base Q l to convert the point-value polynomial into a coefficient polynomial
[0008] S2: Perform a modulo lifting kernel function on the coefficient polynomial to split the base Q of the coefficient polynomial into multiple parts and convert them respectively to the residue number system base T, obtaining multiple groups of coefficient polynomials l
[0009] S3: Perform a number-theoretic transform kernel function on the multiple groups of coefficient polynomials to convert them into point-value polynomials
[0010] S4: Perform an inner product kernel function on the point-value polynomials and the matrices A and B in the key swk = (A, B) for key switching respectively, obtaining the point-value polynomial obtained by taking the inner product with matrix A and the point-value polynomial obtained by taking the inner product with matrix B where and are both multiple groups of point-value polynomials on base T;
[0011] S5: Perform a number-theoretic inverse transform kernel function on the point-value polynomials respectively, to convert the multiple groups of point-value polynomials into coefficient polynomials s a , s b ;
[0012] S6: Perform a modulo lifting kernel function on the coefficient polynomial s a to merge the multiple groups of coefficient polynomials s a and convert them from base T to base PQ l to obtain the polynomial corresponding to s a Perform a modulo lifting kernel function on the coefficient polynomial s to merge the multiple groups of polynomials s b and convert them from base T to base PQ b to obtain the polynomial corresponding to s l b
[0013] S7: Perform respectively on sa The corresponding polynomial and s b The corresponding polynomial Execute the modular reduction kernel function with the coefficient polynomials and Convert from the basis PQ l to the basis Q l .
[0014] In one embodiment, the S1 includes:
[0015] Suppose the input k point-value polynomials Each point-value polynomial has a dimension of N, and the number of modular numbers in the basis B is num; the k point-value polynomials The element [a i B in them are all sliced into a matrix of N = N 1 *N 2 . Take as the input of the number-theoretic inverse transform kernel function for N points on the GPU, and calculate according to the formula . And is calculated using the number-theoretic inverse transform kernel function, which is the naive matrix Hadamard product;
[0016] The calculation process of the number-theoretic inverse transform kernel function is as follows:
[0017] Start the first kernel function, allocate a three-dimensional GPU block with dimensions {N 1 / 2 t , num, k}, allocate 128 threads in each block, each block reads in N 2 *2 t data, the x dimension of the block reads data addressed by offset within the polynomial N, the y dimension addresses data by offset of the polynomial modulus, and the z dimension addresses data by offset of the number of polynomials; allocate a shared memory space of 256 * data bit width for each block, and read the N 2 data into the shared memory within the block in row-major order; calculate the N 2 point number-theoretic transform within each block, where each thread performs the calculation of the radix-2 butterfly unit, and the intermediate data of the iteration is stored using the shared memory; read the data from the shared memory and perform the Hadamard product with the constant matrix W 2 in the global memory, and write the x dimension of the GPU block back to the global memory from the register in transposed form and in column-major order;
[0018] Start the second kernel function, allocate a dimension of {N 2 / 2 t , a 3D GPU block of {num, k}, with 128 threads allocated in each block, and each block reads in N 1 *2 t data. The x - dimension of the block reads data addressed by offset within the polynomial N, the y - dimension addresses data by polynomial modulus, and the z - dimension addresses data by the number of polynomials; allocate N 1 * data - bit - width of shared memory space for each block. Each block reads N 1 data in each row of the N data into the shared memory of the block in row - major order; calculate the N 1 - point number - theoretic transform within each block, where each thread performs the calculation of the radix - 2 butterfly unit, and the intermediate data during iteration is stored using shared memory; finally, write the calculation result back to the global memory in row - major form.
[0019] In one of the embodiments, the S3 includes:
[0020] Suppose the input k coefficient polynomials [a 0 B , [a 1 B , …, [a k-1 B each have a coefficient polynomial dimension of N, and the number of moduli in the basis B is num; divide each coefficient polynomial [a i B into a matrix X of N = N 1 *N 2 , take X as the input of the N - point number - theoretic transform, and perform the number - theoretic transform calculation according to the formula . W 1 , W 2 , W 3 are matrices of N 1 ×N 1 , N 1 ×N 2 , N 2 ×N 2 respectively. ×W 1 and ×W 3 are implemented using the calculation of the number - theoretic transform, and ⊙W 2 is the naive matrix Hadamard product;
[0021] The calculation process of the number - theoretic transform is as follows:
[0022] Start the first kernel function, allocate a 3D GPU block of {N 2 / 2 t , num, k}, with 128 threads allocated in each block, and each block reads in N 1 *2 t For each block, the data is read in the x - dimension by offset addressing within polynomial N, in the y - dimension by offset addressing with polynomial modulus, and in the z - dimension by offset addressing with the number of polynomials; allocate a shared memory space of 256 * data bit - width for each block, and read N 1 data into the shared memory within the block in row - major order; calculate the number - theoretic transform of N 1 points within each block, where each thread performs the calculation of the radix - 2 butterfly unit, and the intermediate data during iteration is stored using the shared memory; read the data from the shared memory and perform the Hadamard product with the constant matrix W 2 in the global memory; write the data back to the global memory from the register in column - major order with the x - dimension of the GPU block in transposed form.
[0023] Launch the second kernel function, allocate a three - dimensional GPU block with dimensions {N 1 / 2 t , num, k}, allocate 128 threads within each block, and each block reads N 2 *2 t data. The data is read in the x - dimension of the block by offset addressing within polynomial N, in the y - dimension by offset addressing with polynomial modulus, and in the z - dimension by offset addressing with the number of polynomials; allocate a shared memory space of N 2 * data bit - width for each block, and each block reads N data for each row of the N 2 data into the shared memory within the block in row - major order; calculate the number - theoretic transform of N 2 points within each block, where each thread performs the calculation of the radix - 2 butterfly unit, and the intermediate data during iteration is stored using the shared memory, and write the calculation result back to the global memory in row - major form.
[0024] In one embodiment, the S4 includes:
[0025] Receiving multiple point - value polynomials of the residual number system basis T from the GPU global memory data and the key swk = U = [A, B] for key switching from the GPU global memory, where both A and B are matrices of Calculating the inner - product operation between the input point - value polynomial and the key swk, and storing the output result of the inner - product operation in the GPU global memory;
[0026] The inner - product operation between the input point - value polynomial and the key swk includes:
[0027] Launching a kernel function, allocating dimensions of A three-dimensional GPU block, where block_size threads are allocated within each block, and block_size columns of data are read. The x-dimension of the block reads data addressed by offset within polynomial N, the y-dimension reads data addressed by offset with polynomial modulus, and the z-dimension reads data addressed by offset with the number of key column blocks.
[0028] Read the ciphertext matrix of the residue number system and the key switching key in batches from the global memory into the register file; calculate [a i T The multiplication result with the element A i,block_dim.z in matrix A, and accumulate it d l times, where 0 ≤ i < d l . Store the intermediate accumulation result and the output result in the register file to obtain multiple matrices of dimension r'*N; calculate [a i T The multiplication and accumulation result with the element B i,block_dim.z in matrix B, and accumulate it d l times to obtain multiple matrices of dimension r'*N; then separately reduce the accumulated matrices below T to obtain and And write the two results and back to the global memory.
[0029] In one embodiment, the S7 includes:
[0030] Respectively execute the modulo reduction kernel function on the input k coefficient polynomials of the basis PQ l from the GPU global memory, where the basis PQ and . Calculate the result of the polynomial data in the basis PQ l ={p 0 ,…,p r-1 ,p r ,…,p r+s-1} under its sub-basis I l ={p 1 ,…,p 0 ,…,p r-1}, and output multiple coefficient polynomials under the basis Q l and store them in the GPU global memory. The number of moduli in the basis PQ l is r + s, and the number of moduli in the sub-basis I 1 is r;
[0031] Execute the modulo reduction kernel function as follows:
[0032] Start the fourth kernel function, allocate a three-dimensional GPU block with dimensions {N / block_size, r, k}, allocate block_size threads within each block, and read block_size columns of data. The x-dimension of the block reads data addressed by offset within the polynomial N, the y-dimension addresses data by offset with respect to the polynomial modulus, and the z-dimension addresses data by offset with respect to the number of polynomials; Batch-read k s*N-dimensional ciphertext matrices of the residual number system from global memory into the register file, calculate the multiplication between the ciphertext and the pre-computed modulus constant vector, output the s*N-dimensional matrix and store it in the register file; Read the s*N-dimensional matrix from the register file, perform floating-point division with the corresponding modulus, sum the results, and finally round to output a 1*N-dimensional integer vector; Read the 1*N-dimensional integer vector from the register file and multiply it with the input basis in global memory to output the fourth r*N-dimensional matrix and store the result in the register file; Read the fourth r*N-dimensional matrix from the register file and multiply it with the basis transformation matrix in global memory to output the fifth r*N-dimensional matrix;
[0033] Read the fourth r*N-dimensional matrix and the fifth r*N-dimensional matrix from the register file, perform subtraction, reduce the output to below the basis O, obtain the sixth r*N-dimensional matrix and save it in the register file; Read the sixth r*N-dimensional matrix from the register file, denoted as Calculate the subtraction between the values of the original coefficient polynomials [a 0 I , [a 1 I , …, [a k-1 I on the basis I 1 and the coefficient polynomial , and calculate the Shoup modular multiplication operation between the result and the constant to obtain and write it back to global memory as the final result of the basis transformation.
[0034] In one embodiment, the S2 includes: performing a modular lifting kernel function calculation on the input k coefficient polynomials [a 0 I , [a 1 I , …, [a k-1 I from the basis I in GPU global memory to calculate the polynomial data [a 0 , …, p r-1} in the basis I = {p 0 I , [a 1 I ,…,[a k-1 I In the basis O = {q 0 ,…,q s-1}, the expressions are obtained as multiple coefficient polynomials in the basis O and stored in the GPU global memory, where the number of moduli in basis I is r and the number of moduli in basis O is s; among them, the input coefficient polynomial is The basis I is the basis Q l ; the basis O is the basis T;
[0035] The process of performing the modular lifting kernel function calculation is as follows:
[0036] Start the third kernel function, allocate a three-dimensional GPU block with dimensions {N / block_size, s, k}, allocate block_size threads within each block, and read block_size columns of data. The x-dimension of the block reads data addressed by offset within the polynomial N, the y-dimension reads data addressed by offset of the polynomial modulus, and the z-dimension reads data addressed by offset of the number of polynomials;
[0037] Batch read l r*N-dimensional residual number system ciphertext matrices from the global memory into the register file, calculate the multiplication between the ciphertext and the pre-computed modulus constant vector to obtain an r*N-dimensional matrix, and store it in the register file; read the r*N-dimensional matrix from the register file, perform floating-point division with the corresponding modulus and sum the results to obtain a 1*N-dimensional integer vector; read the 1*N-dimensional integer vector from the register file and multiply it with the input basis in the global memory to obtain the first s*N-dimensional matrix, and store the result in the register file; read the first s*N-dimensional matrix from the register file and multiply it with the basis transformation matrix in the global memory to obtain the second s*N-dimensional matrix; read the first s*N-dimensional matrix and the second s*N-dimensional matrix from the register file, perform subtraction, and then reduce the output to below the basis O to obtain the third s*N-dimensional matrix and write it back to the global memory.
[0038] In one embodiment, the S6 includes: for the input k coefficient polynomials [a 0 I ,[a 1 I ,…,[a k-1 I perform the modular lifting kernel function calculation to calculate the polynomial data [a 0 ,…,p r-1} in the basis I = {p 0 I ,[a 1 I ,…,[ak-1 I The expression under the basis O = {q 0 , …, q s-1} obtains multiple coefficient polynomials under the basis O and stores them in the GPU global memory, where the number of moduli in the I basis is r, and the number of moduli in the O basis is s; among them, the input coefficient polynomials are s a and s b ; the basis I is the basis T; the basis O is the basis PQ l ;
[0039] The process of executing the modular lifting kernel function calculation is as follows:
[0040] Start the third kernel function, allocate a three-dimensional GPU block with dimensions {N / block_size, s, k}, allocate block_size threads within each block, and read block_size columns of data. The x dimension of the block reads data addressed by offset within the polynomial N, the y dimension addresses data by offset of the polynomial modulus, and the z dimension addresses data by offset of the number of polynomials;
[0041] Batch read l r*N-dimensional residual number system ciphertext matrices from the global memory into the register file, calculate the multiplication between the ciphertext and the pre-computed modulus constant vector to obtain an r*N-dimensional matrix, and store it in the register file; read the r*N-dimensional matrix from the register file, perform floating-point division with the corresponding modulus and sum the results to obtain a 1*N-dimensional integer vector; read the 1*N-dimensional integer vector from the register file and multiply it with the input basis in the global memory to obtain the first s*N-dimensional matrix, and store the result in the register file; read the first s*N-dimensional matrix from the register file and multiply it with the basis conversion matrix in the global memory to obtain the second s*N-dimensional matrix; read the first s*N-dimensional matrix and the second s*N-dimensional matrix from the register file, perform subtraction, and then reduce the output to below the basis O to obtain the third s*N-dimensional matrix and write it back to the global memory.
[0042] According to another aspect of the present invention, a key switching system is provided, including a central processing unit and a graphics processing unit. The memory stores a computer program. The central processing unit is used to schedule the data stream between the central processing unit and the graphics processing unit and allocate a computing control flow for the graphics processing unit; the graphics processing unit is used to implement the steps of the above key switching method when executing the computer program.
[0043] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects can be achieved:
[0044] (1) The present invention provides a key switching method based on a graphics processing unit. The number theory transform kernel function, the exact basis conversion kernel function, and the inner product kernel function implemented on the GPU are assembled in a corresponding order according to the needs of key switching to implement key switching in a high-throughput fully homomorphic encryption technology based on the GPU. On the one hand, the present invention realizes large-scale parallelism among the coefficients and moduli of the fully homomorphic encryption ciphertext polynomials by implementing three kernel functions on the GPU, thereby realizing large-scale parallelism and high-throughput design of the fully homomorphic encryption key switching kernel function and improving the computing efficiency. On the other hand, the present invention supports flexible parameter selection, allowing users to achieve high-throughput key switching of any cryptographic parameters without any code modification.
[0045] (2) The present invention provides a key switching method based on a graphics processing unit. For the three core kernel functions of number theory transform, exact basis conversion, and inner product in the fully homomorphic encryption key switching kernel function, a 3D kernel function startup method is designed. By minimizing the computing tasks of each thread of the GPU and the parallel characteristics of the GPU single instruction multiple thread architecture, the parallelism of the three kernel functions of number theory transform, exact basis conversion, and inner product among polynomials, between moduli, and between polynomials is greatly improved, and the throughput of the three core kernel functions is greatly improved. Finally, through the kernel function assembly and scheduling program, the throughput of the key switching algorithm on the GPU is greatly improved, the latency of a single calculation is reduced, and the application promotion of fully homomorphic encryption calculation is promoted. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the high-throughput key switching method based on the GPU for the solution of the present invention;
[0047] Figure 2 Schematic diagram of the high-throughput number theory transform method based on the GPU for the solution of the present invention;
[0048] Figure 3 Schematic diagram of the high-throughput exact basis conversion method based on the GPU for the solution of the present invention;
[0049] Figure 4 Schematic diagram of the high-throughput inner product method based on the GPU for the solution of the present invention;
[0050] Figure 5 Architecture diagram of the high-throughput key switching device based on the GPU for the solution of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0052] Symbol description:
[0053]
[0054]
[0055] As Figure 1 shown, a high-throughput fully homomorphic encryption key switching method based on GPU according to the present invention, the method includes: implementing high-throughput number theory transformation, exact basis conversion, and inner product kernel function on the GPU, and assembling them to achieve high-throughput fully homomorphic encryption key switching. Determine the cryptographic parameters of the key switching kernel function, perform pre-computation in the CPU, and store the results in the GPU global memory, including the following steps.
[0056] Step S1: Input the point value polynomial on the basis Q l and execute the number theory inverse transform kernel function to convert the polynomial into the coefficient state
[0057] Step S2: Execute the ModUp kernel function to split the polynomial output in step S with the basis Q l into parts, and respectively convert them to T, and output groups of polynomials
[0058] Step S3: Execute the number theory transform kernel function to convert the groups of polynomials output in step S3 into point value polynomials, and output
[0059] Step S5: Execute the inner product kernel function to respectively execute the inner product kernel function on the groups of point value polynomials output in step S4 and the key switching key swk=(A, B), and output where and are respectively groups of polynomials on T, and K is the number of moduli of the key parameter P.
[0060] Step S4: Execute the number theory inverse transform kernel function on the d groups of polynomials output in step S5 Convert them to coefficient polynomials respectively and output s a , s b .
[0061] Step S5: Execute the ModUp kernel function to merge the d groups of polynomials s a , s b output by Step S6, and convert them from the d groups of bases T to the base PQ l , and output polynomials and
[0062] Step S6: Execute the ModDown kernel function to convert the polynomials and output by Step S7 from the base PQ l to the base Q l .
[0063] As Figure 2 shown, a high-throughput number-theoretic transform method based on GPU according to the present invention includes two types: number-theoretic transform and number-theoretic inverse transform.
[0064] Regarding the number-theoretic inverse transform: Input the residue number system representation (Residue Number System) polynomials of several point-value polynomials to be calculated , where the dimension of each polynomial is N, the number of moduli of the base B where the residue number system polynomials are located is num, and the number of polynomials is k. Divide each polynomial [a i B into a matrix with N = N 1 * N 2 . Use as the input of the N-point number-theoretic inverse transform, and perform number-theoretic inverse transform calculation according to the formula . W i and are the inverses of the matrix W in the number-theoretic transform calculation steps respectively. can be calculated by number-theoretic inverse transform iteration, and
[0065] is the naive matrix Hadamard product.
[0066] The calculation steps of the number-theoretic inverse transform are as follows:
[0066] Start the first kernel function, allocate a three-dimensional GPU block with dimensions {N 1 / 2 t , num, k}, allocate 128 threads in each block, and each block reads N 2 * 2 tThe x dimension of the block reads data with offset addressing in polynomial N, the y dimension reads data with offset addressing in polynomial modulus, and the z dimension reads data with offset addressing in polynomial number. A shared memory space of 256*data bit width is allocated to each block, and N 2 The data is read into the shared memory in the block in row-major order; N is calculated in each block. 2 Point number theory transformation, where each thread performs the calculation of the butterfly unit of radix-2, and the iterative intermediate data is stored using shared memory; the data output from the previous step is read from the shared memory and compared with the constant matrix W in the global memory 2 Do the Hadamard product and then write the x dimension of the GPU block from the register back to global memory in column-major order in transposed form. Start the second kernel function and allocate the dimension {N 2 / 2 t ,num,k} three-dimensional GPU blocks, each block is assigned 128 threads, and each block reads N 1 *2 t The x dimension of the block reads data with polynomial N offset addressing, the y dimension addresses data with polynomial modulus, and the z dimension addresses data with the number of polynomials. N is allocated for each block. 1 * Data bit width shared memory space, each block will be N of each row of N data output in the previous step 1 The data is read into the shared memory in the block in row-major order. In each block, N 1 Point number theory transformation, in which each thread performs the calculation of the butterfly unit of radix-2, stores the iterative intermediate data using shared memory, and writes the final calculation result back to the global memory in row-major form.
[0067] About number theory transformation: Input several polynomials with coefficients to be calculated, and represent them in residual number system (Residue Number System) polynomial [a 0 ] B ,[a 1 ] B ,…,[a k-1 ] B , where each polynomial dimension is N, the number of moduli of the base B in which the residual number system polynomial is located is num, the number of polynomials is k, and each polynomial [a i ] B For segmentation N=N 1 *N 2 The matrix X is used as the input of the N-point number theory transformation, and according to the formula Perform number theory transformation calculation, W 1 ,W 2 ,W 3 N 1 ×N 1,N 1 ×N 2 ,N 2 ×N 2 matrix of ×W 1 and ×W 3 can be implemented by the calculation of number-theoretic transform, ⊙W 2 is the naive matrix Hadamard product.
[0068] The process of number-theoretic transform calculation is as follows:
[0069] Start the first kernel function, allocate a three-dimensional GPU block with dimensions {N 2 / 2 t , num, k}, allocate 128 threads in each block, and each block reads N 1 *2 t data. The x-dimension of the block reads data addressed by offset within polynomial N, the y-dimension reads data addressed by offset with polynomial modulus, and the z-dimension reads data addressed by offset with the number of polynomials; allocate a shared memory space of 256 * data bit width for each block, and read N 1 data into the shared memory within the block in row-major order; calculate the number-theoretic transform of N 1 points in each block, where each thread performs the calculation of the radix-2 butterfly unit, and the intermediate data of the iteration is stored using the shared memory; read the data from the shared memory and perform the Hadamard product with the constant matrix W 2 in the global memory, and then write the x-dimension of the GPU block back to the global memory from the register in transposed form and column-major order. Start the second kernel function, allocate a three-dimensional GPU block with dimensions {N 1 / 2 t , num, k}, allocate 128 threads in each block, and each block reads N 2 *2 t data. The x-dimension of the block reads data addressed by offset within polynomial N, the y-dimension reads data addressed by polynomial modulus, and the z-dimension reads data addressed by the number of polynomials. Allocate a shared memory space of N 2 * data bit width for each block. Each block reads N 2 data in each row of the N data output in the previous step into the shared memory within the block in row-major order. Calculate the number-theoretic transform of N 2 points in each block, where each thread performs the calculation of the radix-2 butterfly unit, and the intermediate data of the iteration is stored using the shared memory, and write the calculation result output in the previous step back to the global memory in row-major form.
[0070] As Figure 3 shown, it is a high-throughput and accurate base conversion method based on GPU according to the present invention, and the method includes ModUP and ModDown.
[0071] Regarding ModUP, in S2, the following steps are included: For the k input coefficient polynomials [a 0 I , [a 1 I , …, [a k-1 I Execute the modular lifting kernel function calculation to calculate the polynomial data [a 0 , …, p r-1} in the basis I = {p 0 I , [a 1 I , …, [a k-1 I Obtain multiple coefficient polynomials in the basis O = {q 0 , …, q s-1} and store them in the GPU global memory. The number of moduli in the I basis is r, and the number of moduli in the O basis is s; among them, the input coefficient polynomial is The basis I is the basis Q l ; The basis O is the basis T.
[0072] The process of executing the modular lifting kernel function calculation is as follows: Start the third kernel function, allocate a three-dimensional GPU block with dimensions {N / block_size, s, k}, allocate block_size threads within each block, and read block_size columns of data. The x dimension of the block reads data addressed by offset within the polynomial N, the y dimension addresses data by offset of the polynomial modulus, and the z dimension addresses data by offset of the number of polynomials. Batch read the k r*N-dimensional residual number system ciphertext matrices from the global memory into the register file, calculate the multiplication between the ciphertext and the pre-computed modulus constant vector to obtain an r*N-dimensional matrix, and store it in the register file; Read the r*N-dimensional matrix from the register file, perform floating-point division with the corresponding modulus and sum the results to obtain a 1*N-dimensional integer vector; Read the 1*N-dimensional integer vector from the register file, and multiply it by the input basis in the global memory to obtain the first s*N-dimensional matrix, and store the result in the register file; Read the first s*N-dimensional matrix from the register file, and multiply it by the basis conversion matrix in the global memory to obtain the second s*N-dimensional matrix; Read the first s*N-dimensional matrix and the second s*N-dimensional matrix from the register file, perform subtraction, and then reduce the output to below the basis O to obtain the third s*N-dimensional matrix and write it back to the global memory.
[0073] Regarding ModUP, in S6, the following steps are included: For the k input coefficient polynomials [a 0 I ,[a 1 I ,…,[a k-1 I Execute the modular lifting kernel function calculation to calculate the polynomial data [a 0 ,…,p r-1} in the basis I = {p 0 I ,[a 1 I ,…,[a k-1 I The expression under the basis O = {q 0 ,…,q s-1} obtains multiple coefficient polynomials under the basis O and stores them in the GPU global memory, where the number of moduli in the I basis is r, and the number of moduli in the O basis is s; among them, the input coefficient polynomials are s a and s b ; the basis I is the basis T; the basis O is the basis PQ l .
[0074] The process of the above-mentioned execution of the modular lifting kernel function calculation is as follows: Start the third kernel function, allocate a three-dimensional GPU block with dimensions {N / block_size, s, k}, allocate block_size threads in each block, and read block_size columns of data. The x dimension of the block reads data addressed by offset within the polynomial N, the y dimension reads data addressed by offset of the polynomial modulus, and the z dimension reads data addressed by offset of the number of polynomials. Batch read k r*N-dimensional residual number system ciphertext matrices from the global memory into the register file, calculate the multiplication between the ciphertext and the pre-calculated modulus constant vector to obtain an r*N-dimensional matrix, and store it in the register file; read the r*N-dimensional matrix from the register file, perform floating-point division with the corresponding modulus and sum the results to obtain a 1*N-dimensional integer vector; read the 1*N-dimensional integer vector from the register file, and multiply it by the input basis in the global memory to obtain the first s*N-dimensional matrix, and store the result in the register file; read the first s*N-dimensional matrix from the register file, and multiply it by the basis conversion matrix in the global memory to obtain the second s*N-dimensional matrix; read the first s*N-dimensional matrix and the second s*N-dimensional matrix from the register file, perform subtraction, and then reduce the output to below the basis O to obtain the third s*N-dimensional matrix and write it back to the global memory.
[0075] Regarding ModDown, S7 includes: respectively execute the modular reduction kernel function on the input k coefficient polynomials l of the basis PQ and from the GPU global memory, where the basis PQl = {p 0 , …, p r-1 , p r , …, p r+s-1}, calculate the results of polynomial data in the basis PQ l in its sub - basis I 1 = {p 0 , …, p r-1}, output the basis Q l under multiple coefficient polynomials and store them in the GPU global memory, where the number of moduli in the basis PQ l is r + s, and the number of moduli in the basis I 1 is r.
[0076] Execute the modular reduction kernel function as follows: Start the fourth kernel function, allocate a three - dimensional GPU block with dimensions {N / block_size, r, k}, allocate block_size threads within each block, and read block_size columns of data. The x - dimension of the block reads data addressed by offset within polynomial N, the y - dimension reads data addressed by offset of polynomial modulus, and the z - dimension reads data addressed by offset of the number of polynomials; Batch - read the k s * N - dimensional residual number system ciphertext matrices from the global memory into the register file, calculate the multiplication between the ciphertext and the pre - calculated modulus constant vector, output the s * N - dimensional matrix and store it in the register file; Read the s * N - dimensional matrix from the register file, perform floating - point division with the corresponding modulus, sum the results, and finally round to output a 1 * N - dimensional integer vector; Read the 1 * N - dimensional integer vector from the register file and multiply it with the input basis in the global memory, output the fourth r * N - dimensional matrix and store the result in the register file; Read the fourth r * N - dimensional matrix from the register file and multiply it with the basis conversion matrix in the global memory, output the fifth r * N - dimensional matrix; Read the fourth r * N - dimensional matrix and the fifth r * N - dimensional matrix from the register file, and perform subtraction. Reduce the output below the basis O to obtain a sixth r * N - dimensional matrix and save it in the register file; Read the sixth r * N - dimensional matrix from the register file, denoted as Calculate the subtraction between the values of the original coefficient polynomials [a 0 I , [a 1 I , …, [a k-1 I on the basis I 1 and the coefficient polynomial , and calculate the Shoup modular multiplication operation with the constant to obtain and write it back to the global memory as the final result of the basis conversion.
[0077] As shown Figure 4 in the figure, a high-throughput inner product method based on GPU according to the present invention, the method includes the following steps.
[0078] Receiving multiple point value polynomials of the residual number system base T from the GPU global memory data and the key swk = U = [A, B] for key switching from the GPU global memory, where both A and B are matrices of Calculating the inner product operation of the input point value polynomial and the key swk, and storing the output inner product operation result in the GPU global memory.
[0079] The inner product operation of the input point value polynomial and the key swk includes: starting the kernel function, allocating a three-dimensional GPU block with a dimension of , allocating block_size threads in each block, and reading block_size columns of data. The x dimension of the block reads data addressed by offset within the polynomial N, the y dimension addresses data by offset of the polynomial modulus, and the z dimension addresses data by offset of the number of key column blocks. Reading the residual number system ciphertext matrix and the key switching key in batches from the global memory into the register file; calculating the multiplication result of [a i T and A i,block_dim.z , and accumulating d l times, where 0 ≤ i < d l , storing the intermediate accumulation result and the output result in the register file to obtain multiple matrices of r'*N dimension; calculating the multiplication accumulation result of [a i T and B i,block_dim.z , and accumulating d l times to obtain multiple matrices of r'*N dimension; then respectively reducing the accumulated matrices below T to obtain and and writing the two results and back to the global memory. Among them, A i,block_dim.z is the matrix A of the key, with a dimension of d*~d, and each element of the matrix is a polynomial. Addressing according to i and the z dimension of the GPU block, the obtained matrix element is a polynomial. B i,block_dim.z is the matrix B, and B is also a matrix composed of polynomials, with a dimension of d*~d, and each element of the matrix is a polynomial.
[0080] More specifically, receiving Coefficient polynomials of a residue number system base T from the GPU global memory, and a key switching key swk = U = [A, B] from the GPU global memory, where both A and B are matrices of Calculate the inner product operation of the input point value polynomial and multiple keys, and output the inner product operation results of one or more polynomials to the GPU global memory. The inner product operation process is as follows: Start the kernel function, allocate a three-dimensional GPU block with a dimension of , allocate block_size threads in each block, and read block_size columns of data. The x dimension of the block reads data addressed by offset within polynomial N, the y dimension addresses data by offset with respect to the polynomial modulus, and the z dimension addresses data by offset with the number of key column blocks. Read the residue number system ciphertext matrix and the key switching key in batches from the global memory into the register file, and calculate [a i T The multiplication result with A i,block_dim.z , and accumulate it d l times, where 0 ≤ i < d l . Store the intermediate accumulation result and the output result in two uint64 register files in the form of high and low bits. This step outputs matrices of dimension r' * N. Similarly, calculate the multiplication and accumulation result of [a i T with B i,block_dim.z . Use the Barrett reduction algorithm to reduce the output of the previous step to below T, and write the result back to the global memory.
[0081] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A key switching method based on a graphics processor, characterized in that: include: The following steps are performed in the graphics processor: S1: The residual number system basis Q of the input l Point-valued polynomials on Perform the inverse number-theoretic kernel function to transform the point-valued polynomial Convert to coefficient polynomial S2: Coefficient polynomial Perform the modular lifting kernel function to transform the coefficient polynomial Q l Divide into multiple parts and convert them to the residual number system basis T respectively to obtain multiple sets of coefficient polynomials S3: For multiple sets of coefficient polynomials Perform a number-theoretic transformation kernel function to convert it into a valued polynomial S4: Point value polynomials The inner product kernel function is performed on the matrix A and the matrix B in the key swk=(A,B) used for key switching to obtain the point value polynomial obtained by the inner product with the matrix A The point-valued polynomial obtained by the inner product of the sum and the matrix B in and They are all multiple sets of point-valued polynomials on the basis T; S5: Point-valued polynomials Execute the inverse number theory transform kernel function respectively, with multiple sets of point value polynomials Correspondingly converted into coefficient polynomial s a ,s b ; S6: coefficient polynomial s a Perform modular lifting kernel function to transform multiple sets of coefficient polynomials s a Merge and convert from basis T to basis PQ l On, get s a The corresponding polynomial For the coefficient polynomial s b Perform modular lifting kernel function to transform multiple sets of polynomials s b Merge and convert from basis T to basis PQ l On, get s b The corresponding polynomial S7: respectively for s a The corresponding polynomial and b The corresponding polynomial Perform a modular descent kernel function with coefficients polynomial and From the base PQ l Convert to basis Q l superior.
2. The key switching method based on a graphics processor according to claim 1, characterized in that: The S1 includes: Suppose the input k-point value polynomial Each point value polynomial in the dimension is N, and the number of moduli in the base B is num; k point value polynomials middle element[a i ] B Evenly divided into matrices of N = N1 * N2 Will As the input of the number theory inverse transform kernel function of N points on the GPU, and according to the formula Perform calculations, and It is calculated using the inverse transform kernel function of number theory. is the naive matrix Hadamard product; The calculation process of the inverse transform of number theory is: Start the first kernel function and assign the dimension {N1 / 2 t ,num,k} three-dimensional GPU blocks, each block is assigned 128 threads, and each block reads N2*2 t The x dimension of the block reads data with polynomial N offset addressing, the y dimension offset addressing data with polynomial modulus offset addressing, and the z dimension offset addressing data with the number of polynomials; a shared memory space of 256*data bit width is allocated to each block, and N2 data are read into the shared memory in the block in row-major order; N2 point number theory transformations are calculated in each block, where each thread performs the calculation of the radix-2 butterfly unit, and iterative intermediate data is stored using shared memory; data is read from the shared memory, and the Hadamard product is performed with the constant matrix W2 in the global memory, and the x dimension of the GPU block is written back to the global memory from the register in the form of transposition in column-major order; Start the second kernel function and assign the dimension {N2 / 2 t ,num,k} three-dimensional GPU blocks, each block is assigned 128 threads, and each block reads N1*2 t The x dimension of the block reads data addressed by the offset in the polynomial N, the y dimension addresses data by the modulus of the polynomial, and the z dimension addresses data by the number of polynomials. A shared memory space of N1*data bit width is allocated to each block, and each block reads N1 data in each row of N data into the shared memory in the block in row-major order. N1 point number theory transformations are calculated in each block, where each thread performs the calculation of the butterfly unit of radix-2, and the iterative intermediate data is stored in shared memory. Finally, the calculation results are written back to the global memory in row-major form.
3. The key switching method based on a graphics processor as claimed in claim 2, characterized in that: The S3 includes: Let the input k coefficient polynomial [a0] B ,[a1] B ,…,[a k-1 ] B Each coefficient polynomial has dimension N, and the number of moduli in base B is num; each coefficient polynomial [a i ] B Cut into a matrix X of N = N1*N2, use X as the input of the N-point number theory transformation, and follow the formula Perform number theory transformation calculations, W1, W2, W3 are matrices of N1×N1, N1×N2, N2×N2 respectively, ×W1 and ×W3 are implemented using number theory transformation calculations, ⊙W2 is a naive matrix Hadamard product; The calculation process of number theory transformation is as follows: Start the first kernel function and assign the dimension {N2 / 2 t ,num,k} three-dimensional GPU blocks, each block is assigned 128 threads, and each block reads N1*2 t The x dimension of the block reads data with polynomial N offset addressing, the y dimension offset addressing data with polynomial modulus offset addressing, and the z dimension offset addressing data with the number of polynomials; a shared memory space of 256*data bit width is allocated to each block, and N1 data are read into the shared memory in the block in row-major order; N1 point number theory transformations are calculated in each block, where each thread performs the calculation of the butterfly unit of radix-2, and the iterative intermediate data is stored in shared memory; data is read from the shared memory, and the Hadamard product is performed with the constant matrix W2 in the global memory, and the x dimension of the GPU block is written back to the global memory from the register in the form of transposition in column-major order; Start the second kernel function and assign the dimension {N1 / 2 t ,num,k} three-dimensional GPU blocks, each block is assigned 128 threads, and each block reads N2*2 t The x dimension of the block reads data addressed by the polynomial N internal offset, the y dimension addresses data by the polynomial modulus, and the z dimension addresses data by the number of polynomials; a shared memory space of N2*data bit width is allocated to each block, and each block reads N2 data in each row of N data into the shared memory within the block in row-major order; N2 point number theory transformations are calculated in each block, where each thread performs the calculation of the radix-2 butterfly unit, and the iterative intermediate data is stored using shared memory, and the calculation results are written back to the global memory in row-major form.
4. The key switching method based on a graphics processor as claimed in claim 1, characterized in that: The S4 includes: Receives multiple point-valued polynomials of the residual number system basis T from GPU global memory data and the key for key switching swk=U=[A,B] from the GPU global memory, where A and B are The matrix of Calculate the inner product of the input point value polynomial and the key swk, and store the output inner product result in the GPU global memory; The inner product operation of the input point value polynomial and the key swk includes: Start the kernel function and assign dimensions to 3D GPU blocks, block_size threads are allocated in each block, and block_size columns of data are read. The x dimension of the block reads data with offset addressing in polynomial N, the y dimension reads data with offset addressing in polynomial modulus, and the z dimension reads data with offset addressing in key column blocks. Read the residual number system ciphertext matrix and key switching key from global memory into the register file in batches; calculate [a i ] T and the element A in the matrix A i,block_dim.z The multiplication result of d is accumulated l times, where 0≤i <d l , store the intermediate accumulation results and output results in the register file, and obtain multiple r'*N-dimensional matrices; calculate [a i ] T and the element B in matrix B i,block_dim.z The multiplication and accumulation results of d l Multiple matrices of r'*N dimensions are obtained; then the accumulated matrices are reduced to below T to obtain and And the two results and Write back to global memory.
5. The key switching method based on a graphics processor as claimed in claim 1, characterized in that: The S7 includes: For each of the k input base PQs from the GPU global memory l The coefficient polynomial and Perform a modular descent kernel function, where the basis PQ l ={p0,…,p r-1 ,p r ,…,p r+s-1 }, calculate the basis PQ l The polynomial data in its sub-base I1={p0,…,p r-1 }, output basis Q l Multiple coefficient polynomials are generated and stored in GPU global memory, where the basis PQ l The number of modules in the base is r+s, and the number of modules in the base I1 is r; The execution modulus descent kernel function is as follows: Start the fourth kernel function, allocate a three-dimensional GPU block of dimension {N / block_size,r,k}, allocate block_size threads in each block, and read block_size columns of data. The x dimension of the block reads data with offset addressing within the polynomial N, the y dimension reads data with offset addressing by the polynomial modulus, and the z dimension reads data with offset addressing by the number of polynomials; reads k s*N-dimensional residual number system ciphertext matrices from the global memory into the register file in batches, and calculates the multiplication between the ciphertext and the pre-calculated modulus constant vector, outputs an s*N-dimensional matrix and stores it in the register file; reads the s*N-dimensional matrix from the register file, performs floating-point division on it and the corresponding modulus, sums the results, and finally rounds them off to output a 1*N-dimensional integer vector; reads a 1*N-dimensional integer vector from the register file and adds it to the input basis in the global memory Multiply, output the fourth r*N dimensional matrix and store the result in the register file; read the fourth r*N dimensional matrix from the register file, and multiply it with the basis transformation matrix in the global memory, and output the fifth r*N dimensional matrix; Read the fourth r*N dimensional matrix and the fifth r*N dimensional matrix from the register file, perform subtraction, reduce the output to below the basis O, obtain the sixth r*N dimensional matrix and save it in the register file; read the sixth r*N dimensional matrix from the register file, record it as Calculate the original coefficient polynomial [a0] I ,[a1] I ,…,[a k-1 ] I Values and coefficients of polynomials on basis I1 Subtract between and calculate the constant Shoup modular multiplication operation between them gives And write back to global memory as the final result of the base conversion.
6. The key switching method based on a graphics processor as claimed in claim 5, characterized in that: The S2 includes: k coefficient polynomials [a0] from the basis I in the GPU global memory for the input I ,[a1] I ,…,[a k-1 ] I Perform modular lifting kernel function to calculate the basis I = {p0, ..., p r-1 } in polynomial data[a0] I ,[a1] I ,…,[a k-1 ] I In the basis O={q0,…,q s-1 } is used to obtain multiple coefficient polynomials in basis O and store them in GPU global memory, where the number of moduli in basis I is r and the number of moduli in basis O is s; the input coefficient polynomial is Basis I is Basis Q l ; The base O is the base T; The process of performing modular lifting kernel function calculation is as follows: Start the third kernel function, allocate a three-dimensional GPU block of dimension {N / block_size,s,k}, allocate block_size threads in each block, and read block_size columns of data. The x dimension of the block reads data with polynomial N offset addressing, the y dimension reads data with polynomial modulus offset addressing, and the z dimension reads data with polynomial number offset addressing. k r*N dimensional residual number system ciphertext matrices are read from the global memory into the register file in batches, and the multiplication between the ciphertext and the pre-calculated modulus constant vector is calculated to obtain an r*N dimensional matrix, and stored in the register file; the r*N dimensional matrix is read from the register file, and floating-point division is performed on it and the corresponding modulus and the results are summed to obtain a 1*N dimensional integer vector; the 1*N dimensional integer vector is read from the register file, and it is combined with the input basis vector in the global memory. Multiply to obtain the first s*N dimensional matrix and store the result in the register file; read the first s*N dimensional matrix from the register file and multiply it with the basis transformation matrix in the global memory to obtain the second s*N dimensional matrix; read the first s*N dimensional matrix and the second s*N dimensional matrix from the register file, perform subtraction, and then reduce the output to below the basis O to obtain the third s*N dimensional matrix and write it back to the global memory.
7. The key switching method based on a graphics processor as claimed in claim 5, characterized in that: The S6 includes: inputting k coefficient polynomials [a0] of basis I in the GPU global memory I ,[a1] I ,…,[a k-1 ] I Perform modular lifting kernel function to calculate the basis I = {p0, ..., p r-1 } in polynomial data[a0] I ,[a1] I ,…,[a k-1 ] I In the basis O={q0,…,q s-1 } is used to obtain multiple coefficient polynomials in basis O and store them in GPU global memory, where the number of moduli in basis I is r and the number of moduli in basis O is s; the input coefficient polynomial is s a and b ; Base I is base T; Base O is base PQ l ; The process of performing modular lifting kernel function calculation is as follows: Start the third kernel function, allocate a three-dimensional GPU block of dimension {N / block_size,s,k}, allocate block_size threads in each block, and read block_size columns of data. The x dimension of the block reads data with polynomial N offset addressing, the y dimension reads data with polynomial modulus offset addressing, and the z dimension reads data with polynomial number offset addressing. k r*N dimensional residual number system ciphertext matrices are read from the global memory into the register file in batches, and the multiplication between the ciphertext and the pre-calculated modulus constant vector is calculated to obtain an r*N dimensional matrix, and stored in the register file; the r*N dimensional matrix is read from the register file, and floating-point division is performed on it and the corresponding modulus is performed and the results are summed to obtain a 1*N dimensional integer vector; the 1*N dimensional integer vector is read from the register file and is added to the input basis vector in the global memory. Multiply to obtain the first s*N dimensional matrix and store the result in the register file; read the first s*N dimensional matrix from the register file and multiply it with the basis transformation matrix in the global memory to obtain the second s*N dimensional matrix; read the first s*N dimensional matrix and the second s*N dimensional matrix from the register file, perform subtraction, and then reduce the output to below the basis O to obtain the third s*N dimensional matrix and write it back to the global memory.
8. A key switching system, comprising a central processing unit and a graphics processing unit, wherein the memory stores a computer program, It is characterized by: The central processing unit is used to schedule the data flow between the central processing unit and the graphics processing unit, and allocate the computing control flow to the graphics processing unit; The graphics processor is configured to implement the steps of any one of the methods of claims 1 to 7 when executing the computer program.