Method and device for executing homomorphic encryption algorithm on GPU

By integrating the vector transformation and element-level calculation of the homomorphic encryption algorithm on the GPU, the problem of large calculation overhead of the total homomorphic encryption algorithm is solved, and the computing efficiency and resource utilization are improved.

CN120335993APending Publication Date: 2025-07-18ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389307.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The calculation overhead of the fully homomorphic encryption algorithm is huge, 5 to 6 orders of magnitude higher than plaintext calculations. How to reduce the calculation overhead has become an important technical issue.

Method used

When executing homomorphic encryption algorithm on the GPU, the vector transformation is decomposed into two sets of subvector transformations through kernel splitting, and the kernel fusion is performed according to predetermined fusion rules, including fusing element-level calculations with a nearby set of subvector transformations to a GPU core, and the second set of subvector transformations of two adjacent vector transformations and the first set of subvector transformations to a GPU core to reduce memory access.

Benefits of technology

Through kernel fusion, the utilization rate of computing resources is improved, memory access is reduced, the computing overhead of homomorphic encryption is reduced, and computing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335993A_ABST
    Figure CN120335993A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and a device for executing a homomorphic encryption algorithm on a GPU (Graphics Processing Unit), in the process of executing the homomorphic encryption algorithm on the GPU, vector transformation can be decomposed into two groups of multi-point sub-vector transformation through kernel splitting, and kernel fusion is carried out according to a preset fusion rule so as to start corresponding GPU kernel execution. Wherein the preset fusion rule can comprise the following steps: fusing element-level calculation in a homomorphic encryption algorithm and an adjacent group of sub-vector transformation into a GPU kernel; in two adjacent times of vector transformation, the second group of sub-vector transformation depending on the forward vector transformation and the first group of sub-vector transformation depending on the backward vector transformation are fused into one GPU kernel, so that memory access is reduced, and data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and apparatus for executing a homomorphic encryption algorithm on a GPU. Background Art

[0002] With the development of artificial intelligence, it has been widely applied in various fields, such as intelligent customer service, information consultation, navigation, shopping platforms, and so on. The application of artificial intelligence is accompanied by more data security problems. For example, in the scenario of multi-party secure computing, multiple participating parties each use local data and jointly perform model training, business processing, etc. Each participating party usually does not want to disclose local data to other participating parties. For another example, in various sensitive fields such as medical, finance, and personal assistants, users can provide local data (such as symptom description information, questions about weather inquiries, etc.) to the model, and the service provider provides corresponding business processing results based on the artificial intelligence model. The user's local data may not be willing to be disclosed to the service provider as private information.

[0003] To protect data privacy, various encryption algorithms can usually be adopted in these scenarios, such as secret sharing in multi-party secure computing, fully homomorphic encryption (FHE), etc. Among them, FHE is a feasible solution applicable to the above scenarios. It allows operations to be directly performed on ciphertext without decryption. Specifically, any number of addition and multiplication operations can be performed in the encrypted state, and the same result as the encryption of the plaintext result can be obtained. This means that complex calculations can be performed on encrypted data without decrypting the data. For example, in the sensitive fields mentioned above, the user can first encrypt the data and send it to the service provider. After the service provider calculates the encrypted data, the result is returned to the user in encrypted form, and the user then decrypts the result using their own private key. This process ensures that the model and data are not leaked between the two service parties.

[0004] However, although fully homomorphic encryption has remarkable privacy protection effects, it has a complex and intensive computational form and a huge computational overhead. Compared with plaintext computing, the computational amount may even be 5 to 6 orders of magnitude higher (10 5 to 10 6 ). Therefore, how to reduce the computational overhead of fully homomorphic encryption is an important technical problem that needs to be solved during the execution of fully homomorphic encryption. Summary of the Invention

[0005] One or more embodiments of this specification describe a method and apparatus for executing a homomorphic encryption algorithm on a GPU to solve one or more problems mentioned in the background art.

[0006] According to a first aspect, a method for performing a homomorphic encryption algorithm on a GPU is provided. The current algorithms for homomorphic encryption include at least two vector transformations and element-wise computations. The method includes: decomposing each vector transformation into two sets of sub-vector transformations respectively through kernel splitting; performing kernel fusion according to a predetermined fusion rule to start the execution of the current algorithm by corresponding GPU kernels. The predetermined fusion rule includes: fusing the element-wise computation in the current algorithm with an adjacent set of sub-vector transformations into one GPU kernel; in two adjacent vector transformations, fusing the second set of sub-vector transformations of the earlier vector transformation with the first set of sub-vector transformations of the later vector transformation into one GPU kernel.

[0007] In one embodiment, the fusion rule further includes: fusing computations with logical parallelism into one GPU kernel.

[0008] In one embodiment, the performing kernel fusion according to a predetermined fusion rule includes: constructing a computation graph according to the computing units in the current algorithm. In the computation graph, each node corresponds to each computing unit respectively. A single computing unit corresponds to a basic GPU kernel function. A single set of sub-vector transformations decomposed from a single vector transformation corresponds to a single computing unit. Edges with corresponding connection attributes are connected between the nodes; traversing the computation graph, using the connection edges with the attribute of decomposition edges as splitting points to split the computation graph. The decomposition edges are the edges connecting the nodes corresponding to the two sets of sub-vector transformations decomposed from a single vector transformation; creating GPU kernels according to the splitting results to perform corresponding operations.

[0009] In a further embodiment, the creating GPU kernels according to the splitting results to perform corresponding operations includes: constructing respective execution graphs for each part obtained by splitting; for each execution graph, sequentially performing corresponding operations through each GPU kernel.

[0010] In one embodiment, the vector transformation is one of the following: Discrete Fourier Transform (DFT), Fast Fourier Transform (FFT), Number Theoretic Transform (NTT), Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Discrete Wavelet Transform (DWT), or their inverse transforms.

[0011] In one embodiment, the decomposing each vector transformation into two sets of sub-vector transformations respectively through kernel splitting adopts a four-step decomposition method. The two sets of sub-vector transformations obtained by decomposing an N-point vector transformation are respectively: N1 N2-point vector transformations and N2 N1-point vector transformations, where N = N1 × N2, and both N1 and N2 are powers of 2. The matrix transpose and Hadamard product calculation steps in the decomposition process are respectively fused into any one of the adjacent sets of sub-vector transformations.

[0012] According to a second aspect, there is provided an apparatus for performing a homomorphic encryption algorithm on a GPU. The current algorithms for homomorphic encryption include at least two vector transformations and element-wise computations. The apparatus includes:

[0013] A decomposition unit configured to decompose each vector transformation into two sets of sub-vector transformations respectively through kernel splitting;

[0014] A fusion unit configured to perform kernel fusion according to a predetermined fusion rule to start a corresponding GPU kernel to execute the current algorithm. The predetermined fusion rule includes: fusing the element-wise computation in the current algorithm with a neighboring set of sub-vector transformations into a single GPU kernel; in two adjacent vector transformations, fusing the second set of sub-vector transformations of the earlier vector transformation with the first set of sub-vector transformations of the later vector transformation into a single GPU kernel.

[0015] In one embodiment, the fusion unit is further configured to: construct a computation graph according to the computing units in the current algorithm. In the computation graph, each node corresponds to each computing unit respectively, a single computing unit corresponds to a single GPU kernel function, a single set of sub-vector transformations decomposed from a single vector transformation corresponds to a single computing unit, and the nodes are connected by edges with corresponding connection attributes; traverse the computation graph, and use the connection edges with the attribute of decomposition edges as split points to split the computation graph. The decomposition edges are the edges connecting the nodes corresponding to the two sets of sub-vector transformations decomposed from a single vector transformation; create GPU kernels according to the splitting results to perform corresponding operations.

[0016] According to a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method of the first aspect.

[0017] According to a fourth aspect, there is provided a computing device, including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0018] Through the methods and devices provided in the embodiments of this specification, in view of the computational characteristics of the transformations in different domains in the homomorphic encryption algorithm, the primitives therein can be subjected to kernel fusion during the computational process. Generally, an FHE primitive may include: at least two vector transformations, and element-wise computations, where the two vector transformations are inverse to each other. Under the technical concept of this specification, the vector transformation can be decomposed into two sets of multi-point sub-vector transformations through kernel splitting. In this way, kernel fusion can be performed according to a predetermined fusion rule to initiate the execution of the corresponding GPU kernel. The predetermined fusion rule may include: fusing the element-wise computations in the homomorphic encryption algorithm with a neighboring set of sub-vector transformations into one GPU kernel; in two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation and the first set of sub-vector transformations of the later vector transformation are fused into one GPU kernel, thereby storing data through the shared memory of the kernel and reducing memory access.

[0019] The above-described fusion is usually the fusion between computations with a sequential order, which can be denoted as vertical fusion (or longitudinal fusion). In some embodiments, computations with logical parallelism can also be fused into one GPU kernel. For example, decomposing large numbers in each transformation into small numbers, and the parallel computations between the small numbers can be horizontally fused (which can also be denoted as horizontal fusion), such that similar operations at the same level are performed in one kernel of the GPU, improving the utilization rate of computing resources and the operation efficiency at the same time, thereby reducing the computational overhead of homomorphic encryption. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 Schematic diagram showing a specific example of a 32-point NTT transformation;

[0022] Figure 2 Schematic diagram showing the two-dimensional decomposition architecture of the NTT vector transformation;

[0023] Figure 3 Schematic diagram showing the element-wise operations in the GPU;

[0024] Figure 4 Schematic diagram showing the process of executing the homomorphic encryption algorithm on the GPU according to an embodiment of this specification;

[0025] Figure 5 Schematic diagram showing the principle of plaintext-ciphertext multiplication in the homomorphic encryption algorithm according to a specific example;

[0026] Figure 6 is a schematic diagram of a fusion architecture that performs kernel fusion according to Figure 5 the plaintext-ciphertext multiplication shown;

[0027] Figure 7 is a schematic diagram of an architecture that automatically performs kernel fusion through a computational graph according to an embodiment of the specification;

[0028] Figure 8 shows a schematic diagram of an execution graph constructed by splitting a computational graph according to a specific example;

[0029] Figure 9 is a block diagram of the structure of a device for performing a homomorphic encryption algorithm on a GPU according to an embodiment of the present specification. Detailed implementation manners

[0030] The following describes the solution provided in the present specification in conjunction with the accompanying drawings.

[0031] First, clarify several professional terms that may be involved in the present specification:

[0032] GPU, (Graphic Processing Unit), is translated into Chinese as "graphics processing unit". The GPU is usually the "brain" of a graphics card, determining the grade and most of the performance of the graphics card. The GPU uses a large number of computing units (such as SMs) and a very long pipeline, often having a simpler control logic and omitting the Cache. It usually faces a large-scale data environment with a highly unified type, no mutual dependence, and a pure computing environment that does not need to be interrupted. Both the CPU and the GPU can be used for layout computing. In application scenarios of large-scale basic computing (such as model training, addition, multiplication, modulo operation, etc.), the GPU usually has higher processing performance.

[0033] SM (Streaming Multiprocessor) is a very important concept in the GPU architecture. It is the basic unit for performing parallel computing in the GPU. A single SM can contain multiple SPs (Streaming Processors), also known as CUDA Cores. These cores are responsible for executing specific computing tasks, such as addition, multiplication, modulo calculation, etc. There are many CUDA Cores in the GPU, and the parallel computing power of the GPU can be utilized by these CUDA Cores. The core components of the SM include CUDA Cores, shared memory, registers, etc. The SM can concurrently execute hundreds of threads. The concurrency ability depends on the number of resources owned by the SM. The basic execution unit is the warp, which contains multiple (e.g., 32) threads. These threads execute the same instruction simultaneously. A single thread contains its own instruction address counter and register state, and also has its own independent execution path.

[0034] Homomorphic operation primitives: The application of FHE is based on homomorphic operation primitives, including addition, multiplication, rotation, etc. These operations will be further decomposed into core GPU kernel functions, such as Number Theoretic Transform (NTT), element-wise addition, multiplication, and modulo operations, etc. During the process of processing business data, starting a GPU kernel usually means starting a GPU kernel function. A GPU kernel function usually can call one or more SMs to execute.

[0035] A specific example of a kernel function call is as follows: Specify the kernel function task to the current GPU device, that is, allocate the Grid to a Device; according to the specified first parameter, tell the Giga Thread Engine how many Blocks to schedule. The Giga Thread Engine allocates each Block to each SM. A Block can only occupy one SM, and one SM can run multiple Blocks simultaneously; when an SM receives a Block task, it will determine how many Threads (corresponding to CUDA Cores) the Warp Scheduler needs to schedule according to the second parameter passed by the kernel function.

[0036] NTT (Number Theoretic Transform): NTT is a similar form of the Fast Fourier Transform (FFT) in the field of number theory, which is calculated in a finite field modulo a prime number. The number theoretic transform is a variant of the discrete Fourier transform that uses modular arithmetic instead of traditional real or complex arithmetic. NTT is particularly suitable for occasions that require modular operations, such as large integer multiplication and homomorphic encryption in cryptography. NTT can be used for polynomial multiplication, and its process is similar to that of the FFT (Fast Fourier Transform), the difference being that the operation object of the FFT is the complex number field, while the operation object of the NTT is the prime number field. NTT can also reduce the complexity of polynomial multiplication from O(N^2) to O(N log N), thus providing a significant acceleration in fields such as cryptography. The NTT algorithm has very wide applications in fields such as lattice-based post-quantum cryptography, fully homomorphic encryption, and zero-knowledge proofs, and is one of the important algorithms in the field of modern cryptography.

[0037] Butterfly Computation: The core operation in NTT, which can also be called butterfly transformation. It performs weighted and multiplication operations on two input data respectively, and then outputs the operation results to two different positions. Usually, different NTT algorithms can be implemented through different combinations of butterfly operations, such as radix-2 butterfly, radix-4 butterfly, etc. Usually, the butterfly operation in the NTT algorithm is a recursive process. The original N data points can be divided into two DFT (Discrete Fourier Transform) calculations of N / 2 data points. After each performs intermediate operations, they are merged into a DFT of length N. Different levels of recursive iterations can be integrated through butterfly operations.

[0038] In scenarios involving business processes such as secure computing (such as homomorphic encryption) and image processing (such as the conversion from the spatial domain to the frequency domain), data processing processes usually involve vector transformations (such as FFT or NTT, etc.). The dimension of the vector targeted by the vector transformation, that is, the number of data points corresponding to the vector transformation. For example, homomorphic encryption in secure computing usually involves polynomial operations. In polynomial multiplication operations, the coefficients of the polynomial form an N-dimensional vector, and it is necessary to transform the N-dimensional vector into an N-dimensional vector in another representation domain (such as a finite field), that is, to calculate the N-point NTT transform to accelerate polynomial processing speed. Image processing usually processes feature maps, and feature maps are usually in matrix form. Since FFT or NTT, etc. are usually vector transformations, for the matrix corresponding to the feature map, it can be regarded as multiple groups (such as one row as a group or one column as a group) of vector transformations, or expressed as a tensor in vector form by flattening and other methods and then transformed. In this way, for vectors or matrices in actual business, FFT or NTT, etc. can be performed.

[0039] As Figure 1 shown, a schematic diagram of a 32-point NTT transform is given. In the figure Denotes the rotation factor in the NTT transform, where N = 16, n = 0, 1, 2... 15. Figure 1 Is only an example of the number of points specification for an NTT calculation. The number of butterfly transforms (corresponding to Figure 1 the two adjacent columns in Figure 1 is different for NTT transforms with different numbers of points. The transform in Figure 1 can be understood as containing 16 two-point butterfly transforms, and the NTT transform of this 32-point vector can contain fifth-order butterfly transforms. In the actual calculation structure, calculation instructions for NTTs with different numbers of points specifications can be supported, and the number of butterfly transforms therein can be much larger than

[0040] the 5 shown in

[0041] Figure 2 . Those skilled in the art can understand that the number of points of the NTT transform can be the specification of the NTT transform and also the dimension of the vector targeted by the transform. In homomorphic encryption calculations, it also represents the number of polynomial terms or the number of coefficients contained in a single polynomial.

[0040] In actual business, in addition to the number-theoretic transform NTT and its inverse transform, there may also be various transformation forms of data in different domains, such as the discrete Fourier transform DFT, the fast Fourier transform FFT, the discrete cosine transform DCT, the discrete sine transform DST, the discrete wavelet transform DWT, and their inverse transforms. In this specification, the transforms involved are all exemplified by the NTT transform or the FFT transform, but the possibility of them being other similar transforms is not excluded.

[0041] Figure 2 Shows the decomposition schematic of the FFT transform. The technical concept of this specification can be based on the four-step split (4-step algorithm, i.e., two-dimensional decomposition) of tensor transforms such as FFT and NTT. The principle of the four-step split is briefly described below.

[0042] Assume that a tensor x(j) with the number of elements n is transformed into another tensor y(s) with the number of elements n, and the rotation factor (or denoted as the transformation matrix) is Then the transform can be denoted as: Among them, the transformation matrix is related to the tensor size n, such as ω n = e -2πi / n , 0 ≤ s < n. In the case where n can be decomposed into the product of two factors n1 and n2, n = n1 × n2, then j and s can be expressed as j = j1 + j2n1, s = s2 + s1n2. And x and y are expressed in two-dimensional form: x(j) = x(j1, j2), y(s) = y(s2, s1), 0 ≤ j1 < n1 - 1, 0 ≤ j2 < n2 - 1, 0 ≤ s1 < n1 - 1, 0 ≤ s2 < n2 - 1. Usually, j1 and j2 are as close as possible, s1 and s2 are as close as possible, and n1 and n2 are as close as possible. Thus, the above tensor transformation process can be expressed as:

[0043]

[0044] Under the four-step splitting method, there are:

[0045] In the first step, perform a tensor transformation of n2 points in n1 groups, such as

[0046] In the second step, perform a matrix transpose, such as x2(s2, j1) = x1(j1, s2);

[0047] In the third step, calculate the Hadamard product, such as

[0048] In the fourth step, perform a tensor transformation of n1 points in n2 groups, such as

[0049] According to the above principle, it can be known that when n, n1, and n2 are determined, each transformation matrix as a rotation factor can be determined, such as ω n = e -2πi / n , can all be determined. Usually, the rotation factor can be related to at least one of the rotation type (such as NTT type, FFT type, etc.), transformation parameters (such as e, i, etc.), and the number of elements (such as n).

[0050] According to the above four-step splitting principle, as Figure 2 shown, for the vector transformation process of N points, N can be decomposed into the product of N1 and N2, so that the vector transformation of N points is decomposed into a group of vector transformations of N2 points in N1 groups and a group of vector transformations of N1 points in N2 groups, that is, a two-dimensional decomposition of the N-point vector transformation. Among them, N1 × N2 = N, and N, N1, and N2 are usually powers of 2. In an alternative embodiment, N1 and N2 can be made as close as possible to improve processing efficiency. Specifically, when N is an even power of 2, N1 can be equal to N2. For example, when N = 4096, N1 = N2 = 64 can be taken. When N is an odd power of 2, one of N1 and N2 is the square root of 2N, and the other is the square root of N / 2. For example, when N = 8192, N1 = 128 and N2 = 64 can be taken, and so on. It can be understood that when N is not a power of 2, methods such as padding with 0 can be used to take the smallest power of 2 greater than N.

[0051] Continue to refer to Figure 2As shown, the N-point vector transformation can be decomposed two-dimensionally. Specifically, under a four-step splitting architecture, in the first step, N2 N1-point vector transformations can be performed to obtain the first intermediate matrix of size N1×N2. In the second step, the first intermediate matrix is transposed to obtain the second intermediate matrix of size N2×N1. In the third step, the Hadamard product of the second intermediate matrix and the corresponding transformation matrix can be performed to obtain the third intermediate matrix of dimension N2×N1. In the fourth step, N1 N2-point vector transformations are performed based on the third intermediate matrix to obtain the transformation result matrix of dimension N2×N1, i.e., the target tensor. Among them, both the matrix transpose in the second step and the Hadamard product calculation in the third step are element-wise operations. The matrix transpose in the second step can be postponed to the third or fourth step. At this time, the other steps are moved forward in turn, and the dimensions of the obtained intermediate matrices are determined according to the actual operation sequence. If the matrix transpose operation is performed in a step after the Hadamard product, the corresponding transformation matrix also needs to be transposed accordingly. In this specification, the case where the matrix transpose is in the second step is described as an example. In practice, the situation where the matrix transpose is swapped to other step orders is essentially the same as the example in the second step above.

[0052] Element-wise operation: An element-wise operation is a basic operation that can operate on each element of a set (such as an array or a vector) individually. Common element-wise operations include, for example, addition, subtraction, multiplication, and division. These operations can be used for one-dimensional arrays, two-dimensional arrays, or arrays of higher dimensions. For example, for two one-dimensional arrays a and b, they can be added using addition. Figure 3 An example of an element-wise operation is shown. As Figure 3 shown, in the element-wise calculation of two large numbers (such as vector0 and vector1 corresponding to Figure 3 ), a single large number is divided into 4 smaller numbers. For example, a 20-byte (bit) large number is divided into 4 5-byte smaller numbers. A single smaller number can be processed by a single SM. When both large numbers are divided into 4 smaller numbers, the element-wise operation of the two smaller numbers corresponding to the same position (the 0th to 3rd elements in vector0 and vector1) is performed by a single SM, and the element-wise operation results of the 0th to 3rd elements of the result vector are output. Both the matrix transpose and the Hadamard product calculation in the four-step decomposition above are element-wise operations.

[0053] Conventional business processing (such as homomorphic encryption) processes involving vector transformations (such as NTT) are mostly focused on optimizing the parallelization of vector transformations and complex operations such as inner products in order to save computing overhead. However, in practical applications, simple operations such as element-level addition and modulo are becoming increasingly important in the total execution time. Preliminary experiments show that during bootstrapping (a statistical estimation method based on resampling), the kernel execution time of element-level calculations can account for up to 53.1%. What's more serious is that such simple kernel calculations are usually subject to latency and usually lack sufficient instruction-level parallelism (ILP) to eliminate latency.

[0054] In order to further improve computing efficiency and memory utilization, this specification provides a fully homomorphic encryption computing architecture based on kernel fusion. Under this computing architecture, for elements and operations that rely on vector transformation, on the basis of the decomposition of the vector transformation process, vertical kernel fusion can be performed on the steps and element-level operations of adjacent vector transformations, thereby improving computing core utilization and reducing memory access. Among them, vertical fusion of kernels refers to the calculation process that is performed sequentially, merging multiple basic calculations and starting the same GPU kernel function for calculation. In this way, since the shared memory in the computing core can store the calculation result data, the access to the device memory is reduced, and the data processing efficiency is improved.

[0055] The technical concept of this specification is described in detail below with reference to the accompanying drawings.

[0056] Figure 4 The process of executing a homomorphic encryption algorithm on a GPU in one embodiment of the present specification is shown. The execution subject of the process can be various computers, devices, servers, etc. integrated with a GPU. It can be understood that due to the characteristics of the computer processing process, in actual business, various business data processes can be divided into many algorithms to complete. A single algorithm can be a calculation process that completes a certain function in a specific business processing process. The business processing process here can be a data processing process performed in an actual application to achieve a certain business goal, such as the process of multiple participants jointly conducting model training in the multi-party secure computing scenario described above, the process of a service provider processing user data based on an artificial intelligence model to obtain business processing results provided to users, and so on. An algorithm may, for example, include at least one operation of addition, multiplication, modulo, etc. In the case of business processing based on homomorphic encryption, various algorithms are based on homomorphic encryption, such as modulo in ciphertext state, ciphertext multiplication, etc. In the homomorphic encryption form, the computational complexity of the algorithm increases.

[0057] As a specific example, Figure 5 The figure shows the calculation principle of a plaintext and ciphertext multiplication in the homomorphic encryption algorithm. Figure 5Among them, In represents a large ciphertext number, which is divided into r + 1 parts (where r is the number of layers of multiplication, and during execution, it can correspond to r SMs as shown in Figure 3 ), and are respectively denoted as In0, In1, In2... Inr. Pt is a large plaintext number, which is also divided into r + 1 parts, and are respectively denoted as Pt0, Pt1, Pt2... Ptr. In the case of performing the multiplication calculation of the plaintext and ciphertext of In and Pt, the respective parts of In and Pt can be respectively corresponding to perform the multiplication calculation, and the multiplications can be respectively denoted as Mul0, Mul1, Mul2... Mulr, and r calculation results are obtained, for example, respectively denoted as In'0, In'1, In'2... In'r.

[0058] To reduce the calculation error, it is usually also necessary to perform rescaling on In'0, In'1, In'2... In'r. During the rescaling process, for the last part In'r, first perform the INTT (inverse operation of NTT) vector transformation, and then the results are respectively processed by r paths for rescaling. On a single path, first perform modulo (mod), then perform the NTT vector transformation operation, and then perform scaling (Scale). The scaling result is fused (sub operation) with the corresponding calculation result (such as one of In'0, In'1 to In'(r - 1)), to obtain a part of the output (out). Thus, the output results OUT0, OUT1, OUT2... OUT(r - 1) of each path constitute the output result of the entire plaintext-ciphertext multiplication algorithm.

[0059] In a business processing process, it may include one or more algorithms such as Figure 5 shown. The current algorithm can be any one of them.

[0060] As described above, during the homomorphic encryption calculation process, it usually involves the conversion of different data domains, that is, vector transformation. The data is mapped to the finite field for calculation, and then converted back to the original data domain to speed up the processing speed. In the current algorithm, it can include at least two vector transformations, and related element-level operations. Based on this, as Figure 4 shown, the process of executing the homomorphic encryption algorithm on the GPU can include the following steps: Step 401, decompose each vector transformation into two groups of sub-vector transformations respectively through kernel splitting; Step 402, perform kernel fusion according to the predetermined fusion rules to start the corresponding GPU kernel to execute the current algorithm. The predetermined fusion rules include: fusing the element-level calculations in the current algorithm with a group of adjacent sub-vector transformations into a GPU kernel; in two adjacent vector transformations, the second group of sub-vector transformations of the earlier vector transformation and the second group of sub-vector transformations of the later vector transformation are fused into a GPU kernel.

[0061] First, in step 401, each vector transformation is decomposed into two sets of sub-vector transformations through kernel splitting.

[0062] Among them, kernel splitting is usually the functional segmentation of the algorithms used in the business processing. For vector transformation, functional segmentation can be performed through two-dimensional decomposition. The two-dimensional decomposition can refer to Figure 2 the four-step decomposition method shown. By decomposing the N-point vector, it is converted into two sets of multi-point vector transformations with a smaller magnitude, which are denoted as two sets of sub-vector transformations here. In this way, the parallelism can be increased. For example, N2 vector transformations of N1 points can be parallel, and N1 vector transformations of N2 points can be parallel. For any one of the vector transformation and its inverse transformation (which is also a vector transformation) in the current algorithm, it can be decomposed to obtain the corresponding two sets of sub-vector transformations.

[0063] Referring to Figure 5 shown, the INTT of In'r can be decomposed into two sets of sub-vector transformations similar to Figure 2 through a four-step decomposition method, which can be denoted as (INTTr)1 and (INTTr)2 respectively. Then, taking the path corresponding to label "0" as an example, for the result of (INTTr)2, first take the modulus with respect to the predetermined prime number of the finite field corresponding to this path, denoted as Mod0, and then perform the NTT0 calculation. The NTT0 is also decomposed into two parts, denoted as (NTT0)1 and (NTT0)2 respectively. After scaling the calculation result of (NTT0)2 through Scale0, the scaled result is fused with the multiplication result In'0 of the first part, and the first part OUT0 of the plaintext-ciphertext multiplication result can be obtained. Similarly, on the other r - 1 paths, there are corresponding vector transformations NTT1, NTT2... NTT(r - 1) respectively, and each of them is also decomposed into two sub-vector transformations, such as Figure 5 (INTT1)1, (INTT1)2, (NTT0)1, (NTT0)2, etc. shown in

[0064] Figure 5 The vector transformations shown are NTT vector transformation and its inverse INTT vector transformation. In practice, the vector transformation can also be one of the following: discrete Fourier transform DFT, fast Fourier transform FFT, discrete cosine transform DCT, discrete sine transform DST, discrete wavelet transform DWT, or their inverses.

[0065] Then, via step 402, kernel fusion is performed according to the predetermined fusion rule to start the corresponding GPU kernel to execute the current algorithm.

[0066] Kernel fusion is generally a technique of merging multiple kernel functions into one kernel function. For example, for two vectors, when performing a dot product, the elements at the corresponding positions need to be multiplied once, and then a summation operation is performed. Here, both the dot product and the summation are element-wise operations. The usual approach is to write two kernels to perform one elementwise_mul (element-wise multiplication) and one reduce_sum (reduction summation) operation. Kernel fusion can merge the two kernels. For example, when fetching numbers for reduce_sum, elementwise_mul can be performed incidentally.

[0067] To perform kernel fusion, the fusion rules can be determined in advance, that is, the predefined fusion rules. For example, it can include: fusing the element-wise calculations in the current algorithm with a neighboring set of sub-vector transformations into one GPU kernel; in two adjacent vector transformations, fusing the second set of sub-vector transformations of the earlier vector transformation with the first set of sub-vector transformations of the later vector transformation into one GPU kernel; and so on.

[0068] To clarify the above kernel fusion rules, refer to Figure 5 for examples. It can be understood that in the decomposition process of the vector transformations (such as NTT and INTT) shown in Figure 5 the element-wise calculations of matrix transpose and Hadamard product can be incorporated into the first part or the second part of the vector transformation, such as Figure 5 shown in (INTTr)1, (INTTr)2, (NTT0)1, (NTT0)2, etc.

[0069] Under one fusion rule, the element-wise calculations in the current algorithm can be fused with a neighboring set of sub-vector transformations into one GPU kernel. In other words, the element-wise calculations fused with a single set of sub-vector transformations into one GPU kernel do not go through other vector transformation operations with respect to this single set of sub-vector transformations. For example Figure 5 the multiplication calculation Mulr shown in is fused into the first set of sub-vector transformations (INTTr)1 of the vector transformation INTTr, and the scaling operation Scale0 is fused into the second set of sub-vector transformations (NTT0)2 of the vector transformation NTT0, and so on.

[0070] Under another fusion rule, in two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation is fused with the first set of sub-vector transformations of the later vector transformation into one GPU kernel. As shown in Figure 5Among them, the first set of sub-vector transforms (NTT0)1 of the vector transform NTT0, the first set of sub-vector transforms (NTT1)1 of the vector transform NTT1, etc., can be fused with the second set of sub-vector transforms (INTTr)2 of the vector transform INTTr into a single GPU kernel, and the intermediate element-wise operations can be automatically fused with them into a single GPU kernel.

[0071] The above are all fusions of sequential-relationship computations, that is, vertical fusion. In the case of vertical fusion, since multiple sequentially executed computing units are processed by the same GPU kernel, the shared memory of the GPU kernel can be fully utilized to store intermediate data, thereby reducing the number of device memory accesses and improving data processing efficiency.

[0072] In some algorithms, there may also be computations with logical parallelism, such as Figure 5 the r resizing paths in. In an alternative embodiment, computations with logical parallelism can also be fused into a single GPU kernel, that is, horizontal fusion. Refer to Figure 5 As shown, the computations on the r resizing paths are computations with logical parallelism. Horizontal fusion can fuse paths with inherent logical parallel relationships, allowing threads within a thread block to cooperate to load data from device memory into shared memory in a merged manner, optimizing the memory access pattern and improving data processing efficiency.

[0073] For ease of description, a computing unit can be used to represent the functional unit that can typically be completed by a basic GPU kernel function in the current algorithm. The basic GPU kernel function is a set of basic computing functions that can be completed by the same computing core, such as multiplication (Mul) computations, a set of vector transforms (such as), modulo operations, etc. Correspondingly, a single multiplication computation can correspond to a computing unit, a set of sub-vector transforms can correspond to a computing unit (such as (INTTr)1, (NTT0)1, etc.), a single modulo operation can correspond to a computing unit, a scaling (Scale) computing unit, etc. In Figure 5 a single circle (including the circle with an '×' drawn inside) corresponds to a single computing unit. Then the kernel fusion in this step 402 can be understood as the fusion of computing units.

[0074] Referring to the above fusion rules, Figure 6 shows a schematic of the fusion result of kernel fusion for the Figure 5 plaintext-ciphertext multiplication. As Figure 6 shown, the cutting lines 601 and 602 are respectively cut between two sets of sub-computing units obtained from the four-step decomposition of INTT and NTT. The single parts cut out by the cutting lines can be placed in the same GPU kernel for execution, thereby achieving kernel fusion.

[0075] Kernel fusion has great potential. However, manual fusion has limitations. Due to the complexity of FHE primitives and numerous fusion possibilities, it is easy to overlook potential fusion opportunities in manual fusion. In addition, for each new FHE scheme or even variant, the manual fusion process needs to be repeated. Therefore, in a possible design, kernel fusion is considered through an automatic fusion method.

[0076] Figure 7 The flowchart of kernel fusion according to a predetermined fusion rule in an embodiment of this specification is shown. As Figure 7 shown, the process may include the following steps: Step 701, construct a computation graph according to the computing units in the current algorithm. In the computation graph, each node corresponds to each computing unit one by one. A single computing unit corresponds to a basic GPU kernel function. A single set of sub-vector transforms decomposed from a single vector transform corresponds to a single computing unit, and there are edges with corresponding connection attributes between the nodes; Step 702, traverse the above-mentioned computation graph, and use the connection edge with the attribute of decomposition edge as the splitting point to split the computation graph. The decomposition edge is the edge connecting the nodes corresponding to two sets of sub-vector transforms decomposed from a single vector transform; Step 703, create a GPU kernel according to the splitting result and execute corresponding operations.

[0077] First, in step 701, construct a computation graph according to the computing units in the current algorithm.

[0078] For the current algorithm, a computation graph can be constructed according to the computing units it includes. Specifically, when constructing the computation graph, a single computing unit can be used as a single node, and the sequentially adjacent computing units can be connected by connection edges. Among them, corresponding attributes can be set for the connection edges, which are represented by different attribute identifiers respectively. There are edges with corresponding connection attributes between the nodes. For example, the attributes of the edges can include: decomposition edge (the computing unit connecting the sub-vector transforms decomposed from the vector transform), transfer edge (the two computing units connecting the sequentially executed ones), and so on. The attribute identifiers of the edges can be described in natural language or represented by characters. For example, the decomposition edge is represented as 1, the transfer edge is represented as 2, and so on.

[0079] As a specific example of the computation graph, refer to Figure 6 shown. Figure 6 In Figure 5 Based on the plaintext and ciphertext calculation principle shown, rearrange according to the association relationship between the computing units. The computation graph starts from the multiplication calculation with Inr and Ptr as inputs. After the multiplication is completed, vector transform INTTr is performed through four-step decomposition. The two nodes corresponding to the two computing units of the vector transform INTTr are connected by a decomposition edge. Then, r links of mod, NTT, Scale, and Sub are calculated in parallel to obtain r outputs OUT0 to OUT(r - 1). As Figure 6As shown, the two sub-vector transformations after the decomposition of a single NTT calculation in a single link respectively correspond to two nodes, and the two nodes can be connected by a decomposition edge (represented by a dashed arrow), and other adjacent nodes can be connected by a transfer edge (represented by a solid arrow).

[0080] Figure 6 It is only the computational graph of a single specific algorithm. For other algorithms, a similar method can also be used to construct a computational graph: vector transformations (such as NTT, INTT, etc.) are decomposed into two sets of sub-vector transformations, which respectively correspond to two computational units. The element-level calculations during the decomposition process can be merged into any one of the computational units; a single computational unit corresponds to a single node in the computational graph; the nodes corresponding to the two sets of sub-vector transformations decomposed from a single vector transformation are connected by a decomposition edge, and other nodes are connected by a transfer edge.

[0081] Then, through step 702, traverse the above computational graph, and use the connection edges with the attribute of decomposition edges as the splitting points to split the computational graph.

[0082] Under the technical concept of this specification, the computational graph can be traversed and split based on the attributes of the connection edges between computational units. Among them, each single part after splitting includes several computational units that can be executed by starting a kernel (GPU kernel function).

[0083] Specifically, it can be split at the decomposition edge. Through this splitting, it can be achieved that: the element-level calculations in the current algorithm and a set of adjacent sub-vector transformations are split into one part; among two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation and the first set of sub-vector transformations of the later vector transformation are split into one part, corresponding to the predetermined fusion rule.

[0084] Such as Figure 6 In, the splitting lines 601 and 602 are respectively split between the two computational units obtained from the four-step decomposition of INTT and NTT. Each part cut out by the splitting line is placed in a different part. In this way, the element-level operations are automatically split into the part corresponding to the adjacent sub-vector transformation. This splitting method can divide the second computational unit obtained from the decomposition of the previous vector transformation and the first computational unit obtained from the decomposition of the later vector transformation in two adjacent vector transformations into one part.

[0085] In an alternative embodiment, in order for two adjacent vector transformations to be merged into one kernel for processing, during the decomposition process of the vector transformation, the first computational unit obtained from the decomposition of the later vector transformation (such as Figure 5 , Figure 6 in, (NTT0)1, (NTT1)1, etc.) can be kept the same as the second computational unit obtained from the decomposition of the previous vector transformation (such as Figure 5 , Figure 6such as (INTTr)2 in it have a consistent number of vector transformation points, for example, they are all 128-point vector transformations.

[0086] By detecting the decomposition edges in the computational graph for splitting, vertical partitioning can be achieved, preparing for vertical fusion. According to the previous description, horizontal fusion can also be performed in the case of logically parallel computations. Therefore, in another embodiment, when detecting that the same node is connected to the edges of different paths, the parallel computations on each path can be merged to prepare for horizontal fusion.

[0087] Reference Figure 6 As shown in the figure, among the r paths led out below the computing unit (INTTr)2, each has a consistent data processing mode and performs similar operations, such as modulo operation, NTT calculation, scaling, and fusion. Therefore, horizontal merging between paths can be performed. Since each path is split at the corresponding decomposition edge, the computing units on each path can be merged separately before and after the decomposition edge.

[0088] Next, in step 703, create a GPU kernel according to the splitting result to perform corresponding operations.

[0089] It can be understood that the GPU kernel here is the kernel function of the GPU. For the splitting result obtained by traversing and splitting the computational graph, integration can be performed within a single part to create a GPU kernel to perform corresponding operations. A single part can start a single GPU kernel function. In this way, various element-level operations can be merged.

[0090] In a possible design, when creating a GPU kernel, an execution graph can be created first according to the splitting result, and then for each execution graph, the corresponding GPU kernel function can be created respectively.

[0091] As a specific example, according to Figure 6 the shown splitting lines 601 and 602, each part obtained by splitting can be integrated to obtain an execution graph as shown in Figure 8 the figure. Figure 8 The execution graph in it includes three parts: (1), (2), and (3), and each part corresponds to Figure 6 the three parts split by the splitting lines 601 and 602 in the figure. It can be understood that according to each execution graph, each GPU core can be started in sequence to perform the operations therein. Generally, for a single execution graph, a single GPU kernel can be started to perform corresponding operations.

[0092] Reference Figure 8As shown, in the first execution graph, the multiplication calculation of INr and Ptr is fused with the first set of vector transformations decomposed from INTTr into a GPU kernel calculation. In the second execution graph, the second set of vector transformations decomposed from INTTr is fused with the modulo (MOD) on the parallel r resizing paths and the first set of vector transformations of the NTT decomposition into a GPU kernel calculation. In the third execution graph, the second set of vectors of the NTT decomposition on the r resizing paths is fused with the scaling calculation, the corresponding multiplication layer calculation, and the fusion calculation of the scaling result and the multiplication result into a GPU kernel for calculation. It should be noted that, as described above, element-level calculations such as matrix transpose and Hadamard product in the INTT or NTT decomposition can be fused into the first set or the second set of vector transformations decomposed for calculation. In this way, the number of GPU kernel launches can be determined according to the number of levels of vector transformation, such as adding 1 to the number of levels of vector transformation. Here, the number of levels of vector transformation can also be understood as the number of vector transformations in the path with the most vector transformations involved. This is because the computational graph itself contains certain parallel factors, and different paths are often generated based on parallelism. Therefore, there may be similar operations in parallel paths, and these similar operations can be horizontally fused. For example Figures 5 to 7 in Figures 5 to 7 , the r NTT operations corresponding to the r paths are parallel NTT operations at the same level.

[0093] Reviewing the above process, during the execution of the homomorphic encryption algorithm on the GPU, for the computational characteristics of the homomorphic encryption algorithm that transform to different domains, the primitives therein can be fused during the computational kernel. Generally, an FHE primitive can include: at least two vector transformations, and element-level calculations, and the two vector transformations are inverse transformations of each other. Under the technical concept of this specification, the vector transformation can be decomposed into two sets of multi-point sub-vector transformations through kernel splitting. In this way, kernel fusion can be performed according to a predetermined fusion rule to start the corresponding GPU kernel for execution. The predetermined fusion rule can include: fusing the element-level calculation in the homomorphic encryption algorithm with a neighboring set of sub-vector transformations into a GPU kernel; in two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation is fused with the first set of sub-vector transformations of the later vector transformation into a GPU kernel, so as to store data in the shared memory of the kernel and reduce memory access.

[0094] In a further embodiment, calculations with logical parallelism can also be fused into a GPU kernel. For example, the large numbers in each transformation are decomposed into small numbers, and the parallel calculations between the small numbers can be horizontally fused (which can also be recorded as horizontal fusion), so that similar operations at the same level are performed on one computational core of the GPU, improving the utilization rate of the computational core and the operation efficiency at the same time, thereby reducing the computational overhead of homomorphic encryption.

[0095] According to an embodiment of another aspect, there is also provided an apparatus for performing a homomorphic encryption algorithm on a GPU. Here, the current algorithms for homomorphic encryption include at least two vector transformations and element-wise computations. The apparatus can be incorporated into any computer, device, or server having GPU computing hardware.

[0096] Reference Figure 9 As shown, the apparatus 900 for performing a homomorphic encryption algorithm on a GPU may include:

[0097] A decomposition unit 901 configured to decompose each vector transformation into two sets of sub-vector transformations respectively through kernel splitting;

[0098] A fusion unit 902 configured to perform kernel fusion according to a predetermined fusion rule to initiate the execution of the current algorithm by corresponding GPU kernels. Among them, the predetermined fusion rule may include: fusing the element-wise computation in the current algorithm with a neighboring set of sub-vector transformations into one GPU kernel; in two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation and the first set of sub-vector transformations of the later vector transformation are fused into one GPU kernel.

[0099] In one embodiment, the fusion unit 902 is further configured to: construct a computation graph according to the computing units in the current algorithm; traverse the above computation graph, and use the connection edges with the attribute of decomposition edges as splitting points to split the computation graph; create GPU kernels according to the splitting results to perform corresponding operations. In the computation graph, each node corresponds to each computing unit respectively, a single computing unit corresponds to a basic GPU kernel function, a single set of sub-vector transformations decomposed from a single vector transformation corresponds to a single computing unit, the nodes are connected by edges with corresponding connection attributes, and the decomposition edges are the edges connecting the nodes corresponding to the two sets of sub-vector transformations decomposed from a single vector transformation.

[0100] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to execute the methods described in conjunction with Figure 4 、 Figure 7 and so on.

[0101] According to an embodiment of still another aspect, there is also provided a computing device including a memory and a processor, where the memory stores executable code, and when the processor executes the executable code, the methods described in conjunction with Figure 4 、 Figure 7 and so on are implemented.

[0102] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0103] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above descriptions are only specific implementation manners of the technical concept of this specification and are not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification should be included within the protection scope of the technical concept of this specification.

Claims

1. A method for performing a homomorphic encryption algorithm on a GPU, where the current algorithm of homomorphic encryption includes at least two vector transformations and element-wise calculations; the method includes: Decomposing each vector transformation into two sets of sub-vector transformations through kernel splitting; Performing kernel fusion according to a predefined fusion rule to start a corresponding GPU kernel to execute the current algorithm, and the predefined fusion rule includes: fusing the element-wise calculation in the current algorithm with an adjacent set of sub-vector transformations into one GPU kernel; In two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation and the first set of sub-vector transformations of the later vector transformation are fused into one GPU kernel.

2. The method according to claim 1, wherein The fusion rule further includes: fusing calculations with logical parallelism into one GPU kernel.

3. The method according to claim 1, wherein, The performing kernel fusion according to the predefined fusion rule includes: Constructing a computation graph according to the computing units in the current algorithm. In the computation graph, each node corresponds to each computing unit respectively, a single computing unit corresponds to a basic GPU kernel function, a single set of sub-vector transformations decomposed from a single vector transformation corresponds to a single computing unit, and the nodes are connected by edges with corresponding connection attributes; Traversing the computation graph, using the connection edges with the attribute of decomposition edges as split points to split the computation graph, and the decomposition edges are the edges connecting the nodes corresponding to the two sets of sub-vector transformations decomposed from a single vector transformation; Creating GPU kernels according to the splitting result to perform corresponding operations.

4. The method according to claim 3, wherein, The creating GPU kernels according to the splitting result to perform corresponding operations includes: Constructing respective execution graphs for each part obtained by splitting; For each execution graph, sequentially performing corresponding operations through each GPU kernel.

5. The method according to claim 1, wherein, The vector transformation is one of the following: discrete Fourier transform DFT, fast Fourier transform FFT, number-theoretic transform NTT, discrete cosine transform DCT, discrete sine transform DST, discrete wavelet transform DWT, or their inverse transforms.

6. The method according to claim 1, wherein, The decomposing each vector transformation into two sets of sub-vector transformations through kernel splitting adopts a four-step decomposition method. The two sets of sub-vector transformations decomposed from an N-point vector transformation are respectively: N1 N2-point vector transformations and N2 N1-point vector transformations, where N = N1 × N2, and both N1 and N2 are powers of 2. The matrix transpose and Hadamard product calculations in the decomposition process are respectively fused into any one of the nearest sets of sub-vector transformations.

7. A device for performing a homomorphic encryption algorithm on a GPU, where the current algorithm of homomorphic encryption includes at least two vector transformations and element-wise calculations; the device includes: A decomposition unit configured to decompose each vector transformation into two sets of sub-vector transformations through kernel splitting; A fusion unit configured to perform kernel fusion according to a predefined fusion rule to start a corresponding GPU kernel to execute the current algorithm, and the predefined fusion rule includes: fusing the element-wise calculation in the current algorithm with an adjacent set of sub-vector transformations into one GPU kernel; In two adjacent vector transformations, the second set of sub-vector transformations of the earlier vector transformation and the first set of sub-vector transformations of the later vector transformation are fused into one GPU kernel.

8. The device according to claim 7, wherein, The fusion unit is further configured to: Construct a computational graph according to the computing units in the current algorithm. In the computational graph, each node corresponds to each computing unit respectively. A single computing unit corresponds to a GPU kernel function. A single set of sub-vector transformations decomposed from a single vector transformation corresponds to a single computing unit, and the nodes are connected by edges with corresponding connection attributes. Traverse the computational graph, and use the connection edges with the attribute of decomposition edges as the splitting points to split the computational graph. The decomposition edges are the edges connecting the nodes corresponding to two sets of sub-vector transformations decomposed from a single vector transformation. Create a GPU kernel according to the splitting result to perform corresponding operations.

9. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-6.

10. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory. When the processor executes the executable code, the method according to any one of claims 1-6 is implemented.