Data processing method, system and device, medium and computer program product
By adopting a hybrid method of secure multi-party computing and homomorphic encryption in artificial intelligence model inference, the problem of user privacy information leakage is solved, and efficient privacy protection and model inference efficiency are achieved.
Patent Information
- Application Number
- CN202510526407.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
During the inference process of artificial intelligence model, user privacy information may be leaked, resulting in security risks.
Through a hybrid method of secure multi-party computing and homomorphic encryption, the computing device receives model parameter shards and ciphertexts of user input data, performs calculations and implements a hybrid encryption mechanism to avoid the leakage of model privacy parameters and user data.
It effectively avoids the leakage of private information during model inference, reduces the traffic between computing devices, and improves the efficiency of model inference.
Smart Images

Figure CN120046176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a data processing method, system, device, medium, and computer program product. Background Art
[0002] With the development of artificial intelligence technologies, the applications of artificial intelligence models are becoming more and more extensive, and the number of model parameters is getting larger and larger. For some model inference tasks with high requirements for inference performance and large numbers of model parameters, the models need to be deployed in large-scale computing clusters to meet the performance requirements. In this case, artificial intelligence may collect users' privacy information without the users' knowledge, bringing security risks to the users.
[0003] How to ensure the security of user data during the inference process of artificial intelligence models is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The present invention provides a data processing method, system, device, medium, and computer program product to at least solve the problem that the inference process of artificial intelligence models in related technologies may lead to the leakage of users' privacy information.
[0005] The present invention provides a data processing method applied to a computing device, including: Calculating a ciphertext of a shard of an intermediate calculation result according to a shard of model parameters sent by a first device and a first ciphertext of inference input data sent by a user device; Encrypting the ciphertext of the shard of the intermediate calculation result to obtain a second ciphertext; Sending the second ciphertext to the user device so that the user device aggregates the second ciphertexts of multiple computing devices and decrypts them using a homomorphic encryption algorithm to obtain a third ciphertext; Receiving the third ciphertext and decrypting it to obtain a plaintext of the intermediate calculation result; Performing the remaining calculations of model inference calculation using the plaintext of the intermediate calculation result to obtain a shard of a local output result; Sending the shard of the output result to the user device so that the user device aggregates the shards of the output results of multiple computing devices to obtain a data processing result.
[0006] The present invention further provides a data processing system, including: a plurality of computing devices and a user device; The computing device is configured to calculate the ciphertext of the intermediate calculation result shards based on the model parameter shards sent by the first device and the first ciphertext of the inference input data sent by the user device; encrypt the ciphertext of the intermediate calculation result shards to obtain a second ciphertext; send the second ciphertext to the user device; receive the third ciphertext sent by the user device and decrypt it to obtain the plaintext of the intermediate calculation result; perform the remaining calculations of the model inference calculation using the plaintext of the intermediate calculation result to obtain the local output result shards; and send the output result shards to the user device. The user device is configured to aggregate the second ciphertexts of multiple computing devices and decrypt them using a homomorphic encryption algorithm to obtain the third ciphertext, and is further configured to aggregate the output result shards of multiple computing devices to obtain a data processing result.
[0007] The present invention further provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above data processing methods when executing the computer program.
[0008] The present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above data processing methods.
[0009] The present invention further provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above data processing methods.
[0010] Through the present invention, the computing device receives the model parameter shards sent by the first device through a secure multi-party computing method, receives the first ciphertext of the inference input data sent by the user device through a homomorphic encryption calculation method, and performs calculations based on the model parameter shards and the first ciphertext, implementing a hybrid encryption mechanism based on the computing device, which can effectively avoid the leakage of model privacy parameters or user inference data. For the remaining calculations, the computing device encrypts the ciphertext of the intermediate calculation result shards and sends it to the user device for aggregation and decryption, so that the remaining calculations of the model inference calculation can be performed based on the plaintext of the intermediate calculation result, reducing the computational amount compared to performing the remaining calculations on the ciphertext basis, and avoiding the leakage of model privacy parameters to the user device. Thus, the present invention can not only avoid the leakage of model privacy parameters or user inference data in model inference, but also make full use of computing resources. The user device only needs to perform encryption and decryption calculations and a small amount of simple calculations without performing model calculations. Through the method of mixed plaintext and ciphertext calculation, the communication volume between computing devices is also reduced, improving the model inference efficiency while ensuring data confidentiality. Description of the Drawings
[0011] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It is an architecture diagram of a data processing system provided by an embodiment of the present invention; Figure 2 It is a flowchart of a data processing method provided by an embodiment of the present invention; Figure 3 It is a structure diagram of an encoder in a Transformer model; Figure 4 It is a schematic diagram of attention calculation. Detailed implementation manners
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0014] It should be noted that in the description of the present invention, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0015] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will further elaborate on the present invention with reference to the drawings and specific implementation manners.
[0016] Some key terms used in the embodiments of the present invention will be explained here first.
[0017] Model privacy inference is a technology for performing inference while protecting data and model privacy, aiming to prevent the leakage of user data and model parameters during the inference process.
[0018] Model privacy inference mainly involves two aspects of privacy protection: data privacy protection to prevent the leakage of user input data during the inference process; model privacy protection to prevent model parameters from being obtained by malicious users or competitors.
[0019] Homomorphic encryption and secure multi-party computation are important technical means to achieve data security and privacy protection. They have provable security and are applied to model privacy inference in related technologies.
[0020] Specifically, homomorphic encryption (HE) supports computing on ciphertext data without prior decryption. The result after computation is still in an encrypted state, and after decryption, it is the result that the user wants, that is, the same as the result of "performing corresponding computations on plaintext data".
[0021] Secure multi-party computation (MPC) can enable multiple parties to collaboratively compute any function on secret data without leaking any other secret information except the function output.
[0022] However, using provably secure privacy computing methods usually means extremely high protocol communication and computational overheads. For example, secure multi-party computation relies heavily on communication and interaction, resulting in high communication costs and communication delays; homomorphic encryption has a high computational cost, and the computational amount of computing on encrypted data is significantly higher than that of computing on plaintext data, which will also lead to an extension of the computing time.
[0023] In addition, with the continuous increase in the number of model parameters and the improvement of users' requirements for inference performance, the solution of encrypting all or part of the model parameters and deploying them on user devices to achieve edge computing can no longer solve the problem. It is inevitable to introduce large-scale computing clusters to provide computing resource support in model privacy inference. In this scenario, solving the problem of ensuring user data security during the inference process of artificial intelligence models is even more urgent.
[0024] In response to this, an embodiment of the present invention provides a privacy inference scheme that combines secure multi-party computing and homomorphic encryption computing. The computing device receives the shards of model parameters sent by the first device through secure multi-party computing, and receives the first ciphertext of the inference input data sent by the user device through homomorphic encryption computing. The computing device performs calculations based on the shards of model parameters and the first ciphertext, implementing a hybrid encryption mechanism based on the computing device, which can effectively avoid the leakage of model privacy parameters or user inference data. For the remaining calculations, the computing device encrypts the ciphertext of the intermediate calculation result shards and sends them to the user device for aggregation and decryption, so that the remaining calculations of the model inference calculation can be performed based on the plaintext of the intermediate calculation result, reducing the computational amount compared to performing the remaining calculations on the ciphertext basis, and avoiding the leakage of model privacy parameters to the user device. Thus, the present invention can not only avoid the leakage of model privacy parameters or user inference data in model inference, but also make full use of computing resources. The user device only needs to perform encryption and decryption calculations and a small amount of simple calculations without performing model calculations, and through the method of mixed plaintext and ciphertext calculation, the communication volume between computing devices is also reduced. Therefore, the present invention improves the model inference efficiency while ensuring data confidentiality.
[0025] It should be noted that the structure of the model adopted in the embodiments of the present invention can be a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM) network, etc., or a model constructed by an attention network, such as a Transformer model, a Bidirectional Encoder Representations from Transformers (BERT) model, a Contrastive Language-Image Pre-training (CLIP) model, etc. The present invention does not make any limitations here. An attention network refers to a network model that uses an attention mechanism for training. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence, so that the model finally obtains a more accurate output.
[0026] Figure 1 It is an architecture diagram of a data processing system provided by an embodiment of the present invention.
[0027] Such as Figure 1As shown in the figure, the data processing system provided by the embodiments of the present invention may include multiple computing devices 101 and user devices 102. The computing device 101 is used to calculate the ciphertext of the intermediate calculation result shards based on the model parameter shards sent by the first device 103 and the first ciphertext of the inference input data sent by the user device 102; encrypt the ciphertext of the intermediate calculation result shards to obtain the second ciphertext; send the second ciphertext to the user device 102; receive the third ciphertext sent by the user device 102 and decrypt it to obtain the plaintext of the intermediate calculation result; perform the remaining calculations of the model inference calculation using the plaintext of the intermediate calculation result to obtain the local output result shards; send the output result shards to the user device 102. The user device 102 is used to aggregate the second ciphertexts of multiple computing devices 101 and decrypt them using the homomorphic encryption algorithm to obtain the third ciphertext, and is also used to aggregate the output result shards of multiple computing devices 101 to obtain the data processing result.
[0028] In the embodiments of the present invention, the first device 103 is the model owner, the user device 102 is the owner of the user inference data, and the computing cluster composed of multiple computing devices 101 (inference servers or contracting servers) provides computing resources for the model inference calculation.
[0029] Since the model training data and the trained model parameters are important assets of the model owner, it is crucial to protect the model privacy parameters in the distributed inference service. The embodiments of the present invention implement a secure outsourcing scheme for the model inference service. The first device 103, as the model owner, distributes the model parameters of the artificial intelligence model to multiple computing devices 101 through the method of secret sharing. The computing device 101 performs secure multi-party calculation on the inference input data of the user and cooperates to complete the model inference task. Each computing device 101 only has the secret share of the model privacy parameters and cannot infer the complete model parameter information.
[0030] In addition, the inference input data of the user may contain personal information, personal privacy, trade secrets, etc. Inputting the inference input data into the model in plaintext form will cause relatively high security risks. Technologies such as homomorphic encryption and secure multi-party calculation can be used to protect the confidentiality of the inference data, but secure multi-party calculation depends more on communication and interaction, and the homomorphic encryption calculation cost is relatively high. How to balance the computational complexity and communication complexity has become an important issue faced by the privacy inference of neural networks.
[0031] Since the amount of data in the user's inference input data is small (for example, it can be a piece of text or a picture), performing homomorphic encryption on the user's inference input data requires less computational effort and subsequent calculations compared to performing homomorphic encryption on the model's privacy parameters. Therefore, in the embodiments of the present invention, the parameter distribution module of the first device 103 distributes the model's privacy parameters to the computing device 101 in a secret sharing manner. Each computing device 101 has a share of the parameters (denoted as a shard of the model parameters) for performing joint inference. The encryption and decryption module of the user device 102 encrypts the user's inference input data and sends it to the computing device 101. The computing device 101 performs plaintext-ciphertext calculations in a secure multi-party computing manner and returns the calculation result (denoted as an output result shard) to the user device 102. The user device 102 then decrypts the calculation result or aggregates the ciphertext data through the encryption and decryption module to obtain the plaintext data processing result (i.e., the plaintext inference result of the model inference task).
[0032] It should be noted that the model privacy parameters sent by the first device 103 to the computing device 101 in a secret sharing manner are not all the model parameters. That is, the model parameters also include some parameters that do not need to be kept secret, such as the model parameters used for vector quantization encoding calculations of the inference input data in the model, the parameters of the activation functions (such as the Softmax function) of each layer in the model, etc. These model parameters that do not need to be kept secret can be sent to each computing device 101 in plaintext in full quantity. Thus, the model parameters owned by a computing device 101 include the shards of the model privacy parameters and the full quantity of model parameters that do not need to be kept secret.
[0033] In model inference calculations, to reduce the computational effort of the calculation steps that need to aggregate intermediate calculation results, the ciphertexts of the intermediate calculation result shards obtained from plaintext-ciphertext calculations are sent to the user device 102 for decryption. To prevent the user device 102 from inferring the model privacy parameters when aggregating the intermediate calculation result shards, the computing device 101 encrypts the ciphertexts of the intermediate calculation result shards locally to obtain second ciphertexts and sends the second ciphertexts to the user device 102. Thus, after the user device 102 aggregates the second ciphertexts and decrypts them using the homomorphic encryption algorithm, it obtains the third ciphertext corresponding to the complete intermediate calculation result and cannot infer the model privacy parameters from it.
[0034] The embodiments of the present invention also provide a data processing method. Referring to the data processing system described in the above embodiments and combining the execution process of the data processing method, the method will be described in detail below.
[0035] Figure 2 It is a flowchart of a data processing method provided by an embodiment of the present invention.
[0036] As Figure 2As shown, when applied to a computing device, the data processing method provided by an embodiment of the present invention includes:
[0037] S201: Calculate the ciphertext of the intermediate calculation result shard according to the model parameter shards sent by the first device and the first ciphertext of the inference input data sent by the user device.
[0038] S202: Encrypt the ciphertext of the intermediate calculation result shard to obtain a second ciphertext.
[0039] S203: Send the second ciphertext to the user device so that the user device aggregates the second ciphertexts of multiple computing devices and uses the homomorphic encryption algorithm to decrypt to obtain a third ciphertext.
[0040] S204: Receive the third ciphertext and decrypt it to obtain the plaintext of the intermediate calculation result.
[0041] S205: Use the plaintext of the intermediate calculation result to perform the remaining calculations of the model inference calculation to obtain the local output result shard.
[0042] S206: Send the output result shard to the user device so that the user device aggregates the output result shards of multiple computing devices to obtain the data processing result.
[0043] It should be noted that the encryption and decryption mechanism provided by the embodiment of the present invention in the model inference calculation can be steps executed sequentially once, or can be circular execution of some or all of the steps to solve the data confidentiality problem of different types of calculations in models with different structures.
[0044] For S201, the first device distributes the model privacy parameters to the computing devices in a secret sharing manner, that is, each computing device has a shard of the model parameters. This model parameter shard only refers to the shard of the model privacy parameters, that is, each computing device can have the full amount of model parameters that do not need to be kept secret. Thus, the model parameters owned by a computing device include the model parameter shards of the model privacy parameters and the full amount of model parameters that do not need to be kept secret.
[0045] For example, when the model is a Transformer model, the model privacy parameters involved are mainly weight matrices, including the query weight matrix for attention calculation in the attention layer , the key weight matrix and the value weight matrix , the matrix for combining the outputs of multiple heads in the multi-head attention calculation , the full connection layer linear transformation parameters of the fully connected layer , , , For these parameters, the first device sends the model parameters to the computing devices in a sharded manner using secret sharing, and each computing device holds a shard of the model parameters.
[0046] The word embedding matrix and the position encoding matrix in the vector encoding layer of the Transformer model are model parameters that do not need to be kept secret. For these parameters, the first device sends the full model parameters to the computing devices, that is, each computing device can hold the full model parameters that do not need to be kept secret.
[0047] The user device encrypts the inference input data using a homomorphic encryption algorithm, and the resulting ciphertext is denoted as the first ciphertext. The user device sends the first ciphertext to each computing device.
[0048] Denote as the homomorphic encryption plaintext encoding algorithm, as the homomorphic encryption plaintext decoding algorithm, and as the homomorphic encryption and decryption algorithms respectively, and m is the inference input data of the user device.
[0049] Then the user device encodes the inference input data using the homomorphic encryption plaintext encoding algorithm and then encrypts it using the homomorphic encryption algorithm (this process can be denoted as ), obtains the first ciphertext, and the user device then sends the first ciphertext to each computing device.
[0050] Taking the fully homomorphic encryption algorithm (Cheon-Kim-Kim-Song, CKKS) as an example, the homomorphic encryption plaintext encoding algorithm encodes the vector into a polynomial over the cyclotomic integer ring , and the ciphertext is a pair of polynomials over .
[0051] Let be a vector, and homomorphic encryption supports plaintext-ciphertext calculation and ciphertext-ciphertext calculation:
[0052] , representing plaintext-ciphertext addition calculation, represents ciphertext polynomial addition, represents element-wise addition of vectors. That is, the result of adding the ciphertext obtained by homomorphically encrypting the homomorphic encryption plaintext encoding of m and the plaintext after homomorphic encryption plaintext encoding of a, decrypting and then decoding, is the same as the result of .
[0053] , representing plaintext-ciphertext multiplication calculation, represents ciphertext polynomial multiplication, Denotes element-wise multiplication of vectors. That is, it is the result of multiplying the ciphertext obtained by homomorphically encrypting the homomorphic encryption plaintext encoding of m with the plaintext after homomorphic encryption plaintext encoding of a, decrypting and then decoding, and the obtained result is the same as the result of
[0054] , denotes ciphertext-ciphertext addition, 、 denotes the inference input data, and are respectively the homomorphic encryption plaintext encodings corresponding to and , denotes ciphertext polynomial addition, denotes element-wise addition of vectors. That is, it is the result of multiplying the ciphertext obtained by homomorphically encrypting the homomorphic encryption plaintext encoding of and the ciphertext obtained by homomorphically encrypting the homomorphic encryption plaintext encoding of , decrypting and then decoding, and the obtained result is the same as the result of
[0055] The above vector operations can be extended to matrix operations. Since the model parameters are partitioned in plaintext form, the computing device only needs to perform plaintext-ciphertext addition calculation, plaintext-ciphertext multiplication calculation, ciphertext-ciphertext addition calculation, and ciphertext shift calculation.
[0056] It should be noted that the homomorphic encryption algorithm adopted by the user device is not limited to CKKS.
[0057] Based on the above principle, the computing device performs plaintext-ciphertext calculation based on the model parameter partition and the first ciphertext, and the obtained is the ciphertext of the intermediate calculation result partition. For any calculation that needs to directly use the model privacy parameter and the inference input data in the model inference calculation, the plaintext-ciphertext calculation method provided by S201 can be adopted.
[0058] If the ciphertext of the intermediate calculation result partition is not decrypted, the computing device can also perform the remaining calculations of the model inference calculation based on the ciphertext of the intermediate calculation result partition. However, the calculations performed on the basis of the homomorphic encryption ciphertext are much more computationally intensive than performing the same calculations on the basis of plaintext. To save computing resources and improve the efficiency of model inference calculation, in the embodiments of the present invention, the computing device sends the ciphertext of the intermediate calculation result partition to the user device for homomorphic encryption decryption, so as to perform the remaining calculations of the model inference calculation based on the plaintext of the intermediate calculation result partition.
[0059] However, if the computing device directly sends the ciphertext of the sharded intermediate calculation result to the user device, and the user device performs aggregation and decryption to obtain the complete intermediate calculation result, it can infer the model privacy parameters (such as model weight parameters) based on the local inference input data, which will lead to the leakage of the model privacy parameters to the user device.
[0060] For S202, the computing device encrypts the ciphertext of the sharded intermediate calculation result to obtain the second ciphertext, and in S203, sends the second ciphertext to the user device. At this time, the user device performs aggregation and homomorphic encryption decryption. After removing the homomorphic encryption, the user device also cannot know what the intermediate calculation result is, so it cannot infer the model privacy parameters from the obtained third ciphertext.
[0061] For S204, the computing device obtains the third ciphertext sent by the user device and then decrypts the third ciphertext based on the local encryption and decryption algorithm to obtain the plaintext of the intermediate calculation result. It should be noted that although the computing device can obtain the plaintext of the complete intermediate calculation result at this time, since the computing device only holds the shards of the model parameters of the model privacy parameters and the computing device does not have the plaintext of the inference input data, the computing device cannot infer the model privacy parameters based on the plaintext of the intermediate calculation result.
[0062] For S205, the computing device can then perform the remaining calculations of the model inference calculation based on the plaintext of the intermediate calculation result, so as to obtain high-performance model inference calculation performance with a relatively small amount of calculation while ensuring the confidentiality of the model privacy parameters and the confidentiality of the inference input data.
[0063] In the model inference calculation, if the adopted model includes a multi-layer network structure, each layer of the network structure can select the encryption and decryption mechanism in the above S201~S205 according to the calculation type. For the network result that is not the first layer, its input is the output of the previous layer of the network, then S201 can be replaced with: calculating the sharded intermediate calculation result according to the shards of the model parameters and the output of the previous layer of the network.
[0064] For S206, the computing device sends the shards of the output result of the last layer of the model to the user device, and the user device aggregates the shards of the output results sent by each computing device to obtain the complete data processing result.
[0065] The data processing method provided by the embodiments of the present invention is such that the computing device receives the shards of model parameters sent by the first device through the method of secure multi-party computation, and receives the first ciphertext of the inference input data sent by the user device through the method of homomorphic encryption computation. It performs computations based on the shards of model parameters and the first ciphertext, implementing a hybrid encryption mechanism based on the computing device, which can effectively avoid the leakage of model privacy parameters or user inference data. For the remaining computations, the computing device encrypts the ciphertext of the shards of the intermediate computation results and sends them to the user device for aggregated decryption, so that the remaining computations of the model inference computation can be performed based on the plaintext of the intermediate computation results, reducing the computational amount compared to performing the remaining computations on the ciphertext basis and avoiding the leakage of model privacy parameters to the user device. Thus, the present invention can not only avoid the leakage of model privacy parameters or user inference data during model inference, but also make full use of computing resources. The user device only needs to perform encryption and decryption computations and a small amount of simple computations without performing model computations. Through the method of mixed plaintext and ciphertext computation, the communication volume between computing devices is also reduced, improving the model inference efficiency on the premise of ensuring data confidentiality.
[0066] As introduced in the above embodiments, the first device can send all the model parameters that do not need to be kept secret to the computing device. Then the computing device can perform the computations other than the model privacy inference in the model inference computation based on all the model parameters that do not need to be kept secret, realizing that no additional model inference computations need to be performed on the user device or other devices.
[0067] Then, the calculation of the ciphertext of the shards of the intermediate computation results according to the shards of model parameters sent by the first device and the first ciphertext of the inference input data sent by the user device in S201 may include: performing vector encoding computation on the first ciphertext according to the model vector encoding computation parameters sent by the first device to obtain a first vector encoding result; calculating the ciphertext of the shards of the intermediate computation results according to the shards of model parameters and the first vector encoding result.
[0068] For the model vector encoding layer of models such as the Transformer model, the model vector encoding computation parameters may include a word embedding model and a position encoding matrix. Then, performing vector encoding computation on the first ciphertext according to the model vector encoding computation parameters sent by the first device to obtain a first vector encoding result may include: performing word embedding computation on the first ciphertext using the word embedding model, and adding the computation result to the position encoding matrix to obtain the first vector encoding result.
[0069] In a specific implementation, the first device sends the word embedding model and the position encoding matrix of the model to the computing device. After receiving the first ciphertext of the inference input data of the user device, the computing device first performs word embedding computation on the first ciphertext, and then adds it to the position encoding matrix to obtain the input X for attention computation.
[0070] Based on the above embodiments, the embodiments of the present invention further describe the steps of transmitting the intermediate calculation result between the computing device and the user device.
[0071] In the embodiments of the present invention, encrypting the ciphertext of the sharded intermediate calculation result in S202 to obtain the second ciphertext may include: adding a locally corresponding random number to the ciphertext of the sharded intermediate calculation result to obtain the second ciphertext.
[0072] In a specific implementation, the computing device encrypts the ciphertext of the sharded intermediate calculation result with a locally generated random number to obtain the second ciphertext. Then what the user device receives is the ciphertext of the sharded intermediate calculation result with the random number added. Even if the shards obtained by decrypting with the decryption algorithm of the homomorphic encryption algorithm are the shards of the intermediate calculation result with the random number added, the aggregated result is the intermediate calculation result with the random number added, and the user device cannot deduce the model privacy parameters therefrom.
[0073] In some alternative embodiments of the embodiments of the present invention, when the user device sends the intermediate calculation result with the random number added to the computing device, it may adopt the secret sharing method, that is, the user device aggregates the second ciphertexts of multiple computing devices and decrypts them using the homomorphic encryption algorithm to obtain the third ciphertext, which may include: the user device shards the result of aggregating the second ciphertexts of multiple computing devices and decrypting them using the homomorphic encryption algorithm to obtain the third ciphertext.
[0074] It should be noted that the sharding method by which the user device sends the shards of the plaintext of the intermediate calculation result with the random number added (i.e., the third ciphertext) to each computing device in the secret sharing method has nothing to do with the sharding method by which the first device sends the sharded model parameters of the model privacy parameters to each computing device in the secret sharing method. After each computing device receives the third ciphertext, it removes the local random number from the third ciphertext, and then aggregates the results of removing the random number from each computing device to obtain the plaintext of the complete intermediate calculation result.
[0075] Then in S204 at this time, receiving the third ciphertext and decrypting it to obtain the plaintext of the intermediate calculation result may include: removing the locally corresponding random number from the third ciphertext to obtain the first intermediate calculation result; aggregating the first intermediate calculation results of multiple computing devices to obtain the plaintext of the intermediate calculation result.
[0076] In the embodiments of the present invention, for the convenience of calculation, when aggregating the first intermediate calculation results, the computing device may select one or more target computing devices to perform the aggregation calculation.
[0077] For a computing device, aggregating the first intermediate computation results of multiple computing devices to obtain the plaintext of the intermediate computation result may include: the target computing device in the computing device aggregates the first intermediate computation results of multiple computing devices; receiving the plaintext of the intermediate computation result sent by the target computing device. That is to say, the computing device can use one or more target computing devices to complete the operation of aggregating to obtain the plaintext of the intermediate computation result, and then the computing device can use the plaintext of the intermediate computation result to perform the remaining computations of the model inference computation in plaintext.
[0078] Alternatively, aggregating the first intermediate computation results of multiple computing devices to obtain the plaintext of the intermediate computation result may further include: the target computing device in the computing device aggregates the first intermediate computation results of multiple computing devices to obtain the plaintext of the intermediate computation result. And using the plaintext of the intermediate computation result to perform the remaining computations of the model inference computation to obtain the local output result shard may include: the target computing device performs an activation function computation after computing the ciphertext of the intermediate computation result shard based on the model parameter shard sent by the first device and the first ciphertext of the inference input data sent by the user device, and broadcasts the activation function computation result to the computing devices; the computing device performs the remaining computations of the model inference computation based on the activation function computation result to obtain the local output result shard. That is to say, after the computing device uses one or more target computing devices to complete the aggregation to obtain the plaintext of the intermediate computation result, the target computing device can also perform the activation function computation using the plaintext of the intermediate computation result. For example, if the ciphertext of the intermediate computation result shard is calculated based on the model parameter shard sent by the first device and the first ciphertext of the inference input data sent by the user device in S201 is the computation of the attention layer, and the obtained intermediate computation result is the attention weight parameter, then the target computing device can perform the activation function computation of the attention layer based on the attention weight parameter, and then broadcast the activation function computation result to each computing device, and then the computing device performs the remaining computations in the model inference computation.
[0079] On this basis, if the ciphertext of the intermediate calculation result shard calculated according to the model parameter shards sent by the first device and the first ciphertext of the inference input data sent by the user device in S201 is a multi-head calculation, the target calculation device can correspond to the calculation heads of the multi-head calculation. After the target calculation device performs the activation function calculation after calculating the ciphertext of the intermediate calculation result shard according to the model parameter shards sent by the first device and the first ciphertext of the inference input data sent by the user device based on the plaintext of the intermediate calculation result, and broadcasts the activation function calculation result to the calculation devices, it may include: multiple target calculation devices perform the activation function calculation in parallel based on the plaintext of the intermediate calculation result, and broadcast the activation function calculation result to the calculation devices. The calculation device performs the remaining calculations of the model inference calculation based on the activation function calculation result to obtain the local output result shard, which may include: the calculation device splices the activation function calculation results of multiple calculation devices, and performs the remaining calculations of the model inference calculation according to the splicing result of the activation function calculation results to obtain the local output result shard.
[0080] That is to say, when using the target calculation device to perform the activation function calculation of the multi-head attention calculation, parallel calculation of the activation function calculations of each attention head can be realized, and then each target calculation device broadcasts the local activation function calculation result to each calculation device, and each calculation device splices the activation function calculation results and uses the full connection layer weight shards of the local attention layer to perform a linear transformation on the splicing result to obtain the attention layer output result shard, and then perform the remaining calculations in the model inference calculation based on the attention layer output result shard.
[0081] The data processing method provided by the embodiments of the present invention may further include: re-determining the target calculation device before performing the calculation of different layers of the model inference calculation. That is to say, after each time it is necessary to send the ciphertext of the intermediate calculation result shard to the user device for decryption and obtain the third ciphertext, the target calculation device can be re-selected from the calculation devices.
[0082] In a specific implementation, re - determining the target computing device may include: determining the current target computing device according to the previously selected target computing device and the state parameters of each computing device at the current moment. Further, a scoring method can be set for these two factors. For a computing device, if it was the target computing device in the previous calculation, a larger penalty parameter can be set; if the state parameters of the current computing device indicate better performance and lower load of the computing device, a larger priority value can be set. Subtract the penalty parameter of the computing device from the priority value of the computing device to obtain the score of the computing device. Sort the computing devices in descending order of the scores, and select the corresponding number of computing devices at the front according to the number of computing heads for multi - head calculation as the target computing devices. In this way, it is possible to avoid selecting the same computing device as the target computing device too many times, and select computing devices with better performance and lower load as the target computing devices, ensuring the computing performance and security of each multi - head calculation.
[0083] In some other alternative implementation manners of the embodiments of the present invention, after the user device aggregates the second ciphertexts of multiple computing devices and decrypts them using the homomorphic encryption algorithm to obtain the third ciphertext, in S204, when the computing device receives the third ciphertext and decrypts it to obtain the plaintext of the intermediate calculation result, it may further include: performing a random - number removal process on the third ciphertext according to the random numbers of multiple computing devices to obtain the plaintext of the intermediate calculation result. That is to say, after the computing device obtains the intermediate calculation result of the homomorphic encryption decryption by the user device, it can collect the random numbers of other computing devices, so as to locally remove all the random numbers from the third ciphertext to obtain the plaintext of the intermediate calculation result. At this time, the computing device does not need to use the target computing device to complete the random - number removal calculation, and can directly obtain the plaintext of the intermediate calculation result and perform the remaining calculations in the model inference calculation based on the plaintext of the intermediate calculation result.
[0084] To enable the computing device to conveniently obtain the random numbers of other computing devices, the same random - number generation algorithm can be pre - configured for each computing device, so that the computing device can calculate the random numbers of other computing devices through the random - number generation algorithm and the number of computing devices.
[0085] Alternatively, the selected target computing device can also generate random numbers and broadcast them to each computing device, so that each computing device can calculate all the random numbers added to the third ciphertext according to the locally stored random numbers and the number of computing devices after obtaining the third ciphertext.
[0086] In an embodiment of the present invention, in S205, performing the remaining calculations of the model inference calculation using the plaintext of the intermediate calculation result to obtain the local output result shard may include: receiving the inference input data shard sent by the user device; performing the remaining calculations of the model inference calculation according to the inference input data shard, the plaintext of the intermediate calculation result, and the model parameter shard to obtain the local output result shard.
[0087] In a specific implementation, for the remaining calculations in the model inference calculation that still require the inference input data, at this time, the user device may send the inference input data shard to each computing device in a secret sharing manner, and each computing device performs the remaining calculations of the model inference calculation based on the inference input data shard, the plaintext of the intermediate calculation result, and the model parameter shard in a secure multi-party calculation manner to obtain the local output result shard.
[0088] Specifically, performing the remaining calculations of the model inference calculation according to the inference input data shard, the plaintext of the intermediate calculation result, and the model parameter shard to obtain the local output result shard may include: calculating the local first fully connected layer output shard using the fully connected layer linear transformation parameter shard in the model parameter shard and the activation function calculation result obtained according to the plaintext of the intermediate calculation result; performing residual connection calculation and normalization calculation according to the inference input data shard and the first fully connected layer output shard to obtain the normalization calculation result shard; performing a feedforward neural network calculation according to the normalization calculation result shard and the feedforward neural network parameter shard in the model parameter shard to obtain the secret sharing share of the current layer output; if the current layer is not the last layer of the model, using the secret sharing share of the current layer output as the input data of the next layer of the model; if the current layer is the last layer of the model, performing an output layer calculation according to the secret sharing share of the previous layer output to obtain the local output result shard.
[0089] Since the model usually includes a multi-layer network structure, for example, Transformer may include multiple encoding layers and multiple decoding layers, and CNN may include multiple convolutional layers, for the network structure that is not the first layer, its input is the output of the previous layer network, that is, the steps provided in the embodiment of the present invention are repeatedly executed until the last layer. The last layer of the model is usually called the output layer. If the current layer is the last layer of the model, then perform an output layer calculation according to the secret sharing share of the previous layer network output to obtain the local output result shard.
[0090] Based on the above embodiments, an embodiment of the present invention provides a specific model privacy inference solution taking the Transformer model as an example.
[0091] Figure 3 It is a structural diagram of an encoder in a Transformer model; Figure 4 It is a schematic diagram of attention calculation.
[0092] As a deep learning model architecture for natural language processing (NLP) tasks, the self-attention mechanism of the Transformer has become a key innovation in the Transformer architecture. The self-attention mechanism allows the model to assign different attention weights according to different parts of the input sequence, thereby better capturing semantic relationships, and thus performs excellently in processing sequence data. In the embodiments of the present invention, the Transformer model is taken as an example to illustrate the model privacy inference steps based on the hybrid encryption mechanism, and this method can be extended to other neural network models.
[0093] In a model based on the Transformer architecture, information such as the weight matrices in the encoder and decoder is the private asset of the model owner (i.e., the first device in the above embodiments). To protect the model privacy parameters, the first device sends the weight information shares (i.e., the model parameter shards in the above embodiments) to multiple computing devices in the way of additive secret sharing.
[0094] The structure of the encoder in the Transformer model is as Figure 3 shown. First, perform multi-head attention calculation on the inference input data; secondly, add the output of the multi-head attention calculation to the input data and then perform normalization calculation; then, perform the calculation of the fully connected layer on the output of the normalization calculation; finally, add the output result of the normalization calculation to the output result of the fully connected layer and then perform normalization calculation to obtain the output result.
[0095] Then the model parameter shards that the first device needs to send to the computing device in the way of secret sharing include the weight matrices in the attention mechanism , , , and the linear transformation parameters in the fully connected layer , , , .
[0096] Suppose the number of inference servers is , the first device randomly selects matrix , calculates , and sends the secret shares to computing devices respectively.
[0097] The weight matrices , , , , and vectors 、 are processed similarly.
[0098] In addition, the first device sends the word embedding model and the position encoding matrix to all computing devices.
[0099] After receiving the first ciphertext of the inference input data sent by the user device, the computing device first performs word embedding calculation on the first ciphertext, and then adds it to the position encoding matrix to obtain the input for the attention operation 。
[0100] The computing device performs plaintext-ciphertext attention calculation using the attention weight matrix shards in the model parameter shards. For the th computing device , the computing device performs matrix multiplication calculation, and the ciphertext of the intermediate calculation result shards includes the share of the query matrix , the share of the key matrix , the share of the value matrix 。
[0101] As Figure 3 shown, the calculation of the attention layer of the Transformer includes using the weight matrix 、 、 to perform matrix multiplication, scaling, Softmax function calculation, etc.
[0102] Because , , , so 、 、 。Since the ciphertext of the intermediate calculation result shards 、 、 are all ciphertexts, the secure multi-party calculation of the Softmax function on the homomorphic ciphertext by multiple computing devices requires a large amount of computing and communication overhead. Therefore, the embodiments of the present invention adopt the method of decrypting first and then calculating to reduce the communication volume and computing volume of the computing devices.
[0103] The computing device randomly selects plaintext matrices 、 、 as random numbers, and performs plaintext-ciphertext matrix addition 、 、 , and records the result as the second ciphertext and sends it to the user device.
[0104] After the user equipment receives the second ciphertext, it performs an addition calculation , , , and obtains , , , where , , .
[0105] Then, the user equipment decrypts , , to obtain the query matrix in plaintext form with a random number added, the key matrix , the key matrix , and the value matrix , which is the third ciphertext in the above embodiment. The presence of the random number prevents the user equipment from inferring information about the model weight matrix.
[0106] The user equipment performs additive secret sharing on the query matrix , the key matrix , and the value matrix and sends them to computing devices. Each computing device obtains the shares , , and , where , , . Note that the fragmentation method here is independent of the fragmentation method of the first device.
[0107] After each computing device obtains the shares , the key matrix , and the value matrix of , , and , it locally computes , , .
[0108] After that, the computing device sends , , to the target computing device , where is the value agreed upon by computing devices.
[0109] The target computing device performs an addition calculation to obtain the query matrix in plaintext form without the random number, the key matrix , and the value matrix and the value matrix , i.e., the plaintext of the intermediate calculation result.
[0110] Since the target computing device only has the model weight matrix shares, it cannot infer the inference input data of the user device based on , and . Also, since the target computing device does not have the plaintext of the inference input data, it cannot infer the complete model weight matrix based on , and .
[0111] The target computing device completes the remaining calculations of the attention calculation in plaintext based on , and , i.e., calculates , where is the number of columns of matrix , represents the th attention calculation, represents the calculation result of the th activation function.
[0112] The above steps are one attention calculation. In the multi-head attention calculation, the above steps are executed t times, where t is the number of attention heads. Different attention calculations can negotiate to select different target computing devices to execute to achieve distributed computing. After the multi-head attention calculation is completed, each target computing device broadcasts the calculation result of the activation function , and all computing devices obtain the concatenated matrix .
[0113] Then, the computing device uses the matrix shares to calculate and obtains the additive secret sharing of .
[0114] The computing device performs the remaining calculations in the model inference calculation, including the secure multi-party calculations of residual connection, normalization, and fully connected layers.
[0115] Specifically, the computing device uses the additive secret sharing shares and to calculate in a secure multi-party calculation manner and obtains the additive secret sharing share , where , , the The secret share (inference input data shard) of the inference input data sent by the user device to the computing device.
[0116] The computing device utilizes the additive secret sharing shares , , , , to calculate in a secure multi-party computing manner and obtain the secret sharing share of the encoder output .
[0117] The above steps complete the privacy-preserving calculation of one encoder. Repeating the above steps completes the stacked calculation of multiple encoders.
[0118] The calculation process of the decoder is similar.
[0119] The computing device sends the secret share (output result shard) of the prediction output of the final Softmax function to the user device, and the user device aggregates the secret shares of each computing device to obtain the inference result (data processing result).
[0120] It can be understood that, in addition to the Transformer model, the data processing method provided in each of the above embodiments of the present invention can also be applied to distributed inference calculations based on other models.
[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner.
[0122] An embodiment of the present invention also provides a data processing device, which is applied to a computing device and includes: A first calculation unit, configured to calculate the ciphertext of the intermediate calculation result shard according to the model parameter shard sent by the first device and the first ciphertext of the inference input data sent by the user device; A first encryption unit, configured to encrypt the ciphertext of the intermediate calculation result shard to obtain a second ciphertext; A first sending unit, configured to send the second ciphertext to the user device, so that the user device aggregates the second ciphertexts of multiple computing devices and decrypts them using the homomorphic encryption algorithm to obtain a third ciphertext; A first decryption unit, configured to receive the third ciphertext and decrypt it to obtain the plaintext of the intermediate calculation result; A second calculation unit, configured to perform the remaining calculations of the model inference calculation using the plaintext of the intermediate calculation result to obtain the local output result shard; A second sending unit, configured to fragmentarily send the output result to a user device, so that the user device aggregates the output result fragments of multiple computing devices to obtain a data processing result.
[0123] In an embodiment of the present invention, the first computing unit calculates a ciphertext of an intermediate computing result fragment according to the model parameter fragments sent by the first device and the first ciphertext of the inference input data sent by the user device, which may include: performing vector encoding calculation on the first ciphertext according to the model vector encoding calculation parameters sent by the first device to obtain a first vector encoding result; calculating the ciphertext of the intermediate computing result fragment according to the model parameter fragments and the first vector encoding result.
[0124] In an embodiment of the present invention, the model vector encoding calculation parameters may include a word embedding model and a position encoding matrix; the first computing unit performs vector encoding calculation on the first ciphertext according to the model vector encoding calculation parameters sent by the first device to obtain a first vector encoding result, which may include: performing word embedding calculation on the first ciphertext by using the word embedding model, and adding the calculation result to the position encoding matrix to obtain the first vector encoding result.
[0125] In an embodiment of the present invention, the first encryption unit encrypts the ciphertext of the intermediate computing result fragment to obtain a second ciphertext, which may include: adding a locally corresponding random number to the ciphertext of the intermediate computing result fragment to obtain the second ciphertext.
[0126] In an embodiment of the present invention, the user device aggregates the second ciphertexts of multiple computing devices and decrypts them by using a homomorphic encryption algorithm to obtain a third ciphertext, which may include: the user device fragments the result of aggregating the second ciphertexts of multiple computing devices and decrypting them by using the homomorphic encryption algorithm to obtain the third ciphertext. The first decryption unit receives the third ciphertext and decrypts it to obtain the plaintext of the intermediate computing result, which may include: removing the locally corresponding random number from the third ciphertext to obtain a first intermediate computing result; aggregating the first intermediate computing results of multiple computing devices to obtain the plaintext of the intermediate computing result.
[0127] In an embodiment of the present invention, the first decryption unit aggregates the first intermediate computing results of multiple computing devices to obtain the plaintext of the intermediate computing result, which may include: the target computing device in the computing device aggregates the first intermediate computing results of multiple computing devices; receiving the plaintext of the intermediate computing result sent by the target computing device.
[0128] In an embodiment of the present invention, the first decryption unit aggregates the first intermediate calculation results of multiple computing devices to obtain the plaintext of the intermediate calculation results. It may further include: the target computing device in the computing devices aggregates the first intermediate calculation results of multiple computing devices to obtain the plaintext of the intermediate calculation results. The second calculation unit uses the plaintext of the intermediate calculation results to perform the remaining calculations of the model inference calculation to obtain the local output result shard, which may include: the target computing device performs an activation function calculation after calculating the ciphertext of the intermediate calculation result shard based on the model parameter shard sent by the first device and the first ciphertext of the inference input data sent by the user device based on the plaintext of the intermediate calculation result, and broadcasts the activation function calculation result to the computing devices; the computing devices perform the remaining calculations of the model inference calculation based on the activation function calculation result to obtain the local output result shard.
[0129] In an embodiment of the present invention, the first calculation unit calculates the ciphertext of the intermediate calculation result shard according to the model parameter shard sent by the first device and the first ciphertext of the inference input data sent by the user device, which is multi-head calculation; the target computing device corresponds to the calculation head of the multi-head calculation. Then, the target computing device performs an activation function calculation after calculating the ciphertext of the intermediate calculation result shard based on the model parameter shard sent by the first device and the first ciphertext of the inference input data sent by the user device based on the plaintext of the intermediate calculation result, and broadcasts the activation function calculation result to the computing devices, which may include: multiple target computing devices perform the activation function calculation in parallel based on the plaintext of the intermediate calculation result and broadcast the activation function calculation result to the computing devices. The computing devices perform the remaining calculations of the model inference calculation based on the activation function calculation result to obtain the local output result shard, which may include: the computing devices splice the activation function calculation results of multiple computing devices and perform the remaining calculations of the model inference calculation according to the splicing result of the activation function calculation results to obtain the local output result shard.
[0130] The data processing device provided in the embodiment of the present invention may further include: a determination unit, configured to re-determine the target computing device before performing the calculations of different layers of the model inference calculation.
[0131] In an embodiment of the present invention, the second calculation unit uses the plaintext of the intermediate calculation results to perform the remaining calculations of the model inference calculation to obtain the local output result shard, which may include: receiving the inference input data shard sent by the user device; performing the remaining calculations of the model inference calculation according to the inference input data shard, the plaintext of the intermediate calculation results, and the model parameter shard to obtain the local output result shard.
[0132] In an embodiment of the present invention, the second computing unit performs the remaining calculations of the model inference calculation based on the inference input data shards, the plaintext of the intermediate calculation result, and the model parameter shards, and obtains the local output result shards, which may include: calculating the local first fully connected layer output shard by using the fully connected layer linear transformation parameter shards in the model parameter shards and the activation function calculation result obtained according to the plaintext of the intermediate calculation result; performing residual connection calculation and normalization calculation based on the inference input data shards and the first fully connected layer output shards to obtain the normalization calculation result shards; performing a feedforward neural network calculation based on the normalization calculation result shards and the feedforward neural network parameter shards in the model parameter shards to obtain the secret sharing share of the current layer output; if the current layer is not the last layer of the model, using the secret sharing share of the current layer output as the input data of the next layer of the model; if the current layer is the last layer of the model, performing an output layer calculation based on the secret sharing share of the previous layer output to obtain the local output result shards.
[0133] For the description of the features in the corresponding embodiment of the data processing device, reference may be made to the relevant description in the corresponding embodiment of the data processing method, which will not be elaborated here.
[0134] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data processing method embodiments.
[0135] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned data processing method embodiments when running.
[0136] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0137] An embodiment of the present invention further provides a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned data processing method embodiments.
[0138] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned data processing method embodiments.
[0139] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0140] The above has introduced in detail a data processing method, system, device, medium, and computer program product provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A data processing method, characterized in that: Applied to computing devices, including: Obtaining a ciphertext of an intermediate calculation result slice according to the model parameter slice sent by the first device and the first ciphertext of the inference input data sent by the user device; Encrypting the ciphertext of the intermediate calculation result slice to obtain a second ciphertext; Sending the second ciphertext to the user device, so that the user device aggregates the second ciphertexts of multiple computing devices and decrypts them using a homomorphic encryption algorithm to obtain a third ciphertext; Receiving the third ciphertext and decrypting it to obtain a plaintext of the intermediate calculation result; Utilize the plain text of the intermediate calculation result to perform the remaining calculation of the model reasoning calculation to obtain a local output result fragment; The output result slices are sent to the user device, so that the user device aggregates the output result slices of multiple computing devices to obtain a data processing result.
2. The data processing method according to claim 1, characterized in that: The ciphertext of the intermediate calculation result slice is obtained by calculating according to the model parameter slice sent by the first device and the first ciphertext of the inference input data sent by the user device, including: Performing vector encoding calculation on the first ciphertext according to the model vector encoding calculation parameters sent by the first device to obtain a first vector encoding result; The ciphertext of the intermediate calculation result slice is calculated according to the model parameter slice and the first vector encoding result.
3. The data processing method according to claim 2, characterized in that: The model vector encoding calculation parameters include a word embedding model and a position encoding matrix; Performing vector encoding calculation on the first ciphertext according to the model vector encoding calculation parameter sent by the first device to obtain a first vector encoding result includes: The word embedding model is used to perform word embedding calculation on the first ciphertext, and the calculation result is added to the position coding matrix to obtain the first vector coding result.
4. The data processing method according to claim 1, characterized in that: Encrypting the ciphertext of the intermediate calculation result fragment to obtain a second ciphertext includes: Add a locally corresponding random number to the ciphertext of the intermediate calculation result fragment to obtain the second ciphertext.
5. The data processing method according to claim 4, characterized in that: The user device aggregates the second ciphertexts of the plurality of computing devices and decrypts them using a homomorphic encryption algorithm to obtain a third ciphertext, including: The user device aggregates the second ciphertexts of the plurality of computing devices and decrypts the result by using a homomorphic encryption algorithm to slice the result to obtain the third ciphertext; The third ciphertext is received and decrypted to obtain the plaintext of the intermediate calculation result, including: Remove the locally corresponding random number from the third ciphertext to obtain a first intermediate calculation result; Aggregate the first intermediate calculation results of the plurality of the computing devices to obtain a plain text of the intermediate calculation result.
6. The data processing method according to claim 5, characterized in that: Aggregating the first intermediate calculation results of the plurality of the computing devices to obtain a plain text of the intermediate calculation result includes: A target computing device among the computing devices aggregates the first intermediate computing results of a plurality of the computing devices; Receive the plain text of the intermediate calculation result sent by the target computing device.
7. The data processing method according to claim 5, characterized in that: Aggregating the first intermediate calculation results of the plurality of the computing devices to obtain a plain text of the intermediate calculation result includes: A target computing device among the computing devices aggregates the first intermediate computing results of a plurality of the computing devices to obtain a plain text of the intermediate computing results; The remaining calculations of the model inference calculation are performed using the plain text of the intermediate calculation result to obtain local output result fragments, including: The target computing device performs, based on the plain text of the intermediate computing result, the activation function calculation after the ciphertext of the intermediate computing result slice is obtained by calculating the model parameter slice sent by the first device and the first ciphertext of the inference input data sent by the user device, and broadcasts the activation function calculation result to the computing device; The computing device performs the remaining calculations of the model inference calculation based on the activation function calculation result to obtain the local output result fragment.
8. The data processing method according to claim 7, characterized in that: The ciphertext of the intermediate calculation result slice obtained by calculating the model parameter slice sent by the first device and the first ciphertext of the inference input data sent by the user device is multi-head calculation; the target calculation device corresponds to the calculation head of the multi-head calculation; The target computing device performs, based on the plain text of the intermediate computing result, the activation function calculation after the ciphertext of the intermediate computing result slice is obtained by calculating the model parameter slice sent by the first device and the first ciphertext of the inference input data sent by the user device, and broadcasts the activation function calculation result to the computing device, including: The plurality of target computing devices execute the activation function calculation in parallel based on the plain text of the intermediate calculation result, and broadcast the activation function calculation result to the computing device; The computing device performs the remaining calculations of the model inference calculation based on the activation function calculation result to obtain the local output result fragment, including: The computing device splices the activation function calculation results of multiple computing devices, performs the remaining calculations of the model inference calculation according to the splicing results of the activation function calculation results, and obtains the local output result fragment.
9. The data processing method according to any one of claims 6 to 8, characterized in that: Also includes: Before executing different layers of calculation of the model inference calculation, the target computing device is re-determined.
10. The data processing method according to claim 1, characterized in that: The remaining calculations of the model inference calculation are performed using the plain text of the intermediate calculation result to obtain local output result fragments, including: Receiving the inference input data slice sent by the user equipment; The remaining calculations of the model inference calculation are performed according to the inference input data slices, the plain text of the intermediate calculation results and the model parameter slices to obtain the local output result slices.
11. The data processing method according to claim 10, characterized in that: The remaining calculation of the model reasoning calculation is performed according to the reasoning input data slice, the plain text of the intermediate calculation result and the model parameter slice to obtain the local output result slice, including: Calculate a local first fully connected layer output slice using the fully connected layer linear transformation parameter slice in the model parameter slice and the activation function calculation result obtained according to the plain text of the intermediate calculation result; Performing residual connection calculation and normalization calculation according to the inference input data slice and the first fully connected layer output slice to obtain a normalized calculation result slice; Perform feedforward neural network calculation according to the normalized calculation result slice and the feedforward neural network parameter slice in the model parameter slice to obtain the secret sharing share of the current layer output; If the current layer is not the last layer of the model, the secret sharing share output by the current layer is used as input data for the next layer of the model; If the current layer is the last layer of the model, the output layer calculation is performed according to the secret sharing share output by the previous layer to obtain the local output result fragment.
12. A data processing system, characterized in that: include: multiple computing devices and user devices; The computing device is used to calculate the ciphertext of the intermediate computing result slice according to the model parameter slice sent by the first device and the first ciphertext of the inference input data sent by the user device; Encrypting the ciphertext of the intermediate calculation result slice to obtain a second ciphertext; Sending the second ciphertext to the user equipment; receiving and decrypting the third ciphertext sent by the user equipment to obtain a plaintext of the intermediate calculation result; Utilize the plain text of the intermediate calculation result to perform the remaining calculation of the model reasoning calculation to obtain a local output result fragment; Sending the output result slices to the user equipment; The user device is used to aggregate the second ciphertext of multiple computing devices and decrypt them using a homomorphic encryption algorithm to obtain the third ciphertext, and is also used to aggregate the output result fragments of multiple computing devices to obtain a data processing result.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Method for realizing privacy protection convolutional neural network reasoning based on model conversion
CN114912132A
Privacy reasoning method and system based on Transform network model, medium and electronic equipment
CN117077162A
Data homomorphic encryption method, system, device, equipment, medium and product
CN118118155A
Privacy protection information processing method and device based on large language model
CN118228302A
Inference engine creation method, product, equipment and computer readable storage medium
CN118469024A