Neural network model inference method and computer device
Patent Information
- Application Number
- PCT/CN2025/131195
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2025-10-30
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025131195_01102026_PF_FP_ABST
Abstract
Description
A neural network model inference method and computer device
[0001] This application claims priority to Chinese Patent Application No. 2025103944605, filed on March 28, 2025, entitled "A Neural Network Model Reasoning Method and Computer Device", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The embodiments in this specification belong to the field of computer application technology, and in particular relate to a neural network model inference method and computer equipment. Background Technology
[0003] With the development of machine learning technology, neural network models are being applied in an increasing number of fields. In some cases, there is a need for two parties to jointly utilize neural network models for inference. For example, there is a model provider and a data provider. The model provider possesses a neural network model, such as one used for facial recognition; the data provider possesses the data to be processed. Both parties need to complete the data inference and obtain the inference result while ensuring that their respective data or models are not leaked. Summary of the Invention
[0004] The purpose of this specification is to provide a neural network model inference method and computer equipment.
[0005] The first aspect of this specification provides a neural network model inference method, applied to the data side, including:
[0006] The receiving model sends a first encrypted neural network model with a first encrypted linear layer and N-1 second encrypted neural network models with second encrypted linear layers. The first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model have the same structure. The first encrypted linear layer and the N-1 second encrypted linear layers are respectively encrypted using N fragment keys, which are generated based on a distributed encryption protocol that supports homomorphic operations. The first encrypted linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameters of any second encrypted linear layer is 0.
[0007] The target data is inferred using the first ciphertext neural network model and the second ciphertext neural network model respectively, resulting in N ciphertext outputs;
[0008] The inference result of the target data is obtained by summing the N ciphertext outputs.
[0009] The second aspect of this specification provides a neural network model inference method, applied to a model having a first neural network model, including:
[0010] For the first neural network model, N-1 second neural network models with the same structure as the first neural network model are generated; the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0;
[0011] Obtain N fragment keys generated using the first encryption algorithm, and use the N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;
[0012] A first encrypted neural network model and N-1 second encrypted neural network models are sent to the data party so that the data party can perform inference using the first encrypted neural network model and N-1 second encrypted neural network models.
[0013] A third aspect of this specification provides a neural network model inference apparatus for use on a data side, comprising:
[0014] The receiving module is used to receive a first encrypted neural network model with a first encrypted linear layer and N-1 second encrypted neural network models with second encrypted linear layers sent by the model provider. The first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model have the same structure. The first encrypted linear layer and the N-1 second encrypted linear layers are respectively encrypted using N fragmentation keys, which are generated based on a distributed encryption protocol that supports homomorphic operations. The first encrypted linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameters of any second encrypted linear layer is 0.
[0015] The inference module is used to infer the target data using the first encrypted neural network model and the second encrypted neural network model respectively, and obtain N encrypted outputs; the sum of the N encrypted outputs is used to obtain the inference result of the target data.
[0016] This specification provides a fourth aspect of a neural network model inference apparatus, applied to a model having a first neural network model, comprising:
[0017] The obfuscation module is used to generate N-1 second neural network models with the same structure as the first neural network model for the first neural network model; the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0.
[0018] The encryption module is used to obtain N fragment keys generated by the first encryption algorithm, and to encrypt the first linear layer and N-1 second linear layers using the N fragment keys respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;
[0019] The sending module is used to send a first encrypted neural network model and N-1 second encrypted neural network models to the data party, so that the data party can use the first encrypted neural network model and N-1 second encrypted neural network models to perform inference.
[0020] The fifth aspect of this specification provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the neural network model inference method.
[0021] A sixth aspect of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the neural network model inference method.
[0022] A seventh aspect of this specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the neural network model inference method.
[0023] This specification provides a method for inference using a neural network model. The data provider receives N encrypted neural network models from the model provider. These N encrypted neural network models are obtained by obfuscating a first neural network model. The true neural network model includes a first linear layer, and the N-1 obfuscated pseudo neural network models each include a second linear layer with parameters set to 0. The first linear layer and the N-1 second linear layers are each encrypted using N fragment keys, which are generated based on a distributed encryption protocol supporting homomorphic operations. The data provider uses these N encrypted neural network models to infer the target data, obtaining a sum of N encrypted outputs. The inference result for the target data can then be obtained from this sum of the N encrypted outputs.
[0024] The encryption model enhances its security by using N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively. Furthermore, the sum of the outputs of the N linear layers can be decrypted using the aggregate key corresponding to the N fragment keys. Additionally, generating a fake neural network model to obfuscate the model further improves its security. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 is a schematic diagram of the system architecture in one embodiment of this specification;
[0027] Figure 2 is a flowchart of a neural network model inference method in one embodiment of this specification;
[0028] Figure 3 is a flowchart of a neural network model inference method in another embodiment of this specification;
[0029] Figure 4 is a schematic diagram of matrix-vector multiplication in one embodiment of this specification;
[0030] Figure 5 is a block diagram of a neural network model inference device in one embodiment of this specification;
[0031] Figure 6 is a block diagram of a neural network model inference device according to another embodiment of this specification. Detailed Implementation
[0032] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0033] First, the architecture of the neural network model inference method provided in this manual will be explained.
[0034] As shown in Figure 1, the method in this specification involves a model side and a data side. The model side possesses a trained neural network model. The data side possesses the data to be used for inference. In the method described in this specification, the model from the model side is deployed offline to the data side in encrypted form, and the data side uses this model for inference. In this way, the model side cannot obtain the data from the data side, ensuring the security of the data from the data side. At the same time, the neural network model from the model side is encrypted before being deployed to the data side, and the data side cannot obtain the plaintext parameters of the neural network model, ensuring the security of the model.
[0035] The neural network model can be any model that includes linear layers, such as Convolutional Neural Networks (CNNs). Neural network models can be used for image recognition, natural language processing, speech recognition, etc., and this specification does not limit the application of the neural network model. Optionally, the neural network model can be an image recognition model for face recognition, where the data provider possesses the face data to be recognized.
[0036] The following section will provide a general explanation of the method described in this specification, referring to the system architecture diagram in Figure 1. During execution, the model first obfuscates the first neural network model, generating N-1 fake neural network models. The structure of the fake neural network models is the same as that of the first neural network model, but the parameter values can be random numbers. For ease of description, the fake neural network models will be referred to as the second neural network model.
[0037] As shown in Figure 1, the first neural network model can include a first linear layer. Correspondingly, the pseudo-linear layer corresponding to the first linear layer in the second neural network model can be called a second linear layer. Unlike other pseudo-layers in the second neural network model, the parameters of the second linear layer can take the value of 0.
[0038] For the first and second linear layers, the model can generate N fragment keys based on a distributed encryption protocol that supports homomorphic operations. Data encrypted with N fragment keys can support homomorphic operations. Fragment keys have the following property: the sum of data encrypted by N fragment keys can be decrypted using the aggregate key corresponding to the N fragment keys. However, data encrypted with a single fragment key cannot be decrypted using only the aggregate key; similarly, the sum of data encrypted with N fragment keys cannot be decrypted using any fragment key. Homomorphic operations can include homomorphic addition and homomorphic multiplication. Here, we will explain one type of homomorphic operation involved in this specification. Suppose that [[a]] represents the data obtained by encrypting a with a certain fragment key, and b is the plaintext data. Then, the homomorphic multiplication of [[a]] and b results in [[a*b]]. Furthermore, the homomorphic addition of [[a]] and [[b]] results in [[a+b]].
[0039] As shown in Figure 1, the model uses N fragment keys to encrypt the first linear layer and N-1 fake linear layers respectively, resulting in the first ciphertext linear layer and N-1 second ciphertext linear layers.
[0040] This method utilizes multiple second neural network models to obfuscate the first neural network model, making it impossible for the data provider to identify which model is the first neural network model even after obtaining the encrypted model. Furthermore, the data provider cannot know that the first and second encrypted linear layers are encrypted using N fragment keys. Therefore, if the data provider wants to crack the encrypted linear layers, it needs to crack N fragment keys to obtain all the linear layers. It is evident that encryption using N fragment keys can improve the security of the encrypted neural network model.
[0041] The value of N can be determined based on the preset level of confusion. A higher level of confusion indicates a higher requirement for confidentiality on the model side, so N can be set slightly higher. However, N cannot be set indefinitely high, as increasing N will increase the number of homomorphic multiplication operations performed by the data side, which are time-consuming. Therefore, the size of N should be reasonably determined based on the computational efficiency requirements of the data side and the level of confusion required by the model side.
[0042] Furthermore, as shown in Figure 1, the first neural network model may also include a third layer. Correspondingly, the third layer in the second neural network model is called the fourth layer. The third layer and the first encrypted linear layer can form the first encrypted neural network model, and the second encrypted linear layer and the fourth layer can form the second encrypted neural network model. It should be noted that Figure 1 uses the third layer as an example. It can be understood that the first neural network model may also include a fifth layer, etc. The processing method for the fifth layer is the same as that for the third layer, and will not be repeated here. This example uses the third layer before the first linear layer as an example, and this embodiment does not represent a limitation of this specification.
[0043] Optionally, to ensure the security of other layers, a second encryption algorithm can be used to further encrypt the first and second neural network models, ensuring that each layer of the neural network model is ciphertext and thus guaranteeing the model's security. For example, the second encryption algorithm can be used to encrypt the third and fourth layers.
[0044] Furthermore, for the model's inference program, secure compilation, a trusted execution environment (TEE), and fully encrypted inference can be used to ensure that the data party performs inference on the model in encrypted form. This specification will use secure compilation as an example to illustrate how the data party implements encrypted inference.
[0045] Secure compilation refers to the use of specific methods and techniques during software development to ensure the security of the compilation process, preventing malicious attacks and the introduction of potential vulnerabilities. Secure compilation aims to improve the reliability and security of software, preventing security flaws from occurring during the building, compilation, and deployment of applications.
[0046] As shown in Figure 1, the inference program can be encrypted through secure compilation, ensuring that the data provider can only obtain the ciphertext version of the inference program, not the plaintext version. Furthermore, during the execution of the inference program by the data provider, the data generated by the inference program cannot be accessed by the data provider. This ensures the secure inference capability of the data provider.
[0047] As shown in Figure 1, the model provider can send an encrypted neural network model and a securely compiled encrypted inference program to the data provider.
[0048] The data provider can perform inference using a securely compiled inference program. The specific operations and data generated during the inference process cannot be accessed by the data provider. Specifically, in the case of a second encryption algorithm, the first encrypted neural network model can be decrypted using the key of the second encryption algorithm written into the inference program.
[0049] To further ensure the data security of neural network models, they can be decrypted in memory. If the decryption result is stored as a file on the hard drive, the data provider can obtain the plaintext parameters of the neural network model by reading the file. Decryption in memory ensures that the plaintext is not written to disk, thus guaranteeing model security.
[0050] As shown in Figure 1, the data provider can input the target data into N encrypted neural network models, namely the third layer of the first neural network model and the fourth layer of N-1 second neural network models. Then, the outputs of the third and fourth layers are input into the corresponding first encrypted linear layers and N-1 second encrypted linear layers, respectively. Within the encrypted linear layers, based on the homomorphic multiplication between the plaintext of the intermediate results and the encrypted parameters of the linear layers, N encrypted outputs corresponding to the first and second linear layers are calculated.
[0051] After obtaining N ciphertext outputs, these N ciphertext outputs can be summed to obtain the final ciphertext output. Since the parameters of the second ciphertext linear layer are 0, the plaintext output corresponding to the ciphertext output obtained by the homomorphic multiplication of the parameters of the second ciphertext linear layer and the intermediate results is also 0. Furthermore, adding the N ciphertext outputs yields the ciphertext output corresponding to the first linear layer of the original neural network. Moreover, since the parameters of the second linear layer are 0, the output of the second linear layer remains unaffected by the parameters of other layers preceding it in the second neural network model.
[0052] After obtaining the ciphertext output, if there is only one non-linear layer after the first linear layer, such as a single activation function, then the first and second neural network models mentioned above do not need to confuse the activation function; that is, both the first and second neural network models correspond to the same activation function. Correspondingly, the ciphertext output can be directly decrypted and input into the activation function to obtain the model's inference result.
[0053] Furthermore, if the final output requires processing through other layers after the first linear layer, similar to the above processing, these other layers can also be left unobfuscated. The specific processing method is similar to the above: the ciphertext output can be decrypted and input into an unobfuscated layer to obtain the inference result. Alternatively, these layers can be obfuscated, but the parameters can be set to 0 (even if all are 0, the encrypted values will be different). Correspondingly, the ciphertext output can be decrypted, and the decryption results can be input into the subsequent layers of the first linear layer of the first neural network model, and the subsequent layers of the second linear layers of N-1 second neural networks. The outputs of each layer are then summed to obtain the final inference result. The above two examples are not intended to limit this specification.
[0054] The above method ensures the confidentiality and non-theft of the entire inference process of the neural network model by using encrypted deployment of the neural network model, inference of the encrypted inference program, and decryption with aggregated keys, thus protecting the model's security. The method also obfuscates the neural network model, increasing the difficulty for data providers to crack it.
[0055] Meanwhile, this method employs an efficient and secure compilation approach, resulting in minimal computation time. Although homomorphic encryption is used, it only encrypts a single linear layer using a fragmented key, requiring a minimum of N homomorphic multiplication operations. This ensures efficient inference, guaranteeing that end-to-end processing time is less than 1.5 times that of plaintext processing.
[0056] Furthermore, this method requires no modification to the model, only the generation of an additional N-1 second linear layers.
[0057] The following sections will explain the neural network model inference method provided in this specification from the perspectives of both the model and data sides, using the flowcharts shown in Figures 2 and 3.
[0058] First, the method described in this specification will be explained from the perspective of the model owner who has a neural network model. As shown in Figure 2, it includes the following steps:
[0059] Step 201: For the first neural network model, generate N-1 second neural network models with the same structure as the first neural network model.
[0060] In this model, the second linear layer corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0.
[0061] Specifically, firstly, for the first linear layer in the neural network model, N-1 pseudo-linear layers are generated corresponding to this first linear layer. These pseudo-linear layers are used for obfuscation; their parameter matrices have the same size as the first linear layer, but the parameter values are all zero. The parameter values are zero to facilitate obtaining the original obfuscated output from the first linear layer later by summing the N obfuscated outputs. Furthermore, pseudo-obfuscated layers are also generated for the other layers of the first neural network model. These obfuscated layers together form N-1 second neural network models.
[0062] In this context, a linear layer refers to a layer that performs a linear transformation on the input, typically represented as y = Wx + b, where W is the weight matrix, b is the bias vector, x is the input, and y is the output. This specification treats linear layers in this way because they can perform linear operations. However, the first encryption algorithm is based on a distributed encryption protocol that supports homomorphic operations, which generally only support linear operations and not nonlinear operations.
[0063] The scope of obfuscation can be the entire neural network model or only a portion of its layers. In other words, the first neural network model refers to a portion of the neural network model that needs to be obfuscated. For the latter, some layers of the neural network model can be obfuscated while others remain unobfuscated. For example, if the layers before and after a certain layer are obfuscated, then the input to that layer can be the sum of the outputs of N neural network models from the previous layer, and the output of that layer can be input to the subsequent N neural network models. How to ensure the correctness of the inference result in this case can be seen in the example above, and will not be repeated here.
[0064] The first linear layer can be any linear layer in the neural network model, such as a fully connected layer, a convolutional layer, or an embedding layer. This specification does not limit the specific form of the first linear layer.
[0065] In one optional implementation, the first linear layer is the last fully connected layer of the neural network model. In the neural network model, the output of the last fully connected layer is typically input into an activation function, and the output of the activation function is the output of the neural network model. Thus, during inference, after the last fully connected layer, the sum of the N ciphertext outputs can be decrypted to obtain the output of the last fully connected layer, and this plaintext output can then be input into the activation function to obtain the output of the neural network model. This allows for more flexible inference.
[0066] It should be noted that, in order to ensure the confidentiality of the reasoning result and the reasoning process, the above decryption process can be implemented through a confidential reasoning program, thereby ensuring that the reasoning process is unknown to the data party.
[0067] Furthermore, in this case, the parameters of the second neural network model, except for the second linear layer, can be any value. Since the parameters of the second linear layer are 0, regardless of how the preceding layers are calculated, the plaintext corresponding to the output of the second linear layer will also be 0, without affecting the correctness of the inference result. Of course, the parameters of the second neural network model, except for the second linear layer, can also be 0. Later, a second encryption algorithm can be used to encrypt the first and second neural network models. Even if all parameters are 0, the encrypted values will be different, and privacy will not be leaked.
[0068] Regarding the parameter values for other layers in the second neural network model, we will use the third layer of the first neural network model as an example. The fourth layer in the second neural network model is generated based on the third layer, and the parameters of this fourth layer can be random numbers.
[0069] As mentioned above, when the first linear layer is the last layer of the first neural network model, taking a random value for the fourth layer will not affect the correctness of the inference result. In another scenario, the first neural network model may also include a fifth layer, and the sixth layer is the layer corresponding to that fifth layer in the second neural network model. The fifth layer can be the last layer of the first neural network model. In this case, the parameters of the sixth layer can be set to 0, so taking a random value for the fourth layer will not affect the correctness of the inference result. The above two examples are not intended to limit this specification.
[0070] Step 203: Obtain N fragment keys generated using the first encryption algorithm, and use the N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N-1 second ciphertext neural network models with a second ciphertext linear layer.
[0071] The first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations.
[0072] Specifically, firstly, N fragment keys generated using a distributed encryption protocol supporting homomorphic operations, and the corresponding aggregate key for each of the N fragment keys, were obtained. The characteristics of the distributed encryption protocol are detailed above and will not be repeated here. The first linear layer and N-1 fake second linear layers are then encrypted using the N fragment keys, resulting in the first encrypted linear layer and N-1 second encrypted linear layers. Different linear layers use different encryption fragment keys, and each of the N fragment keys corresponds one-to-one with one of the N linear layers.
[0073] Furthermore, although the values of the second linear layer are all 0, the data encrypted using different fragmentation keys are different. It's impossible to distinguish which is the original linear layer simply by looking at the first and second ciphertext linear layers. While the first and second ciphertext linear layers are used to differentiate the two sources, the data sender receives N ciphertext linear layers. These N linear layers do not indicate which was encrypted by the first linear layer and which by the second; therefore, the data sender cannot distinguish between the first and second ciphertext linear layers.
[0074] The specific implementation method of the first encryption algorithm can be determined based on the distributed key generation protocol corresponding to homomorphic encryption. Two examples will be given here, but these examples are not intended to limit this specification. One is a first encryption algorithm based on semi-homomorphic encryption: a distributed key protocol based on Paillier. The other is a first encryption algorithm based on fully homomorphic encryption: a distributed key protocol based on Cheon-Kim-Kim-Song (CKKS), where the polynomial coefficients are the same when different fragment keys are encrypted.
[0075] The difference between fully homomorphic and semi-homomorphic operations lies in whether they support the following: the homomorphic multiplication of [[a]] and [[b]] results in [[a*b]]. Both support homomorphic multiplication of plaintext and ciphertext, as well as homomorphic addition between ciphertexts. That is, the homomorphic multiplication of [[a]] and b results in [[a*b]]. Furthermore, the homomorphic addition of [[a]] and [[b]] results in [[a+b]].
[0076] The specific implementations of these two encryption algorithms will be detailed later, and will not be elaborated here.
[0077] In an optional implementation, in addition to obfuscating and encrypting the first linear layer, other layers of the first and second neural network models can also be encrypted. In other words, the first and second neural network models are further encrypted based on a second encryption algorithm. The second encryption algorithm can be any encryption algorithm; in an optional implementation, the second encryption algorithm can be a symmetric encryption algorithm, such as the Advanced Encryption Standard (AES).
[0078] The specific encryption range of the second encryption algorithm will be explained using the first neural network model. It is understood that the encryption model of the second neural network model is the same, so it will not be elaborated further.
[0079] In one alternative implementation, the layers of the neural network model other than the first linear layer can be encrypted. In this case, the aggregation key can be written into the encrypted inference program. Although the data provider is unaware that the first and second linear layers are encrypted using a distributed key, to ensure data security, writing the aggregation key into the encrypted inference program prevents the data provider from obtaining the plaintext aggregation key. This prevents the data provider from performing homomorphic addition operations on the first linear layer and N-1 second linear layers, and then decrypting the result using the aggregation key to obtain the parameters of the first linear layer.
[0080] In another alternative implementation, all layers of the neural network model can be encrypted. Specifically, all layers except the first linear layer can be encrypted once using the second encryption algorithm. The first and second linear layers are first encrypted using a fragmentation key, and then further encrypted using the second encryption algorithm. In this way, the aggregate key can be written into the encrypted inference program or sent directly to the data provider in plaintext. Since the first and second linear layers are encrypted twice for the encrypted neural network model, even if the data provider obtains the plaintext of the aggregate key, it cannot decrypt the first linear layer.
[0081] Furthermore, in both of the above cases, the decryption key of the second encryption algorithm can be written into the encrypted inference program to prevent the data party from obtaining the parameters of the plaintext neural network model.
[0082] In an alternative implementation, as described above, the inference program can be processed to prevent the data provider from obtaining the plaintext inference program and its intermediate results. The specific processing procedure of the inference program will be explained here using a secure compilation algorithm as an example.
[0083] The inference program can be obfuscated through secure compilation, making it unknown to the data provider. Furthermore, the inference program can be implemented based on the Triton Server framework. The core function of Triton Server is to achieve efficient deployment and scalable expansion of neural network inference services through unified model management and optimized resource scheduling, ensuring high-concurrency, low-latency inference performance.
[0084] In some cases, secure compilation can only process programs written in a specific programming language, such as C++, while the Triton server framework is written in another programming language, such as Python. In such cases, it's inconvenient to directly perform secure compilation on the Triton server framework. Therefore, the inference part of the neural network model can be separated and performed in a separate process. This process can be written in a specific programming language and can call the Triton server framework written in another language. In other words, the inference program calls the unobfuscated Triton service framework during runtime.
[0085] This allows for secure compilation of the code to protect the inference program (i.e., the process) and the keys written into it.
[0086] Furthermore, since the inference program and the inference program itself are independent of the Triton server framework, they need to communicate. To ensure data security, the Triton service framework communicates via pipes during operation. Pipe communication is a method of inter-process communication that allows communication to be implemented through memory, ensuring that the communication content is not written to disk and guaranteeing the security of the intermediate results output by the inference program.
[0087] Through the above steps, an encrypted neural network model and an inference program for the encrypted neural network model obtained by obfuscation using secure compilation methods can be obtained.
[0088] Step 205: Send the first encrypted neural network model and N-1 second encrypted neural network models to the data party so that the data party can use the first encrypted neural network model and N-1 second encrypted neural network models for inference.
[0089] After obtaining the first and second encrypted neural network models, the models can be sent to the data provider to deploy the encrypted neural network models offline. This allows the data provider to perform inference using the offline-deployed first and second encrypted neural network models.
[0090] In an alternative implementation, an inference program for the encrypted neural network model, obtained by obfuscation using a secure compilation method, may also be sent to the data provider; wherein inference of the target data to obtain the inference result is implemented based on the inference program. This protects the security of the inference program.
[0091] The following section, using Figure 3 as an example, will explain the model inference method from the data provider's perspective. As shown in Figure 3, the method includes the following steps:
[0092] Step 301: Receive the first encrypted neural network model with a first encrypted linear layer and N-1 second encrypted neural network models with a second encrypted linear layer sent by the receiving model.
[0093] The first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model have the same structure. The first encrypted linear layer and N-1 second encrypted linear layers are obtained by encryption based on N fragment keys. The N fragment keys are generated based on a distributed encryption protocol that supports homomorphic operations. The first encrypted linear layer is obtained by encrypting the first linear layer included in the first neural network model. The plaintext of the parameter of any second encrypted linear layer is 0.
[0094] This step corresponds to step 205. The explanations of the first ciphertext linear layer, the second ciphertext linear layer, the first ciphertext neural network model, the second ciphertext neural network model, the fragmented key, and the distributed key encryption protocol are detailed above and will not be repeated here.
[0095] In an alternative implementation, the data provider may also receive an inference program for the encrypted neural network model obtained by obfuscation using a secure compilation method. A detailed description of this inference program is provided above and will not be repeated here.
[0096] Step 303: Use the first ciphertext neural network model and the second ciphertext neural network model to infer the target data and obtain N ciphertext outputs.
[0097] Specifically, the data provider can use a first encrypted neural network model and a second encrypted neural network model to perform inference on the target data. Each of the N encrypted outputs corresponds one-to-one with a encrypted neural network model. In an optional implementation, the N encrypted outputs can be the outputs of the first encrypted linear layer and N-1 of the second encrypted linear layers.
[0098] The specific implementation of step 303 will be explained below with reference to the embodiments described above.
[0099] As mentioned earlier, the inference program used to perform the inference is obfuscated and encrypted using a secure compilation method. The key for the second encryption algorithm is written into this inference program. Therefore, since the neural network model is also encrypted using the second encryption algorithm, the inference program can use the key for the second encryption algorithm written into the inference program to decrypt both the first and second encrypted neural network models.
[0100] It's important to note that decryption is performed using memory-based decryption, ensuring the plaintext is not written to disk. This prevents the data provider from directly retrieving the decrypted result from a file stored on the hard drive. Content stored in memory is often difficult to access. Furthermore, the decryption key for the second encryption algorithm written in the inference program is also inaccessible to the data provider.
[0101] After decrypting the encrypted neural network model, the target data can be input into the decrypted layers of the N models to obtain intermediate results from the outputs of the first linear layer and the N-1 second linear layers.
[0102] Furthermore, the plaintext of the intermediate result can be input into N ciphertext linear layers. Each ciphertext linear layer uses a homomorphic multiplication operation between the ciphertext weight matrix W and the input X, and a homomorphic addition operation between the result of this operation and the bias vector b, to obtain the ciphertext output of that ciphertext linear layer. This results in N ciphertext outputs.
[0103] It should be noted that although step 303 describes inputting the intermediate results into the first encrypted neural network model and the second encrypted neural network model respectively, the data-side inference program cannot determine which is the first encrypted neural network model. That is, the data-side inference program can input the target data into N encrypted neural network models, and the N encrypted neural network models cannot distinguish which is the first encrypted neural network model.
[0104] Step 305: Use the sum of the N ciphertext outputs to obtain the reasoning result of the target data.
[0105] After obtaining N ciphertext outputs, the sum of these N ciphertext outputs can be obtained using the homomorphic addition operation. Since the plaintext parameter of the second ciphertext linear layer is 0, the plaintext corresponding to the ciphertext output of the second ciphertext linear layer is also 0. Accordingly, the plaintext corresponding to the sum of the N ciphertext outputs is actually equivalent to the output of the first linear layer. Furthermore, the inference result of the target data can be obtained using the sum of the N ciphertext outputs.
[0106] If there are other linear layers after the first linear layer, the specific implementation of step 305 can be as follows: the sum of the N ciphertext outputs is input into the subsequent layers of the N neural network models respectively to obtain the outputs of the N subsequent linear layers. Then, the sum of the N outputs is decrypted, and the decryption result is input into the nonlinear layer following the linear layer to obtain the inference result.
[0107] In this case, in each subsequent linear layer of the ciphertext model, at least the parameter of the last linear layer is 0 in plaintext. This ensures the accuracy of the inference results.
[0108] Furthermore, step 305 can also be implemented as follows: the inference program outputs the inference result of the target data based on the sum of the N ciphertext outputs. Specifically, the inference program performs the following steps: determining the sum of the N ciphertext outputs; and using the aggregation key corresponding to the N shard keys written in the inference program, decrypting the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0109] Specifically, after obtaining the sum of N ciphertext outputs, the aggregate key written in the inference program can be used to decrypt them. Then, the decrypted data is used to perform inference on subsequent layers of the first and second linear layers, yielding N plaintext outputs. The sum of these plaintext outputs is used as the inference result. Similar to the previous approach, in each subsequent linear layer of the ciphertext model, at least the plaintext parameter of the last linear layer is 0. This ensures the accuracy of the inference result.
[0110] This minimizes the number of time-consuming homomorphic operations and reduces the time spent on reasoning.
[0111] Furthermore, as mentioned earlier, the model can also obfuscate only certain layers. For example, the first linear layer is the last fully connected layer in the neural network model. The output of this fully connected layer can be input into an activation function used for classification, such as the softmax function. Here, the layer corresponding to the softmax function does not need to be obfuscated. In this way, the sum of the N ciphertext outputs can be directly decrypted, and the decrypted data can be input into a unique softmax function to obtain the final inference result.
[0112] The implementation of the Paillier distributed key will be explained next.
[0113] As mentioned earlier, the model needs to obfuscate the first linear layer of the neural network model to obtain m-1 second linear layers. It should be noted that the previous mention referred to the generation of N-1 second linear layers; since N will represent other data later, m here is actually the N from the previous context. The number of parameters in the second linear layers is the same as in the first linear layer, but the parameters in W and b of the second linear layers are all 0. The first linear layer and the N-1 second linear layers are then encrypted. The encryption process of the W matrix is used as an example to illustrate the encryption process of the first and second linear layers. It can be understood that the encryption process of the bias vector b is the same as the encryption process of the weight matrix W, and will not be repeated here. The encryption process includes the following steps:
[0114] 1. Key Generation: Given the security parameter π, the model first generates public and private keys for the Paillier encryption algorithm, where the private key SK p =(μ,λ), public key PK p = (N, g), where the order of the public key parameter g is a multiple of N, i.e., ord(g) = kN (if k = 1, g = 1 + N, g N = (1+N) N modN 2 =1,g kN modN 2 =1) Then select a large random integer γ that satisfies γ∈{0,1} π / 2 Given gcd(k,γ)=1, calculate h=g γ modN 2 The publicly disclosed system parameters are PP = <π,N,g,h>.
[0115] 2. Key splitting: splitting the public key into PK keys. p The parameter N is divided into m random numbers {n1, n2, ..., nn}. m},satisfy m is the total number of the first and second linear layers. Then, random numbers are selected for each model confusion task. As the task ID, a sharding key SK is generated for each linear layer. i = and aggregation key SK data =<λ,μ,γ>.
[0116] 3. Model Encryption: The first and second linear layers are encrypted using the m fragment keys obtained above and the public system parameters. Specifically, given the linear layer parameters w... i ∈Z N The model is encrypted to obtain:
[0117] Correspondingly, the data provider performs inference and obtains m ciphertext outputs [[y i By aggregating the m ciphertext outputs, we can obtain...
[0118] Furthermore, using the aggregation key SK data Decrypting [[y]] yields... Where L(x) = (x-1) / N.
[0119] The correctness of the above process will be explained next.
[0120] According to the encryption formula, the ciphertext reasoning result can be transformed as shown in formula (1):
[0121] Correspondingly, the decrypted result is shown in formula (2):
[0122] As can be seen, the plaintext corresponding to [[y]] is y. (i) The sum of.
[0123] The implementation of the MPCKKS distributed key will be explained next. This scheme is based on the CKKS distributed key, a fully homomorphic encryption algorithm.
[0124] Before using CKKS encryption, it is first necessary to define R. q =Z q [X] / (X N +1), a ring of residues of a polynomial. Simply put, R q It is a set of polynomials, where each element is a polynomial with coefficients in {0,1,…,q-1} and a degree not exceeding N.
[0125] The scheme described in this manual involves the following operations in the CKKS algorithm:
[0126] 1. KeyGen: Key generation algorithm, generates private key s←R q ;
[0127] 2. ECD: An encoding algorithm that takes a vector of length less than N as input and transforms it into R. q polynomials in;
[0128] 3. Enc: Encryption algorithm. Given a polynomial and a private key, it encrypts the polynomial into ciphertext. Specifically, for a polynomial m∈R... q The encryption algorithm first came from R q Randomly select polynomial a←R q , and noise polynomial e←R q (The coefficient of e is relatively small), and then calculate (c0,c1):=(-a·s+e+m,a)∈ And output;
[0129] 4. HomMulPt: Plaintext-Ciphertext Multiplication, takes a plaintext polynomial m′∈R as input. q And a secret calculate The new ciphertext encrypts m·m′;
[0130] 5. Dec: Decryption function, input ciphertext (c0, c1) ∈ R q Calculate c0,+c1·s;
[0131] 6. Dcd: Decoding function, takes a polynomial as input and outputs a vector (the inverse process of encoding function).
[0132] The reasoning process of the data side involves calculating matrix-vector multiplication. Matrix-vector multiplication can be achieved using coefficient encoding by calling the plaintext-ciphertext multiplication function HomMulPt once. For example, as shown in Figure 4, calculating two vectors... The inner product can be used to... These correspond sequentially to the coefficients of the polynomial, while the vector... By corresponding them in a certain reverse order, we can ensure that the constant term of the product of the two polynomials is the inner product of the two vectors.
[0133] Furthermore, it should be noted that this specification also includes improvements to CKKS. Specifically, during encryption, different ciphertexts share the same random polynomial a∈R. q Since distributed key encryption requires that the aggregated ciphertext of data encrypted with fragments can be decrypted with the aggregate key, CKKS can implement distributed key encryption by allowing different ciphertexts to share the same polynomial coefficients.
[0134] Encryption is performed by obfuscating the first linear layer (e.g., a fully connected layer). Regarding the parameters of the first and second linear layers, assuming there are two second linear layers and a total of 3 first linear layers, the encryption methods include:
[0135] The model encodes the parameter matrices of the three linear layers into three polynomials m0, m1, m2 ∈ R. q Then the model uses three fragment keys s0, s1, and s2 to encrypt the three polynomials respectively, resulting in three ciphertexts as shown in formula (3).
[0136] During the inference process, the data provider encodes its own sample vector (i.e., the target data) as a polynomial m′∈R. q Then the data provider calls HomMulPt to calculate the plaintext-ciphertext multiplication three times, and obtains the ciphertext output corresponding to the three linear layers, as shown in formula (4).
[0137]
[0138] The data provider performs a homomorphic addition operation on the three ciphertext outputs to obtain the aggregated result of the three ciphertext outputs, (d,a′)=(c0+c1+c2,a′). The data provider uses the aggregated key for decryption, where the aggregated key s=s0+s1+s2. The decryption process is shown in formula (5). Dec(s,(d,a′))=d+a′·s =(c0+c1+c2)+a′·(s0+s1+s2) =m′·(b0+a·s0)+m′·(b1+a·s1)+m′·(b2+a·s2) =m′·(m0+m1+m2)+m′·(e0+e1+e2) (5)
[0139] As can be seen, the data provider can directly decrypt m′·m, m=m0+m1+m2, where m′ is the sample vector and m is the sum of the parameter matrices of the three linear layers. These two polynomial multiplications already include the result of the inner product; therefore, the data provider can then call the decoding function to recover the result.
[0140] It should also be noted that the CKKS homomorphic encryption algorithm may contain noise, which can be eliminated by selecting appropriate parameters.
[0141] As for the correctness, it can be directly obtained through the above formula derivation.
[0142] This specification also provides a neural network model inference device, applied to the data side, as shown in Figure 5:
[0143] The receiving module 510 is used to receive a first encrypted neural network model with a first encrypted linear layer and N-1 second encrypted neural network models with second encrypted linear layers sent by the model provider; the first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model have the same structure; the first encrypted linear layer and the N-1 second encrypted linear layers are respectively encrypted using N fragment keys, and the N fragment keys are generated based on a distributed encryption protocol that supports homomorphic operations; the first encrypted linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameter of any second encrypted linear layer is 0;
[0144] The reasoning module 520 is used to reason about the target data using the first encrypted neural network model and the second encrypted neural network model respectively, to obtain N encrypted outputs; and to obtain the reasoning result of the target data by using the sum of the N encrypted outputs.
[0145] In one alternative embodiment, the first neural network model further includes a third layer, and the second neural network further includes a fourth layer, wherein the fourth layer is a layer with random numbers generated based on the parameters of the third layer.
[0146] In an optional embodiment, the receiving module 510 is further configured to receive an inference program for the encrypted neural network model obtained by obfuscation using a secure compilation method; wherein the inference of the target data to obtain the inference result is implemented based on the inference program.
[0147] In an optional implementation, the inference module 520 is specifically used to output the inference result of the target data based on the sum of the N ciphertext outputs; the inference program executes: determining the sum of the N ciphertext outputs; and using the aggregation key corresponding to the N shard keys written in the inference program to decrypt the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0148] In an optional embodiment, the first neural network model and the second neural network model are further encrypted based on a second encryption algorithm; the inference module 520 is further configured to: decrypt the first ciphertext neural network model and the second ciphertext neural network model using the key of the second encryption algorithm written in the inference program.
[0149] In one alternative implementation, the inference program invokes the unobfuscated Triton service framework during runtime; the inference program and the Triton service framework communicate via piped communication during runtime.
[0150] In one alternative implementation, the first linear layer is the last fully connected layer of the neural network model.
[0151] In one optional implementation, the distributed homomorphic encryption protocol is a Paillier-based distributed key protocol; or, the distributed homomorphic encryption protocol is a CKKS-based distributed key protocol, wherein the polynomial coefficients are the same when different fragment keys are encrypted.
[0152] As shown in Figure 6, this specification also provides a neural network model inference device, applied to a model provider with a neural network model, including:
[0153] The obfuscation module 610 is used to generate N-1 second neural network models with the same structure as the first neural network model for the first neural network model; the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0.
[0154] The encryption module 620 is used to obtain N fragment keys generated by the first encryption algorithm, and to encrypt the first linear layer and N-1 second linear layers using the N fragment keys respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;
[0155] The sending module 630 is used to send a first encrypted neural network model and N-1 second encrypted neural network models to the data party, so that the data party can use the first encrypted neural network model and N-1 second encrypted neural network models to perform inference.
[0156] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0157] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0158] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this specification does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0159] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.
[0160] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0161] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.
[0162] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0163] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0164] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0165] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0166] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0167] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0168] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0169] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0170] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.
Claims
1. A neural network model inference method, applied to the data side, comprising: The receiving model sends a first encrypted neural network model with a first encrypted linear layer and N-1 second encrypted neural network models with a second encrypted linear layer; The first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model have the same structure. The first encrypted linear layer and N-1 second encrypted linear layers are respectively encrypted using N fragment keys, which are generated based on a distributed encryption protocol that supports homomorphic operations. The first encrypted linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameter of any second encrypted linear layer is 0. The target data is inferred using the first ciphertext neural network model and the second ciphertext neural network model respectively, resulting in N ciphertext outputs; The inference result of the target data is obtained by summing the N ciphertext outputs.
2. The method according to claim 1, wherein the first neural network model further includes a third layer, and the second neural network further includes a fourth layer, wherein the fourth layer is a layer whose parameters are random numbers generated based on the third layer.
3. The method according to claim 1, further comprising: An inference program that receives the encrypted neural network model obtained by obfuscation using a secure compilation method; The reasoning of the target data to obtain the reasoning result is implemented based on the reasoning program.
4. The method according to claim 3, wherein obtaining the inference result of the target data by summing the N ciphertext outputs comprises: The reasoning program outputs the reasoning result of the target data based on the sum of the outputs of N ciphertexts; The inference procedure executes: Determine the sum of the N ciphertext outputs; Using the aggregation key corresponding to the N fragment keys written in the inference program, the sum of the N ciphertext outputs is decrypted to obtain the inference result of the target data.
5. The method according to claim 3, wherein the first neural network model and the second neural network model are further encrypted based on the second encryption algorithm; The inference procedure is also used to execute: The first ciphertext neural network model and the second ciphertext neural network model are decrypted using the key of the second encryption algorithm written in the inference program.
6. The method according to claim 4, wherein the inference program calls the unobfuscated Triton service framework during operation; the inference program and the Triton service framework communicate via pipe communication during operation.
7. The method according to claim 1, wherein the first linear layer is the last fully connected layer of the neural network model.
8. The method according to claim 1, wherein the distributed homomorphic encryption protocol is a Paillier-based distributed key protocol; Alternatively, the distributed homomorphic encryption protocol is a distributed key protocol implemented based on CKKS, wherein, The polynomial coefficients are the same when different fragment keys are used for encryption.
9. A neural network model inference method, applied to a model having a first neural network model, comprising: For the first neural network model, generate N-1 second neural network models with the same structure as the first neural network model; The second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0. Obtain N fragment keys generated using the first encryption algorithm, and use the N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations; A first encrypted neural network model and N-1 second encrypted neural network models are sent to the data party so that the data party can perform inference using the first encrypted neural network model and N-1 second encrypted neural network models.
10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-9.