Neural network model reasoning method and computer equipment
By using distributed encryption protocols and secure compilation technology that support homomorphic operations, the cryptographic neural network model is generated and decrypted, and the problem of inference of secure neural network model between the model and the data party is solved, and efficient and secure data and model protection is achieved.
Patent Information
- Application Number
- CN202510394460.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
On the premise of ensuring the security of data and model, how to reason about neural network models between the two parties, especially between the model party and the data party to safely transmit and calculate data.
By using a distributed encryption protocol that supports homomorphic operations, a ciphertext neural network model is generated, and the model is encrypted using shard keys. The data party performs post-inference aggregation and decryption, and combines secure compilation technology to ensure the security and efficiency of the inference process.
It realizes efficient inference of neural network models without leaking data and models, improves the security and computing efficiency of the model, and reduces the inference time.
Smart Images

Figure CN120278274A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of computer application technologies, and particularly relate to a neural network model inference method and a computer device. Background Art
[0002] With the development of machine learning technologies, neural network models have been applied in more and more fields. In some cases, there is a need for two parties to jointly use a neural network model for inference. For example, there is a model party and a data party. The model party has a neural network model, such as a neural network model for face recognition; the data party has data to be processed. The two parties need to complete the data inference and obtain the inference result while ensuring that the data or models they own are not leaked. Summary of the Invention
[0003] The purpose of this specification is to provide a neural network model inference method and a computer device.
[0004] The first aspect of this specification provides a neural network model inference method, which is applied to the data party and includes:
[0005] Receiving a first encrypted neural network model with a first encrypted linear layer and N - 1 second encrypted neural network models with second encrypted linear layers sent by the model party; the first neural network model corresponding to the first encrypted neural network model and the second neural network models corresponding to the second encrypted neural network models have the same structure, the first encrypted linear layer and the N - 1 second encrypted linear layers are respectively encrypted according to N sharding keys, and the N sharding keys are generated based on a distributed encryption protocol that supports homomorphic operations; the first encrypted linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameters of the second linear layer corresponding to any one of the second encrypted linear layers is 0;
[0006] Respectively using the first encrypted neural network model and the second encrypted neural network models to perform inference on target data to obtain N encrypted outputs;
[0007] Using the sum of the N encrypted outputs to obtain the inference result of the target data.
[0008] The second aspect of this specification provides a neural network model inference method, which is applied to the model party that owns the first neural network model and includes:
[0009] Generating N - 1 second neural network models with the same structure as the first neural network model for the first neural network model; the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the value of the second linear layer parameter is 0;
[0010] Obtain N shard keys generated using the first encryption algorithm, and use the N shard keys to encrypt the first linear layer and N - 1 second linear layers respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N - 1 second ciphertext neural network models with second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;
[0011] Send the first ciphertext neural network model and N - 1 second ciphertext neural network models to the data party, so that the data party uses the first ciphertext neural network model and N - 1 second ciphertext neural network models for inference.
[0012] A third aspect of this specification provides a neural network model inference device, which is applied to the data party and includes:
[0013] A receiving module, configured to receive the first ciphertext neural network model with a first ciphertext linear layer and N - 1 second ciphertext neural network models with second ciphertext linear layers sent by the model party; the first neural network model corresponding to the first ciphertext neural network model and the second neural network model corresponding to the second ciphertext neural network model have the same structure, the first ciphertext linear layer and N - 1 second ciphertext linear layers are respectively encrypted according to N shard keys, and the N shard keys are generated based on a distributed encryption protocol that supports homomorphic operations; the first ciphertext linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameters of the second linear layer corresponding to any one of the second ciphertext linear layers is 0;
[0014] An inference module, configured to respectively use the first ciphertext neural network model and the second ciphertext neural network model to perform inference on the target data to obtain N ciphertext outputs; and use the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0015] A fourth aspect of this specification provides a neural network model inference device, which is applied to the model party that owns the first neural network model and includes:
[0016] A confusion module, configured to generate N - 1 second neural network models with the same structure as the first neural network model for the first neural network model; the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the value of the second linear layer parameter is 0;
[0017] An encryption module, configured to obtain N shard keys generated by using a first encryption algorithm, and respectively encrypt the first linear layer and N - 1 second linear layers by using the N shard keys to obtain a first ciphertext neural network model with a first ciphertext linear layer and N - 1 second ciphertext neural network models with second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;
[0018] A sending module, configured to send the first ciphertext neural network model and N - 1 second ciphertext neural network models to a data party, so that the data party performs inference by using the first ciphertext neural network model and the N - 1 second ciphertext neural network models.
[0019] The fifth aspect of this specification provides a computer program product, including a computer program / instructions, which when executed by a processor, implements the steps of the neural network model inference method.
[0020] The sixth aspect of this specification provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the neural network model inference method.
[0021] The seventh aspect of this specification provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the neural network model inference method is implemented.
[0022] This specification provides a neural network model inference method. A data party receives N ciphertext neural network models sent by a model party. The N ciphertext neural network models are obtained by performing confusion processing on a first neural network model. The true neural network model includes a first linear layer, and N - 1 confused false neural network models include second linear layers, and the parameters of the second linear layers are 0. The first linear layer and N - 1 second linear layers are respectively encrypted based on N shard keys, and the N shard keys are generated based on a distributed encryption protocol that supports homomorphic operations. The data party uses the N ciphertext neural network models to perform inference on target data to obtain the sum of N ciphertext outputs, and further, the inference result of the target data can be obtained according to the sum of the N ciphertext outputs.
[0023] The first linear layer and N - 1 second linear layers are respectively encrypted by using N shard keys, and at the same time, the sum of the output results of the N linear layers can be decrypted by using an aggregation key corresponding to the N shard keys, which improves the security of the encrypted model. In addition, by generating false neural network models to confuse the model, the security of the model is further improved. Description of the Drawings
[0024] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1 is a schematic diagram of the system architecture in an embodiment of this specification;
[0026] Figure 2 is a flowchart of the neural network model inference method in an embodiment of this specification;
[0027] Figure 3 is a flowchart of the neural network model inference method in another embodiment of this specification;
[0028] Figure 4 is a schematic diagram of matrix-vector multiplication in an embodiment of this specification;
[0029] Figure 5 is a block diagram of the neural network model inference device in an embodiment of this specification;
[0030] Figure 6 is a block diagram of the neural network model inference device in another embodiment of this specification. Detailed implementation manners
[0031] To enable those skilled in the art of this technology to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0032] First, the architecture of the neural network model inference method provided in this specification will be described.
[0033] As Figure 1 shown, the method in this specification involves the model side and the data side. The model side has a trained neural network model. The data side has the data to be inferred. In the method of this specification, it is selected to deploy the model of the model side offline in ciphertext form to the data side, and the data side uses this model for inference. In this way, the model side cannot obtain the data of the data side, ensuring the security of the data of the data side. At the same time, the neural network model of the model side is encrypted and deployed to the data side, and the data side cannot obtain the plaintext of the parameters of the neural network model, ensuring the security of the model.
[0034] Among them, the neural network model can be any model including a linear layer, such as a Convolutional Neural Network (CNN), etc. The neural network model can be used for models for image recognition, natural language processing, speech recognition, etc. The present specification does not limit the uses of the neural network model. Optionally, the neural network model can be an image recognition model for face recognition, and the data owned by the data party is the data to be subjected to face recognition.
[0035] Next, it will be combined with Figure 1 the system architecture diagram shown in the present specification to give an overall description of the method shown in the present specification. When executed, for example, the model party first confuses the first neural network model to generate N - 1 false neural network models. The structure of the false neural network model is the same as that of the first neural network model, but the parameter values can be random numbers. For the sake of convenience of description, the false neural network model is called the second neural network model.
[0036] For the first neural network model, as Figure 1 shown, it can include a first linear layer. Correspondingly, the false linear layer corresponding to the first linear layer in the second neural network model can be called the second linear layer. For the second linear layer, different from other false layers in the second neural network model, the parameter values of this second linear layer can be 0.
[0037] For the first linear layer and the second linear layer, the model party can generate N shard keys based on a distributed encryption protocol supporting homomorphic operations, and the data encrypted with the N shard keys can support homomorphic operations. The shard keys have the following properties: The sum of the data encrypted with the N shard keys respectively can be decrypted using the aggregation key corresponding to the N shard keys. However, only using the aggregation key cannot decrypt the data encrypted with one shard key. Similarly, using any shard key cannot decrypt the sum of the data encrypted with the N shard keys respectively. Homomorphic operations can include homomorphic addition operations and homomorphic multiplication operations. Here, one kind of homomorphic operation involved in the present specification will be described. Suppose [[a]] represents the data obtained by encrypting a with a certain shard key, and b is the plaintext data. Then the homomorphic multiplication operation result of [[a]] and b is [[a*b]]. In addition, the homomorphic addition operation result of [[a]] and [[b]] is [[a + b]].
[0038] As Figure 1 shown, the model party uses the N shard keys to encrypt the first linear layer and N - 1 false linear layers respectively, obtaining the first encrypted linear layer and N - 1 second encrypted linear layers.
[0039] In this method, multiple second neural network models are used to obfuscate the first neural network model, which makes it impossible for the data party to determine which model is the first neural network model even if it obtains the ciphertext model. Moreover, the data party has no idea that the first ciphertext linear layer and the second ciphertext linear layer are encrypted with N shard keys. Then, if the data party wants to crack the ciphertext linear layer, it needs to crack N shard keys to obtain all the linear layers. It can be seen that encrypting with N shard keys can improve the security of the ciphertext neural network model.
[0040] Among them, the value of N can be determined according to the preset obfuscation degree. A high obfuscation degree indicates that the model party has higher requirements for confidentiality, so N can be set slightly higher. In addition, N cannot be infinitely high because increasing N will increase the number of homomorphic multiplication operations performed by the data party, and homomorphic multiplication operations are time-consuming. Therefore, the size of N can be reasonably determined according to the computational efficiency requirements of the data party and the obfuscation degree requirements of the model party.
[0041] In addition, as Figure 1 shown, the first neural network model may further include a third layer. Correspondingly, the corresponding layer in the second neural network model is called the fourth layer. The third layer and the first ciphertext linear layer can form the first ciphertext neural network model, and the second ciphertext linear layer and the fourth layer can form the second ciphertext neural network model. It should be noted that Figure 1 taking the third layer as an example, it can be understood that the first neural network model may further include a fifth layer, etc. The processing method for the fifth layer is the same as that for the third layer and will not be elaborated here. Here, the third layer is taken as an example of the layer before the first linear layer, and this embodiment does not represent a limitation to this specification.
[0042] Optionally, in order to ensure the security of other layers, the first neural network model and the second neural network model can also be further encrypted using a second encryption algorithm to ensure that all layers of the neural network model are in ciphertext and to ensure the security of the model. For example, the third layer and the fourth layer can be encrypted using the second encryption algorithm.
[0043] In addition, for the inference program of the model, secure compilation, a trusted execution environment (TEE), and fully homomorphic encryption inference can be used to ensure that the data party can perform inference on the model in a ciphertext state. This specification will use secure compilation as an example to illustrate how the data party can achieve ciphertext inference.
[0044] Secure compilation refers to ensuring the security of the compilation process through specific methods and technologies during the software development process to prevent malicious attacks and the introduction of potential vulnerabilities. Secure compilation aims to improve the reliability and security of software and prevent security defects from occurring during the construction, compilation, and deployment of applications.
[0045] As Figure 1 shown, the inference program can be encrypted through secure compilation, so that the data party cannot obtain the plaintext inference program and can only obtain the ciphertext inference program. At the same time, during the process of the data party running the inference program, the data generated by the inference program cannot be obtained by the data party. This ensures the encrypted state inference of the data party.
[0046] Furthermore, as Figure 1 shown, the model party can send the encrypted neural network model and the encrypted inference program after secure compilation to the data party.
[0047] The data party can use the inference program after secure compilation to perform inference. The specific operations and generated data during the inference process cannot be obtained by the data party. For the specific inference process, in the case of encryption with the second encryption algorithm, the first ciphertext neural network model can be decrypted first using the key of the second encryption algorithm written in the inference program.
[0048] To further ensure the data security of the neural network model, the neural network model can be decrypted in memory. Because if the decryption result is stored in the hard disk in the form of a file, then the data party can obtain the plaintext of the parameters of the neural network model by reading the file. Decrypting in memory can ensure that the plaintext does not fall to disk and ensure the security of the model.
[0049] For the specific inference process, as Figure 1 shown, the data party can input the target data into N ciphertext neural network models respectively, that is, the third layer of the first neural network model and the fourth layer of N - 1 second neural network models. Then, the outputs of the third layer and the fourth layer are input into the corresponding first ciphertext linear layer and N - 1 second ciphertext linear layers respectively. Inside the ciphertext linear layer, according to the homomorphic multiplication between the plaintext of the intermediate result and the ciphertext of the linear layer parameters, N ciphertext outputs corresponding to the first linear layer and the second linear layer are calculated respectively.
[0050] After obtaining the N ciphertext outputs, the N ciphertext outputs can be summed to obtain the ciphertext output. Since the parameter of the second ciphertext linear layer is 0, the plaintext corresponding to the ciphertext output obtained by the homomorphic multiplication of the parameter of the second ciphertext linear layer and the intermediate result is also 0. Furthermore, by adding the N ciphertext outputs, the ciphertext output corresponding to the first linear layer of the original neural network can be obtained. In addition, since the parameter of the second linear layer is 0, no matter what the parameters of the other layers before the second linear layer in the second neural network model are, they will not affect the output result of the second linear layer.
[0051] After obtaining the ciphertext output, if there is only one non-linear layer after the first linear layer, such as only one activation function, in this case, the first neural network model and the second neural network model mentioned above may not obfuscate this activation function, that is, the first neural network model and the second neural network model both correspond to the same activation function. Correspondingly, the ciphertext output can be directly decrypted and input into this activation function to obtain the inference result of the model.
[0052] In addition, if other layers of processing are required after the first linear layer to obtain the final output, similar to the above processing, other layers after the first linear layer can also not be obfuscated. Then the specific processing method is similar to the above, the ciphertext output can be decrypted and input into the unobfuscated layers to obtain the inference result. Of course, these layers can also be obfuscated, but the parameters are set to 0 (even if they are all 0, the encrypted values are different). Correspondingly, the ciphertext output can be decrypted, and the decryption results are respectively input into the subsequent layers of the first linear layer of the first neural network model and the subsequent layers of the second linear layer of N-1 second neural networks, and the results output by each layer are summed to obtain the final inference result. The above two examples do not represent a limitation to this specification.
[0053] In the above method, through the ciphertext deployment of the neural network model, the inference of the ciphertext inference program, and the decryption of the aggregation key, the confidentiality and non-stealability of the entire inference process of the model are guaranteed, protecting the security of the model. The above method also obfuscates the neural network model, increasing the difficulty for the data party to crack the neural network model.
[0054] At the same time, this method adopts an efficient secure compilation method, with less time consumption. Although homomorphic encryption is used, only one linear layer is encrypted using the shard key, and the homomorphic multiplication operation only needs to be calculated N times at least, which ensures the efficiency of the inference and can ensure that the end-to-end time consumption is less than 1.5 times that of the plaintext time consumption.
[0055] In addition, this method does not require modifying the model, only an additional N-1 second linear layers need to be generated.
[0056] Next, from the perspectives of the model party and the data party respectively, combined with Figure 2 and Figure 3 the flowcharts shown, the neural network model inference method provided in this specification will be described.
[0057] First, from the perspective of the model party that owns the neural network model, the method shown in this specification will be described. As Figure 2 shown, the following steps are included:
[0058] Step 201: Generate N - 1 second neural network models with the same structure as the first neural network model for the first neural network model.
[0059] Among them, the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0.
[0060] Specifically, first, for the first linear layer in the neural network model, generate N - 1 fake linear layers corresponding to this first linear layer. The fake linear layers are used for confusion. The size of the parameter matrix is the same as that of the first linear layer, but the parameter value is 0. The parameter value of 0 is to facilitate obtaining the ciphertext output that the original first linear layer should obtain by using the sum of the N ciphertext outputs later. In addition, fake layers for confusion are also generated for other layers of the first neural network model accordingly. These layers for confusion altogether constitute N - 1 second neural network models.
[0061] Among them, the linear layer refers to the layer that performs a linear transformation on the input, which can generally be expressed as y = Wx + b, where W is the weight matrix, b is the bias vector, x is the input, and y is the output. The reason for performing such processing on the linear layer in this specification is that the linear layer can perform linear operations. And the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations. Homomorphic operations generally only support linear operations and do not support non - linear operations.
[0062] Among them, the scope of confusion can be the entire neural network model or some layers of the neural network model. In other words, the first neural network model refers to the part of the neural network model that needs to be encrypted. For the latter case, some layers of the neural network model can be confused, and other layers are not confused. For example, if the layers before and after a certain layer are both confused, then the input of this layer can be the sum of the outputs of the N neural network models in the previous layer, and the output of this layer can be input into the subsequent N neural network models. How to ensure the correctness of the inference result in this case can be seen in the previous example and will not be elaborated here.
[0063] Among them, the first linear layer can be any linear layer in the neural network model, such as a fully - connected layer, a convolutional layer, and an embedding layer, etc. This specification does not limit the specific form of the first linear layer.
[0064] In an optional implementation, the first linear layer is the last fully connected layer of the neural network model. In the neural network model, the output of the last fully connected layer is generally input into the activation function, and the output of the activation function is the output of the neural network model. In this way, during the reasoning process, after the last fully connected layer, the sum of the N ciphertext outputs can be decrypted to obtain the output of the last fully connected layer, and the plaintext output is input into the activation function to obtain the output of the neural network model. This allows for more flexible reasoning.
[0065] It should be noted that in order to ensure the confidentiality of the inference results and the inference process, the above decryption process can be implemented through a confidential inference program, thereby ensuring that the inference process is unknown to the data party.
[0066] In addition, in this case, the parameters of the other layers of the second neural network model except the second linear layer can take any value. Since the parameter of the second linear layer is 0, no matter how the previous layer is calculated, the plaintext corresponding to the output of the second linear layer is also 0, which does not affect the correctness of the reasoning result. Of course, the parameters of the other layers of the second neural network model except the second linear layer can also be 0. Later, the second encryption algorithm can be used to encrypt the first neural network model and the second neural network model. Even if all parameters are 0, the encrypted values are different and will not leak privacy.
[0067] The values of the parameters of other layers of the second neural network model are described here by taking the third layer included in the first neural network model as an example. The fourth layer included in the second neural network model is generated based on the third layer, and the parameters of the fourth layer can be random numbers.
[0068] As described above, when the first linear layer is the last layer of the first neural network model, the fourth layer taking a random number will not affect the correctness of the inference result. In another case, the first neural network model may also include a fifth layer, and the sixth layer is the layer corresponding to the fifth layer in the second neural network model. The fifth layer may be the last layer of the first neural network model. In this way, the parameters of the sixth layer can be set to 0, so that the fourth layer parameters in the previous text are random numbers and will not affect the correctness of the inference result. The above two examples do not represent limitations on this specification.
[0069] Step 203: Obtain N sharding keys generated by the first encryption algorithm, and use the N sharding keys to encrypt the first linear layer and N-1 second linear layers respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N-1 second ciphertext neural network models with a second ciphertext linear layer.
[0070] Among them, the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations.
[0071] Specifically, first, N shard keys generated by a distributed encryption protocol based on homomorphic operation support and the aggregation key corresponding to the N shard keys are obtained. The features of the distributed encryption protocol are described in detail above and will not be elaborated here. The first linear layer and N - 1 fake second linear layers are encrypted respectively by the N shard keys to obtain the first encrypted linear layer and N - 1 second encrypted linear layers. Among them, different encrypted shard keys are used for different linear layers, and the N shard keys correspond one-to-one to the N linear layers.
[0072] In addition, although the values of the second linear layer are all 0, the data encrypted by different shard keys are different, and it is impossible to distinguish which is the original linear layer only from the first encrypted linear layer and the second encrypted linear layer. Although the first encrypted linear layer and the second encrypted linear layer are used here to distinguish the linear layers from two sources, for the data party, what it receives are N encrypted linear layers, and the N linear layers are not marked which is encrypted from the first linear layer and which is encrypted from the second linear layer, and the data party cannot distinguish the first encrypted linear layer and the second encrypted linear layer.
[0073] Among them, for the specific implementation method of the first encryption algorithm, it can be determined based on the distributed key generation protocol corresponding to homomorphic encryption. Two examples will be given here, and the two examples do not represent a limitation to this specification. One is the first encryption algorithm implemented based on semi - homomorphic encryption: the distributed key protocol based on Paillier. One is the first encryption algorithm implemented based on fully homomorphic encryption: the distributed key protocol implemented based on Cheon - Kim - Kim - Song (CKKS), where the polynomial coefficients corresponding to different shard key encryptions are the same.
[0074] The difference between fully homomorphic and semi - homomorphic lies in whether it supports: the homomorphic multiplication result of [[a]] and [[b]] is [[a*b]]. Both support the homomorphic multiplication operation of plaintext and ciphertext, and the homomorphic addition operation between ciphertext and ciphertext. That is, the homomorphic multiplication result of [[a]] and b is [[a*b]]. In addition, the homomorphic addition result of [[a]] and [[b]] is [[a + b]].
[0075] The specific implementations of these two encryption algorithms will be elaborated in the following text and will not be elaborated here for the time being.
[0076] In an alternative embodiment, in addition to obfuscating and encrypting the first linear layer, other layers of the first neural network model and the second neural network model can also be encrypted. In other words, the first neural network model and the second neural network model are also encrypted based on a second encryption algorithm. The second encryption algorithm can be any encryption algorithm. In an alternative embodiment, the second encryption algorithm can be a symmetric encryption algorithm, such as the Advanced Encryption Standard (AES).
[0077] Regarding the specific encryption scope of the second encryption algorithm, the first neural network model is used for illustration. It can be understood that the encryption model of the second neural network model is the same and will not be elaborated further.
[0078] In an alternative embodiment, it can be to encrypt other layers of the neural network model except the first linear layer. In this case, the aggregation key can be written into the encrypted inference program. Although the data party does not know that the first linear layer and the second linear layer are encrypted based on the distributed key, in order to ensure data security, by writing the aggregation key into the encrypted inference program, the data party cannot obtain the plaintext aggregation key. This can prevent the data party from performing homomorphic addition operations on the first linear layer and N - 1 second linear layers and decrypting the operation result through the aggregation key to obtain the parameters of the first linear layer.
[0079] In another alternative embodiment, it can be to encrypt all layers of the neural network model, that is, to encrypt other layers except the first linear layer once through the second encryption algorithm, and for the first linear layer and the second linear layer, first encrypt them through the sharding key and then through the second encryption algorithm. In this way, the aggregation key can be written into the encrypted inference program or directly sent to the data party in plaintext. Since the first linear layer and the second linear layer are encrypted twice for the encrypted neural network model, even if the data party obtains the plaintext of the aggregation key, it cannot decrypt the first linear layer.
[0080] In addition, in the above two cases, the decryption key of the second encryption algorithm can be written into the encrypted inference program to prevent the data party from obtaining the parameters of the plaintext neural network model.
[0081] In an alternative embodiment, as described above, the inference program can be processed so that the data party cannot obtain the plaintext inference program and the intermediate results of the inference program. Here, the security compilation algorithm is taken as an example to illustrate the specific processing process of the inference program.
[0082] The inference program can be obfuscated through secure compilation so that the inference program is in an unknown state to the data party. In addition, the inference program can be implemented based on the Triton Server framework. The core function of Triton Server is to achieve the efficient deployment and large-scale expansion of neural network inference services by uniformly managing models and optimizing resource scheduling, ensuring high-concurrency and low-latency inference performance.
[0083] In some cases, secure compilation can only process programs written in a specific programming language. For example, it can only process C++. However, the Triton Server framework is written in another programming language, such as Python. In this case, it is not convenient to directly perform secure compilation on the Triton Server framework. Therefore, the inference part of the neural network model can be separated and inference can be performed through an independent process. This process can be written in a specific programming language and can call the Triton Server framework written in another language. In other words, the inference program calls the unobfuscated Triton service framework during operation.
[0084] In this way, secure compilation can be performed on this to protect the security of the inference program (i.e., this process) and the keys written in the inference program.
[0085] In addition, since the inference program and the inference program are independent of the Triton Server framework, the two need to communicate. To ensure data security, the Triton service framework communicates through pipe communication during operation. Pipe communication is a communication method between processes. Through pipe communication, communication can be achieved through memory, ensuring that the communication content does not fall to disk and protecting the security of the intermediate results output by the inference program.
[0086] Through the above steps, an encrypted neural network model and an inference program of the encrypted neural network model obtained by obfuscation using the secure compilation method can be obtained.
[0087] Step 205: Send the first encrypted neural network model and N - 1 second encrypted neural network models to the data party so that the data party can perform inference using the first encrypted neural network model and the N - 1 second encrypted neural network models.
[0088] After obtaining the above first encrypted neural network model and second encrypted neural network model, the model can be sent to the data party for offline deployment of the encrypted neural network model. In this way, the data party can perform inference using the offline-deployed first encrypted neural network model and second encrypted neural network model.
[0089] In an alternative embodiment, an inference program of the encrypted neural network model obtained by performing obfuscation processing using a secure compilation method may also be sent to the data party; wherein, the inference of the target data to obtain an inference result is implemented based on the inference program. This can protect the security of the inference program.
[0090] Next, the model inference method will be described from the perspective of the data party in conjunction with Figure 3 , as follows. As Figure 3 shown, the method includes the following steps:
[0091] Step 301: Receive a first encrypted neural network model with a first encrypted linear layer and N - 1 second encrypted neural network models with second encrypted linear layers sent by the model party.
[0092] Wherein, the first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model have the same structure. The first encrypted linear layer and the N - 1 second encrypted linear layers are respectively encrypted according to N shard keys, and the N shard keys are generated based on a distributed encryption protocol supporting homomorphic operations; the first encrypted linear layer is obtained by encrypting a first linear layer included in the first neural network model, and the plaintext of the parameters of the second linear layer corresponding to any one of the second encrypted linear layers is 0.
[0093] This step corresponds to step 205. The descriptions of the first encrypted linear layer, the second encrypted linear layer, the first encrypted neural network model, the second encrypted neural network model, the shard key, and the distributed key encryption protocol, etc. can be found in the above text and will not be elaborated here.
[0094] In an alternative embodiment, the data party may also receive an inference program of the encrypted neural network model obtained by performing obfuscation processing using a secure compilation method. The description of this inference program can also be found in the previous text and will not be elaborated here.
[0095] Step 303: Respectively use the first encrypted neural network model and the second encrypted neural network model to perform inference on the target data to obtain N encrypted outputs.
[0096] Specifically, the data party may respectively use the first encrypted neural network model and the second encrypted neural network model to perform inference on the target data. Among them, the N encrypted outputs correspond one-to-one with the encrypted neural network models. In an alternative embodiment, the N encrypted outputs may be the outputs of the first encrypted linear layer and the N - 1 second encrypted linear layers respectively.
[0097] Next, the specific implementation manner of step 303 will be described in conjunction with the foregoing embodiments.
[0098] As described above, the inference program for performing inference has been obfuscated and encrypted by means of secure compilation. The key of the second encryption algorithm is written into the inference program. Therefore, on the basis that the neural network model is still encrypted based on the second encryption algorithm, the inference program can use the key of the second encryption algorithm written in the inference program to decrypt the first ciphertext neural network model and the second ciphertext neural network model.
[0099] It should be noted that during decryption, in-memory decryption is used and the plaintext does not fall to disk. This can prevent the data party from directly obtaining the decrypted result from the file stored on the hard disk. And the content stored in memory is often difficult to obtain. Moreover, the data party cannot obtain the decryption key of the second encryption algorithm written in the inference program.
[0100] After decrypting the encrypted neural network model, the target data can be input into the decrypted layers of the N models, and the intermediate results output before the first linear layer and the N - 1 second linear layers can be obtained respectively.
[0101] Furthermore, the plaintext of the intermediate results can be input into the N ciphertext linear layers. Each ciphertext linear layer uses the homomorphic multiplication operation of the ciphertext weight matrix W and the input X, and the homomorphic addition operation result of the operation result and the bias vector b to obtain the ciphertext output of the ciphertext linear layer. Furthermore, N ciphertext outputs can be obtained.
[0102] It should be noted that although step 303 describes inputting the intermediate results into the first ciphertext neural network model and the second ciphertext neural network model respectively, for the data party's inference program, it cannot determine which is the first ciphertext neural network model. That is to say, the data party's inference program can input the target data into the N ciphertext neural network models, and it is impossible to distinguish which is the first ciphertext neural network model among the N ciphertext neural network models.
[0103] Step 305, using the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0104] After obtaining the N ciphertext outputs, the sum of the N ciphertext outputs can be obtained by using the homomorphic addition operation result of the N ciphertext outputs. Since the plaintext of the parameters of the second ciphertext linear layer is 0, correspondingly, the plaintext corresponding to the ciphertext output by the second ciphertext linear layer is also 0. Correspondingly, the plaintext corresponding to the sum of the N ciphertext outputs is actually equivalent to the output of the first linear layer. Furthermore, using the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0105] In the case where there are other linear layers after the first linear layer, the specific implementation process of step 305 can be to input the sum of the N ciphertext outputs into the subsequent layers of N neural network models respectively, obtain the outputs of the N subsequent linear layers, then decrypt the sum of the N outputs, and input the decrypted result into the non-linear layer subsequent to this linear layer to obtain the inference result.
[0106] In this case, in the subsequent linear layer of each ciphertext model, at least the parameters of the last linear layer are plaintext 0. This can ensure the accuracy of the inference result.
[0107] In addition, the specific implementation process of step 305 can also be that the inference program outputs the inference result of the target data according to the sum of the N ciphertext outputs. Among them, the inference program performs: determining the sum of the N ciphertext outputs; using the aggregation key corresponding to the N shard keys written in the inference program to decrypt the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0108] Specifically, after obtaining the sum of the N ciphertext outputs, it can be decrypted using the aggregation key written in the inference program. Then, the decrypted data is used for inference in the subsequent layers of the first linear layer and the second linear layer to obtain N plaintext outputs, and the sum of the plaintext outputs is used as the inference result. Similar to the previous text, in the subsequent linear layer of each ciphertext model, at least the parameters of the last linear layer are plaintext 0. This can ensure the accuracy of the inference result.
[0109] In this way, the number of relatively time-consuming homomorphic operations can be minimized, and the time consumed by inference can be reduced.
[0110] In addition, as described above, the model party can also only obfuscate some layers. For example, the first linear layer is the last fully connected layer of the neural network model, and the output of this fully connected layer can be input into an activation function for classification, such as the softmax function. Here, the layer corresponding to the softmax function can not be obfuscated. In this way, the sum of the N ciphertext outputs can be directly decrypted, and the decrypted data can be input into the only softmax function to obtain the final inference result.
[0111] Next, the implementation method of the Paillier distributed key will be described.
[0112] As described above, the model party needs to obfuscate the first linear layer of the neural network model to obtain m - 1 second linear layers. It should be noted that it was mentioned above that N - 1 second linear layers were generated. Since N will be used to represent other data later, m here is actually N in the previous text. The number of parameters in the second linear layer is the same as that in the first linear layer, but the parameters in W and b of the second linear layer are all 0. Encrypt the first linear layer and N - 1 second linear layers. Here, taking the encryption process of the W matrix as an example, the encryption processes of the first linear layer and the second linear layer are described. It can be understood that the encryption process of the bias vector b is the same as that of the weight matrix W and will not be elaborated here. The encryption process includes the following steps:
[0113] 1. Key Generation: Given the security parameter π, the model party first generates the public and private keys of the Paillier encryption algorithm, where the private key SK p =(μ, λ), and the public key PK p =(N, g). The order of the public key parameter g is a multiple of N, i.e., ord(g)=kN (if k = 1, g = 1 + N, g N =(1 + N) N mod N 2 =1, g kN mod N 2 =1). Then, a large random integer γ is selected such that γ ∈ {0, 1} π / 2 and gcd(k, γ)=1, and h = g γ mod N 2 , and the public system parameters PP = <π, N, g, h> are made public.
[0114] 2. Key Splitting: Divide N in the public key PK p parameters into m random numbers {n1, n2,..., n m}, satisfying where m is the total number of the first linear layer and the second linear layer. Then, a random number is selected as the task ID for each model obfuscation task, and a fragment key is generated for each linear layer, as well as the aggregated key SK data =<λ, μ, γ>.
[0115] 3. Model Encryption: Encrypt the first linear layer and the second linear layer using the m fragment keys obtained above and the public system parameters. Specifically, given the linear layer parameter w i ∈Z N , the model is encrypted to obtain:
[0116] Correspondingly, the data party performs inference to obtain m ciphertext outputs [[y i, the aggregation of m ciphertext outputs can obtain
[0117] Furthermore, using the aggregation key SK data to decrypt [[y]], we can obtain where L(x) = (x - 1) / N.
[0118] Next, the correctness of the above process will be explained.
[0119] According to the encryption formula, the ciphertext inference result can be transformed as shown in formula (1):
[0120]
[0121] Correspondingly, the decrypted result is shown in formula (2):
[0122]
[0123] It can be seen that the plaintext corresponding to [[y]] is the sum of y (i) .
[0124] Next, the implementation method of the MPCKKS distributed key will be explained. This scheme is based on the CKKS distributed key implementation. CKKS is a fully homomorphic encryption algorithm.
[0125] Before encrypting using CKKS, it is first necessary to define R q = Z q [X] / (X N + 1), a residue class ring of polynomials. Simply put, R q is a set of polynomials, where each element is a polynomial with coefficients in {0, 1,..., q - 1} and degree not exceeding N.
[0126] The scheme in this specification involves the following operations in the CKKS algorithm:
[0127] 1. KeyGen: Key generation algorithm, generating a private key s ← R q ;
[0128] 2. Ecd: Encoding algorithm, inputting a vector with a length less than N, which converts the vector to a polynomial in R q ;
[0129] 3. Enc: Encryption algorithm, inputting a polynomial and a private key, encrypting the polynomial into a ciphertext. Specifically, for the polynomial m ∈ R q , the encryption algorithm first randomly selects a polynomial a ← R q from R q, and the noise polynomial e ← R q (the coefficients of e are relatively small), and then calculate and output;
[0130] 4. HomMulPt: Plaintext-ciphertext multiplication, input a plaintext polynomial m′ ∈ R q and a cipher Calculate The newly encrypted ciphertext is for m * m′;
[0131] 5. Dec: Decryption function, input the ciphertext (c0, c1) ∈ R q , calculate c0 + c1 · s;
[0132] 6. Dcd: Decoding function, input a polynomial and output a vector (the inverse process of the encoding function).
[0133] The reasoning process of the data party is to calculate matrix-vector multiplication, and matrix-vector multiplication can be achieved by using the coefficient encoding method and calling the plaintext-ciphertext multiplication HomMulPt once. For example, as Figure 4 shown, to calculate the inner product of two vectors , we can map to the coefficients of the polynomial in sequence, while the vector is mapped in a certain reverse order, so as to ensure that the constant term after multiplying the two polynomials is the inner product of the two vectors.
[0134] In addition, it should be noted that this specification also improves CKKS. Specifically, during encryption, let different ciphertexts share the same random polynomial a ∈ R q . Since the distributed key requires that the aggregated ciphertext of the data encrypted in fragments can be decrypted with the aggregated key, by letting different ciphertexts share the same polynomial coefficients, CKKS can implement distributed keys.
[0135] During encryption, confusion encryption is performed on the first linear layer (which can be a fully connected layer for example). For the parameters of the first linear layer and the second linear layer, assuming there are two second linear layers and the total number of the first linear layer and the first linear layer is 3, the encryption methods include:
[0136] The model party encodes the parameter matrices of the 3 linear layers into 3 polynomials m0, m1, m2 ∈ R q respectively. Then the model party encrypts these three polynomials with 3 shard keys s0, s1, s2 respectively to obtain three ciphertexts as shown in formula (3).
[0137]
[0138] During the inference process, the data party encodes its sample vector (i.e., the target data) into a polynomial m′ ∈ R q . Then the data party calls HomMulPt to calculate three plaintext-ciphertext multiplications to obtain the ciphertext outputs corresponding to the three linear layers, as shown in formula (4).
[0139]
[0140] The data party calculates the homomorphic addition operation of the three ciphertext outputs to obtain the aggregated result of the three ciphertext outputs, (d,a′) = (c0 + c1 + c2,a′). The data party decrypts using the aggregation key, where the aggregation key s = s0 + s1 + s2. The decryption process is as shown in formula (5).
[0141]
[0142] It can be seen that the data party can directly decrypt to obtain m′·m, m = m0 + m1 + m2, where m′ is the sample vector and m is the sum of the parameter matrices of the three linear layers. The results of these two polynomial multiplications contain the inner product result. Therefore, the data party can call the decoding function to recover the result.
[0143] In addition, it should be noted that the CKKS homomorphic encryption algorithm may contain noise, and the noise can be eliminated by selecting appropriate parameters.
[0144] Regarding the correctness, it can be directly obtained through the above formula derivation.
[0145] This specification also provides a neural network model inference device, which is applied to the data party, as Figure 5 including:
[0146] A receiving module 510, configured to receive a first ciphertext neural network model with a first ciphertext linear layer and N - 1 second ciphertext neural network models with second ciphertext linear layers sent by the model party; the first neural network model corresponding to the first ciphertext neural network model and the second neural network model corresponding to the second ciphertext neural network model have the same structure, the first ciphertext linear layer and the N - 1 second ciphertext linear layers are respectively encrypted according to N sharding keys, the N sharding keys are generated based on a distributed encryption protocol supporting homomorphic operations; the first ciphertext linear layer is obtained by encrypting the first linear layer included in the first neural network model, and the plaintext of the parameters of the second linear layer corresponding to any one of the second ciphertext linear layers is 0;
[0147] An inference module 520, configured to respectively perform inferences on the target data using the first ciphertext neural network model and the second ciphertext neural network models to obtain N ciphertext outputs; and obtain the inference result of the target data using the sum of the N ciphertext outputs.
[0148] In an alternative embodiment, the first neural network model further includes a third layer, and the second neural network further includes a fourth layer, and the fourth layer is a layer with random parameters generated based on the third layer.
[0149] In an alternative embodiment, the receiving module 510 is further configured to receive the inference program of the encrypted neural network model obtained by performing obfuscation processing using a secure compilation method; wherein, the inference of the target data to obtain the inference result is implemented based on the inference program.
[0150] In an alternative embodiment, the inference module 520 is specifically configured to output the inference result of the target data according to the sum of the N ciphertext outputs of the inference program; the inference program performs: determining the sum of the N ciphertext outputs; using the aggregation key corresponding to the N shard keys written in the inference program to decrypt the sum of the N ciphertext outputs to obtain the inference result of the target data.
[0151] In an alternative embodiment, the first neural network model and the second neural network model are further encrypted based on a second encryption algorithm; the inference module 520 is further configured to perform: using the key of the second encryption algorithm written in the inference program to decrypt the first encrypted neural network model and the second encrypted neural network model.
[0152] In an alternative embodiment, the inference program calls an unobfuscated triton service framework during operation; the inference program and the triton service framework communicate in a pipe communication manner during operation.
[0153] In an alternative embodiment, the first linear layer is the last fully connected layer of the neural network model.
[0154] In an alternative embodiment, the distributed homomorphic encryption protocol is a distributed key protocol based on Paillier; or, the distributed homomorphic encryption protocol is a distributed key protocol implemented based on CKKS, wherein the polynomial coefficients corresponding to different shard key encryptions are the same.
[0155] As Figure 6 shown, this specification further provides a neural network model inference device, which is applied to a model party having a neural network model, and includes:
[0156] An obfuscation module 610, configured to generate N - 1 second neural network models with the same structure as the first neural network model for the first neural network model; the second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0;
[0157] An encryption module 620, configured to obtain N shard keys generated by using a first encryption algorithm, and respectively encrypt the first linear layer and N - 1 second linear layers by using the N shard keys, to obtain a first encrypted neural network model with a first encrypted linear layer and N - 1 second encrypted neural network models with second encrypted linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol supporting homomorphic operations;
[0158] A sending module 630, configured to send the first encrypted neural network model and N - 1 second encrypted neural network models to a data party, so that the data party performs inference by using the first encrypted neural network model and N - 1 second encrypted neural network models.
[0159] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by a user's programming of the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain a hardware circuit that implements the logical method flow.
[0160] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0161] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this specification does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0162] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0163] For convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing one or more of this specification, the functions of each module may be implemented in the same or multiple software and / or hardware, or the modules implementing the same function may be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in electrical, mechanical or other forms.
[0164] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the function specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0165] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0167] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0168] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0169] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0170] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0171] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0172] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0173] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included within the scope of the claims.
Claims
1. A neural network model inference method, applied to a data party, includes: Receiving a first encrypted neural network model with a first encrypted linear layer and N - 1 second encrypted neural network models with second encrypted linear layers sent by a model party; The structures of the first neural network model corresponding to the first encrypted neural network model and the second neural network model corresponding to the second encrypted neural network model are the same. The first encrypted linear layer and the N - 1 second encrypted linear layers are respectively encrypted according to N shard keys, and the N shard keys are generated based on a distributed encryption protocol supporting homomorphic operations; the first encrypted linear layer is obtained by encrypting a first linear layer included in the first neural network model, and the plaintext of the parameters of the second linear layer corresponding to any one of the second encrypted linear layers is 0; Respectively using the first encrypted neural network model and the second encrypted neural network models to perform inference on target data to obtain N encrypted outputs; Using the sum of the N encrypted outputs to obtain the inference result of the target data.
2. The method according to claim 1, wherein the first neural network model further includes a third layer, and the second neural network further includes a fourth layer, and the fourth layer is a layer with random numbers as parameters generated based on the third layer.
3. The method according to claim 1, the method further includes: Receiving an inference program of the encrypted neural network model obtained by performing obfuscation processing using a secure compilation method; Wherein, the inference on the target data to obtain the inference result is implemented based on the inference program.
4. The method according to claim 3, the using the sum of the N encrypted outputs to obtain the inference result of the target data includes: The inference program outputs the inference result of the target data according to the sum of the N encrypted outputs; The inference program executes: Determining the sum of the N encrypted outputs; Using the aggregation key corresponding to the N shard keys written in the inference program to decrypt the sum of the N encrypted outputs to obtain the inference result of the target data.
5. The method according to claim 3, the first neural network model and the second neural network model are further encrypted based on a second encryption algorithm; The inference program is further used to execute: Using the key of the second encryption algorithm written in the inference program to decrypt the first encrypted neural network model and the second encrypted neural network models.
6. The method according to claim 4, the inference program calls an unobfuscated triton service framework during operation; the inference program and the triton service framework communicate through pipe communication during operation.
7. The method according to claim 1, the first linear layer is the last fully connected layer of the neural network model.
8. The method according to claim 1, the distributed homomorphic encryption protocol is a distributed key protocol based on Paillier; Or, the distributed homomorphic encryption protocol is a distributed key protocol implemented based on CKKS, where The polynomial coefficients corresponding to different shard keys during encryption are the same.
9. A neural network model inference method, applied to a model party having a first neural network model, includes: Generate N - 1 second neural network models with the same structure as the first neural network model for the first neural network model; The second linear layer of the second neural network model corresponds to the first linear layer of the first neural network model, and the parameter value of the second linear layer is 0; Obtain N shard keys generated by the first encryption algorithm, and use the N shard keys to encrypt the first linear layer and the N - 1 second linear layers respectively, to obtain a first ciphertext neural network model with a first ciphertext linear layer and N - 1 second ciphertext neural network models with second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol supporting homomorphic operations; Send the first ciphertext neural network model and the N - 1 second ciphertext neural network models to the data party, so that the data party uses the first ciphertext neural network model and the N - 1 second ciphertext neural network models for inference.
10. A computing device, including a memory and a processor, where the memory stores executable code, and when the processor executes the executable code, it implements the method according to any one of claims 1 - 9.