Neural network model inference method and computer device

WO2026199920A1PCT designated stage Publication Date: 2026-10-01ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/131198
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2025-10-30
Publication Date
2026-10-01

Smart Images

  • Figure CN2025131198_01102026_PF_FP_ABST
    Figure CN2025131198_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A neural network model inference method and a computer device. A data party first receives an encrypted neural network model sent by a model party. A first linear layer of the encrypted neural network model corresponds to a first ciphertext linear layer and N-1 second ciphertext linear layers, the first ciphertext linear layer being obtained by directly encrypting the first linear layer, and the second ciphertext linear layer being obtained by encrypting a linear layer with all zero parameters. The first ciphertext linear layer and the second ciphertext linear layer are respectively encrypted using N key shares, the N key shares being generated on the basis of a distributed encryption protocol supporting homomorphic operations. The data party uses the encrypted neural network model to perform target data inference to obtain an intermediate result output by another layer preceding the first linear layer, and inputs the intermediate result into the first linear layer and the N-1 second linear layers, respectively, to obtain N outputs respectively corresponding to N ciphertext linear layers. An inference result of the target data is obtained according to a sum of the N outputs.
Need to check novelty before this filing date? Find Prior Art

Description

A neural network model inference method and computer device

[0001] This application claims priority to Chinese Patent Application No. 2025103949308, filed on March 28, 2025, entitled "A Neural Network Model Reasoning Method and Computer Device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The embodiments in this specification belong to the field of computer application technology, and in particular relate to a neural network model inference method and computer equipment. Background Technology

[0003] With the development of machine learning technology, neural network models are being applied in an increasing number of fields. In some cases, there is a need for two parties to jointly utilize neural network models for inference. For example, there is a model provider and a data provider. The model provider possesses a neural network model, such as one used for facial recognition; the data provider possesses the data to be processed. Both parties need to complete the data inference and obtain the inference result while ensuring that their respective data or models are not leaked. Summary of the Invention

[0004] The purpose of this specification is to provide a neural network model inference method and computer equipment.

[0005] The first aspect of this specification provides a neural network model inference method, applied to the data side, including:

[0006] The receiver sends an encrypted neural network model; the encrypted neural network model includes a first ciphertext linear layer and N-1 second ciphertext linear layers, the first ciphertext linear layer and the N-1 second ciphertext linear layers are respectively encrypted using N fragmentation keys, the N fragmentation keys are generated based on a distributed encryption protocol that supports homomorphic operations; the first ciphertext linear layer is obtained by encrypting the first linear layer included in the neural network model, and the plaintext of the parameters of any second ciphertext linear layer is 0;

[0007] The encrypted neural network model is used to reason about the target data to obtain intermediate results of the first ciphertext linear layer and the second ciphertext linear layer to be input.

[0008] The intermediate results are input into the first ciphertext linear layer and N-1 second ciphertext linear layers respectively to obtain N ciphertext outputs;

[0009] The inference result of the target data is obtained by summing the N ciphertext outputs.

[0010] The second aspect of this specification provides a neural network model inference method, applicable to model providers that possess neural network models, including:

[0011] Generate a second linear layer with N-1 parameters set to 0, corresponding to the first linear layer;

[0012] Obtain N fragment keys generated using the first encryption algorithm, and use the N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively, to obtain the first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;

[0013] An encrypted neural network model is sent to the data party so that the data party can use the encrypted neural network model for inference.

[0014] A third aspect of this specification provides a neural network model inference apparatus for use on a data side, comprising:

[0015] A receiving module is used to receive an encrypted neural network model sent by the model provider. The encrypted neural network model includes a first ciphertext linear layer and N-1 second ciphertext linear layers. The first ciphertext linear layer and the N-1 second ciphertext linear layers are respectively encrypted using N fragmentation keys, which are generated based on a distributed encryption protocol that supports homomorphic operations. The first ciphertext linear layer is obtained by encrypting the first linear layer included in the neural network model, and the plaintext of the parameters of any second ciphertext linear layer is 0.

[0016] The first inference module is used to infer the target data using the encrypted neural network model to obtain intermediate results of the first ciphertext linear layer and the second ciphertext linear layer to be input.

[0017] The second inference module is used to input the intermediate results into the first ciphertext linear layer and N-1 second ciphertext linear layers respectively, and obtain N ciphertext outputs;

[0018] The output module is used to obtain the reasoning result of the target data by using the sum of N ciphertext outputs.

[0019] This specification provides a fourth aspect of a neural network model inference apparatus, applied to a model provider possessing a neural network model, comprising:

[0020] The obfuscation module is used to generate a second linear layer with N-1 parameters set to 0, corresponding to the first linear layer.

[0021] The encryption module is used to obtain N fragment keys generated by the first encryption algorithm, and to encrypt the first linear layer and N-1 second linear layers using the N fragment keys respectively, to obtain the first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;

[0022] The sending module is used to send an encrypted neural network model to the data party so that the data party can use the encrypted neural network model for inference.

[0023] The fifth aspect of this specification provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the neural network model inference method described above.

[0024] A sixth aspect of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the aforementioned neural network model inference method.

[0025] A seventh aspect of this specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the above-described neural network model inference method.

[0026] This specification provides a method for inference using a neural network model. The data provider first receives an encrypted neural network model from the model provider. This encrypted neural network model has a first ciphertext linear layer and N-1 second ciphertext linear layers corresponding to its first linear layer. The first ciphertext linear layer is obtained by directly encrypting the first linear layer, and the second ciphertext linear layers are obtained by encrypting linear layers with all parameters set to 0. The first and second ciphertext linear layers are each encrypted using N sharding keys, which are generated based on a distributed encryption protocol supporting homomorphic operations. The data provider uses this encrypted neural network model to infer the target data, obtaining intermediate results from the outputs of other layers before the first linear layer. These intermediate results are then input into the first linear layer and the N-1 second linear layers, respectively, yielding N outputs corresponding to the N ciphertext linear layers. The sum of these N outputs yields the inference result for the target data.

[0027] The encryption model improves security by using N fragment keys to encrypt the true and false linear layers separately, and the sum of the outputs of the true and false linear layers can be decrypted using the aggregate key corresponding to the N fragment keys. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 is a schematic diagram of the system architecture in one embodiment of this specification;

[0030] Figure 2 is a flowchart of a neural network model inference method in one embodiment of this specification;

[0031] Figure 3 is a flowchart of a neural network model inference method in another embodiment of this specification;

[0032] Figure 4 is a schematic diagram of matrix-vector multiplication in one embodiment of this specification;

[0033] Figure 5 is a block diagram of a neural network model inference device in one embodiment of this specification;

[0034] Figure 6 is a block diagram of a neural network model inference device according to another embodiment of this specification. Detailed Implementation

[0035] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0036] First, the architecture of the neural network model inference method provided in this manual will be explained.

[0037] As shown in Figure 1, the method in this specification involves a model side and a data side. The model side possesses a trained neural network model. The data side possesses the data to be used for inference. In the method described in this specification, the model from the model side is deployed offline to the data side in encrypted form, and the data side uses this model for inference. In this way, the model side cannot obtain the data from the data side, ensuring the security of the data from the data side. At the same time, the neural network model from the model side is encrypted before being deployed to the data side, and the data side cannot obtain the plaintext parameters of the neural network model, ensuring the security of the model.

[0038] The neural network model can be any model that includes linear layers, such as Convolutional Neural Networks (CNNs). Neural network models can be used for image recognition, natural language processing, speech recognition, etc., and this specification does not limit the application of the neural network model. Optionally, the neural network model can be an image recognition model for face recognition, where the data provider possesses the face data to be recognized.

[0039] The following section will provide a general explanation of the method described in this specification, referring to the system architecture diagram in Figure 1. During execution, as shown in Figure 1, the model first obfuscates the first linear layer of the neural network model, generating N-1 fake linear layers. The parameter structure of the fake linear layers is the same as that of the first linear layer, and the parameter values ​​are 0. For ease of description, the fake linear layers targeting the first linear layer will be referred to as the second linear layer.

[0040] The model then generates N fragment keys based on a distributed encryption protocol that supports homomorphic operations. Data encrypted with N fragment keys can support homomorphic operations. Fragment keys have the following property: the sum of data encrypted by N fragment keys can be decrypted using the aggregate key corresponding to the N fragment keys. However, data encrypted with a single fragment key cannot be decrypted using only the aggregate key; similarly, the sum of data encrypted with N fragment keys cannot be decrypted using any fragment key. Homomorphic operations can include homomorphic addition and homomorphic multiplication. This specification will describe one type of homomorphic operation. Assuming [[a]] represents data encrypted with a certain fragment key, and b is plaintext data, then the homomorphic multiplication of [[a]] and b results in [[a*b]]. Furthermore, the homomorphic addition of [[a]] and [[b]] ​​results in [[a+b]].

[0041] As shown in Figure 1, the model uses N fragment keys to encrypt the first linear layer and N-1 fake linear layers respectively, to obtain the first ciphertext linear layer and N-1 second ciphertext linear layers, thus obtaining the encrypted neural network model.

[0042] This neural network model uses a second linear layer to obfuscate the first linear layer, making it impossible for the data provider to identify which ciphertext linear layer is the first. Furthermore, the data provider cannot know that the first and second ciphertext linear layers are encrypted using N fragment keys. Therefore, if the data provider wants to crack the ciphertext linear layers, it would need to crack all N fragment keys to obtain the entire linear layer. It is evident that encryption using N fragment keys can improve the security of the ciphertext neural network model.

[0043] The value of N can be determined based on the preset level of confusion. A higher level of confusion indicates a higher requirement for confidentiality on the model side, so N can be set slightly higher. However, N cannot be set indefinitely high, as increasing N will increase the number of homomorphic multiplication operations performed by the data side, which are time-consuming. Therefore, the size of N should be reasonably determined based on the computational efficiency requirements of the data side and the level of confusion required by the model side.

[0044] Optionally, to ensure the security of other layers, a second encryption algorithm can be used to further encrypt the neural network model, ensuring that each layer of the neural network model is encrypted and thus guaranteeing the model's security.

[0045] Furthermore, for the model's inference program, secure compilation, a trusted execution environment (TEE), and fully encrypted inference can be used to ensure that the data party performs inference on the model in encrypted form. This specification will use secure compilation as an example to illustrate how the data party implements encrypted inference.

[0046] Secure compilation refers to the use of specific methods and techniques during software development to ensure the security of the compilation process, preventing malicious attacks and the introduction of potential vulnerabilities. Secure compilation aims to improve the reliability and security of software, preventing security flaws from occurring during the building, compilation, and deployment of applications.

[0047] As shown in Figure 1, the inference program can be encrypted through secure compilation, ensuring that the data provider can only obtain the ciphertext version of the inference program, not the plaintext version. Furthermore, during the execution of the inference program by the data provider, the data generated by the inference program cannot be accessed by the data provider. This ensures the secure inference capability of the data provider.

[0048] As shown in Figure 1, the model provider can send an encrypted neural network model and a securely compiled encrypted inference program to the data provider.

[0049] The data provider can perform inference using a securely compiled inference program. The specific operations of the inference process and the generated data cannot be accessed by the data provider. Regarding the specific inference process, if a second encryption algorithm is used, the neural network model can be decrypted using the key of the second encryption algorithm written into the inference program.

[0050] To further ensure the data security of neural network models, they can be decrypted in memory. If the decryption result is stored as a file on the hard drive, the data provider can obtain the plaintext parameters of the neural network model by reading the file. Decryption in memory ensures that the plaintext is not written to disk, thus guaranteeing model security.

[0051] As shown in Figure 1, the data provider can input the target data into the decrypted plaintext model to obtain intermediate results from the outputs of other layers before the first linear layer. These intermediate results are then input into the first ciphertext linear layer and N-1 second ciphertext linear layers. Within each ciphertext linear layer, based on the homomorphic multiplication between the plaintext of the intermediate results and the ciphertext of the linear layer parameters, N ciphertext outputs corresponding to the first and second linear layers are calculated. As shown in Figure 1, these N ciphertext outputs can be summed to obtain the ciphertext output, from which the inference result can be derived.

[0052] Since the parameters of the second ciphertext linear layer are 0, the plaintext corresponding to the ciphertext output obtained by the homomorphic multiplication of the parameters of the second ciphertext linear layer and the intermediate results is also 0. Therefore, by adding the N ciphertext outputs, we can obtain the ciphertext output corresponding to the first linear layer of the original neural network.

[0053] Regarding specific methods for obtaining inference results from ciphertext output, for example, an encrypted inference program can use the aggregate key corresponding to N fragment keys to decrypt the ciphertext output and input the decrypted data into subsequent layers for further inference. If the subsequent layers are also linear layers, then the ciphertext output can be used directly for inference. This specification does not impose any limitations on this.

[0054] The above method ensures the confidentiality and non-theft of the entire inference process of the model by using encrypted deployment of the neural network model, inference of the encrypted inference program, and decryption of the aggregate key, thus protecting the security of the model.

[0055] Meanwhile, this method employs an efficient and secure compilation approach, resulting in minimal computation time. Although homomorphic encryption is used, it only encrypts a single linear layer using a fragmented key, requiring a minimum of N homomorphic multiplication operations. This ensures efficient inference, guaranteeing that end-to-end processing time is less than 1.5 times that of plaintext processing.

[0056] Furthermore, this method requires no modification to the model, only the generation of an additional N-1 second linear layers.

[0057] The following sections will explain the neural network model inference method provided in this specification from the perspectives of both the model and data sides, using the flowcharts shown in Figures 2 and 3.

[0058] First, the method described in this specification will be explained from the perspective of the model owner who has a neural network model. As shown in Figure 2, it includes the following steps:

[0059] Step 201: Generate a second linear layer with N-1 parameters set to 0, corresponding to the first linear layer.

[0060] Specifically, for the first linear layer in the neural network model, N-1 pseudo-linear layers are generated corresponding to it. These pseudo-linear layers are used for obfuscation; their parameter matrices have the same size as the first linear layer, but the parameter values ​​are all zero. The parameter values ​​are zero to facilitate obtaining the original ciphertext output from the first linear layer later by summing the N ciphertext outputs.

[0061] In this context, a linear layer refers to a layer that performs a linear transformation on the input, typically represented as y = Wx + b, where W is the weight matrix, b is the bias vector, x is the input, and y is the output. This specification treats linear layers in this way because they can perform linear operations. However, the first encryption algorithm is based on a distributed encryption protocol that supports homomorphic operations, which generally only support linear operations and not nonlinear operations.

[0062] The first linear layer can be any linear layer in the neural network model, such as a fully connected layer, a convolutional layer, or an embedding layer. This specification does not limit the specific form of the first linear layer.

[0063] In one optional implementation, the first linear layer is the last fully connected layer of the neural network model. In the neural network model, the output of the last fully connected layer is typically input into an activation function, and the output of the activation function is the output of the neural network model. Thus, during inference, after the last fully connected layer, the sum of the N ciphertext outputs can be decrypted to obtain the output of the last fully connected layer, and this plaintext output can then be input into the activation function to obtain the output of the neural network model. This allows for more flexible inference.

[0064] It should be noted that, in order to ensure the confidentiality of the reasoning result and the reasoning process, the above decryption process can be implemented through a confidential reasoning program, thereby ensuring that the reasoning process is unknown to the data party.

[0065] Step 203: Obtain N fragment keys generated using the first encryption algorithm, and use the N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively to obtain the first ciphertext linear layer and N-1 second ciphertext linear layers.

[0066] The first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations.

[0067] Specifically, firstly, N fragment keys generated using a distributed encryption protocol supporting homomorphic operations, and the corresponding aggregate key for each of the N fragment keys, were obtained. The characteristics of the distributed encryption protocol are detailed above and will not be repeated here. The first linear layer and N-1 fake second linear layers are then encrypted using the N fragment keys, resulting in the first encrypted linear layer and N-1 second encrypted linear layers. Different linear layers use different encryption fragment keys, and each of the N fragment keys corresponds one-to-one with one of the N linear layers.

[0068] Furthermore, although the values ​​of the second linear layer are all 0, the data encrypted using different fragmentation keys are different. It's impossible to distinguish which is the original linear layer simply by looking at the first and second ciphertext linear layers. While the first and second ciphertext linear layers are used to differentiate the two sources, the data sender receives N ciphertext linear layers. These N linear layers do not indicate which was encrypted by the first linear layer and which by the second; therefore, the data sender cannot distinguish between the first and second ciphertext linear layers.

[0069] The specific implementation method of the first encryption algorithm can be determined based on the distributed key generation protocol corresponding to homomorphic encryption. Two examples will be given here, but these examples are not intended to limit this specification. One is a first encryption algorithm based on semi-homomorphic encryption: a distributed key protocol based on Paillier. The other is a first encryption algorithm based on fully homomorphic encryption: a distributed key protocol based on Cheon-Kim-Kim-Song (CKKS), where the polynomial coefficients are the same when different fragment keys are encrypted.

[0070] The difference between fully homomorphic and semi-homomorphic operations lies in whether they support the following: the homomorphic multiplication of [[a]] and [[b]] ​​results in [[a*b]]. Both support homomorphic multiplication of plaintext and ciphertext, as well as homomorphic addition between ciphertexts. That is, the homomorphic multiplication of [[a]] and b results in [[a*b]]. Furthermore, the homomorphic addition of [[a]] and [[b]] ​​results in [[a+b]].

[0071] The specific implementations of these two encryption algorithms will be detailed later, and will not be elaborated here.

[0072] In an optional implementation, in addition to obfuscating and encrypting the first linear layer, other layers of the neural network model can also be encrypted. In other words, the neural network model is also encrypted based on a second encryption algorithm. The second encryption algorithm can be any encryption algorithm; in an optional implementation, the second encryption algorithm can be a symmetric encryption algorithm, such as the Advanced Encryption Standard (AES).

[0073] Regarding the specific encryption scope of the second encryption algorithm, in one optional implementation, it can be used to encrypt all layers of the neural network model except for the first linear layer. In this case, the aggregation key can be written into the encrypted inference program. Although the data provider is unaware that the first and second linear layers are encrypted using a distributed key, to ensure data security, writing the aggregation key into the encrypted inference program prevents the data provider from obtaining the plaintext aggregation key. This prevents the data provider from performing homomorphic addition operations on the first linear layer and N-1 second linear layers, and then decrypting the result using the aggregation key to obtain the parameters of the first linear layer.

[0074] In another alternative implementation, all layers of the neural network model can be encrypted. Specifically, all layers except the first linear layer can be encrypted once using the second encryption algorithm. The first and second linear layers are first encrypted using a fragmentation key, and then further encrypted using the second encryption algorithm. In this way, the aggregate key can be written into the encrypted inference program or sent directly to the data provider in plaintext. Since the first and second linear layers are encrypted twice for the encrypted neural network model, even if the data provider obtains the plaintext of the aggregate key, it cannot decrypt the first linear layer.

[0075] Furthermore, in both of the above cases, the decryption key of the second encryption algorithm can be written into the encrypted inference program to prevent the data party from obtaining the parameters of the plaintext neural network model.

[0076] In an alternative implementation, as described above, the inference program can be processed to prevent the data provider from obtaining the plaintext inference program and its intermediate results. The specific processing procedure of the inference program will be explained here using a secure compilation algorithm as an example.

[0077] The inference program can be obfuscated through secure compilation, making it unknown to the data provider. Furthermore, the inference program can be implemented based on the Triton Server framework. The core function of Triton Server is to achieve efficient deployment and scalable expansion of neural network inference services through unified model management and optimized resource scheduling, ensuring high-concurrency, low-latency inference performance.

[0078] In some cases, secure compilation can only process programs written in a specific programming language, such as C++, while the Triton server framework is written in another programming language, such as Python. In such cases, it's inconvenient to directly perform secure compilation on the Triton server framework. Therefore, the inference part of the neural network model can be separated and performed in a separate process. This process can be written in a specific programming language and can call the Triton server framework written in another language. In other words, the inference program calls the unobfuscated Triton service framework during runtime.

[0079] This allows for secure compilation of the code to protect the inference program (i.e., the process) and the keys written into it.

[0080] Furthermore, since the inference program and the inference program itself are independent of the Triton server framework, they need to communicate. To ensure data security, the Triton service framework communicates via pipes during operation. Pipe communication is a method of inter-process communication that allows communication to be implemented through memory, ensuring that the communication content is not written to disk and guaranteeing the security of the intermediate results output by the inference program.

[0081] Through the above steps, an encrypted neural network model and an inference program for the encrypted neural network model obtained by obfuscation using secure compilation methods can be obtained.

[0082] Step 205: Send the encrypted neural network model to the data party so that the data party can use the encrypted neural network model for inference.

[0083] After obtaining the encrypted neural network model, the model can be sent to the data provider for offline deployment. This allows the data provider to perform inference using the offline-deployed neural network model.

[0084] In an alternative implementation, an inference program for the encrypted neural network model, obtained by obfuscation using a secure compilation method, may also be sent to the data provider; wherein inference of the target data to obtain the inference result is implemented based on the inference program. This protects the security of the inference program.

[0085] The following section, using Figure 3 as an example, will explain the model inference method from the data provider's perspective. As shown in Figure 3, the method includes the following steps:

[0086] Step 301: Receive the encrypted neural network model sent by the model provider.

[0087] The encrypted neural network model includes a first ciphertext linear layer and N-1 second ciphertext linear layers. The first ciphertext linear layer and the N-1 second ciphertext linear layers are respectively encrypted using N fragment keys, which are generated based on a distributed encryption protocol that supports homomorphic operations. The first ciphertext linear layer is obtained by encrypting the first linear layer included in the neural network model, and the plaintext of the parameters of any second ciphertext linear layer is 0.

[0088] This step corresponds to step 205. The explanations of the first ciphertext linear layer, the second ciphertext linear layer, the fragmentation key, and the distributed key encryption protocol are detailed above and will not be repeated here.

[0089] In an alternative implementation, the data provider may also receive an inference program for the encrypted neural network model obtained by obfuscation using a secure compilation method. A detailed description of this inference program is provided above and will not be repeated here.

[0090] Step 303: Use the encrypted neural network model to reason about the target data to obtain intermediate results of the first ciphertext linear layer and the second ciphertext linear layer to be input.

[0091] Specifically, the data provider can use the layers before the first linear layer of the neural network model to perform inference and obtain intermediate results, which are the inputs that were originally intended to be input to the first linear layer.

[0092] The specific implementation of step 303 will be explained below with reference to the embodiments described above.

[0093] As mentioned earlier, the inference program used to perform the inference is obfuscated and encrypted using a secure compilation method. The key for a second encryption algorithm is written into this inference program. Therefore, since the neural network model is also encrypted using the second encryption algorithm, the inference program can use the key for the second encryption algorithm written into the inference program to decrypt the encrypted neural network model.

[0094] It's important to note that decryption is performed using memory-based decryption, ensuring the plaintext is not written to disk. This prevents the data provider from directly retrieving the decrypted result from a file stored on the hard drive. Content stored in memory is often difficult to access. Furthermore, the decryption key for the second encryption algorithm written in the inference program is also inaccessible to the data provider.

[0095] After decrypting the encrypted neural network model, the target data can be input into the decrypted layer to obtain the intermediate results of the layer output before the first linear layer.

[0096] Step 305: Input the intermediate results into the first ciphertext linear layer and N-1 second ciphertext linear layers respectively to obtain N ciphertext outputs.

[0097] Specifically, the plaintext of the intermediate result can be input into N ciphertext linear layers. Each ciphertext linear layer uses a homomorphic multiplication operation between the ciphertext weight matrix W and the input X, and a homomorphic addition operation between the result of this operation and the bias vector b, to obtain the ciphertext output of that ciphertext linear layer. Thus, N ciphertext outputs can be obtained.

[0098] It should be noted that although step 305 describes inputting the intermediate results into the first and second ciphertext linear layers respectively, the data-side inference program cannot determine which is the first ciphertext linear layer. That is, the data-side inference program can input the intermediate results into N ciphertext linear layers, and it is impossible to distinguish which of the N ciphertext linear layers is the first ciphertext linear layer.

[0099] Step 307: Use the sum of the N ciphertext outputs to obtain the reasoning result of the target data.

[0100] After obtaining N ciphertext outputs, the sum of these N ciphertext outputs can be obtained using the homomorphic addition operation. Since the plaintext parameter of the second ciphertext linear layer is 0, the plaintext corresponding to the ciphertext output of the second ciphertext linear layer is also 0. Accordingly, the plaintext corresponding to the sum of the N ciphertext outputs is actually equivalent to the output of the first linear layer. Furthermore, the inference result of the target data can be obtained using the sum of the N ciphertext outputs.

[0101] If there are other linear layers after the first linear layer, the specific implementation of step 307 can be as follows: input the sum of the N ciphertext outputs into the subsequent linear layer to obtain the output of the subsequent linear layer, then decrypt the output, and input the decryption result into the subsequent nonlinear layer of the linear layer to obtain the inference result.

[0102] Furthermore, step 307 can be implemented by the inference program outputting the inference result of the target data based on the sum of the N ciphertext outputs. Specifically, the inference program performs the following steps: determining the sum of the N ciphertext outputs; and using the aggregation key corresponding to the N shard keys written in the inference program to decrypt the sum of the N ciphertext outputs to obtain the inference result of the target data.

[0103] Specifically, after obtaining the sum of N ciphertext outputs, the aggregation key written in the inference program can be used to decrypt them. The decrypted data is then used to perform inference on subsequent layers of the first linear layer to obtain the final inference result.

[0104] For example, the first linear layer is the last fully connected layer in a neural network model. The output of this fully connected layer can be input into an activation function used for classification, such as the softmax function, to obtain the final inference result.

[0105] This minimizes the number of time-consuming homomorphic operations and reduces the time spent on reasoning.

[0106] The implementation of the Paillier distributed key will be explained next.

[0107] As mentioned earlier, the model needs to obfuscate the first linear layer of the neural network model to obtain m-1 second linear layers. It should be noted that the previous mention referred to the generation of N-1 second linear layers; since N will represent other data later, m here is actually the N from the previous context. The number of parameters in the second linear layers is the same as in the first linear layer, but the parameters in W and b of the second linear layers are all 0. The first linear layer and the N-1 second linear layers are then encrypted. The encryption process of the W matrix is ​​used as an example to illustrate the encryption process of the first and second linear layers. It can be understood that the encryption process of the bias vector b is the same as the encryption process of the weight matrix W, and will not be repeated here. The encryption process includes the following steps:

[0108] 1. Key Generation: Given the security parameter π, the model first generates public and private keys for the Paillier encryption algorithm, where the private key SK p =(μ,λ), public key PK p = (N, g), where the order of the public key parameter g is a multiple of N, i.e., ord(g) = kN (if k = 1, g = 1 + N, g N = (1+N) N modN 2 =1,g kN modN 2 =1) Then select a large random integer γ that satisfies γ∈{0,1} π / 2 Given gcd(k,γ)=1, calculate h=gγ modN 2 The publicly disclosed system parameters are PP = <π,N,g,h>.

[0109] 2. Key splitting: splitting the public key into PK keys. p The parameter N is divided into m random numbers {n1, n2, ..., nn}. m},satisfy m is the total number of the first and second linear layers. Then, random numbers are selected for each model confusion task. As the task ID, a sharding key is generated for each linear layer. and aggregation key SK data =<λ,μ,γ>.

[0110] 3. Model Encryption: The first and second linear layers are encrypted using the m fragment keys obtained above and the public system parameters. Specifically, given the linear layer parameters w... i ∈Z N The model is encrypted to obtain:

[0111] Correspondingly, the data provider performs inference and obtains m ciphertext outputs [[y i By aggregating the m ciphertext outputs, we can obtain...

[0112] Furthermore, using the aggregation key SK data Decrypting [[y]] yields... Where L(x) = (x-1) / N.

[0113] The correctness of the above process will be explained next.

[0114] According to the encryption formula, the ciphertext reasoning result can be transformed as shown in formula (1):

[0115] Correspondingly, the decrypted result is shown in formula (2):

[0116] As can be seen, the plaintext corresponding to [[y]] is y. (i) The sum of.

[0117] The implementation of the MPCKKS distributed key will be explained next. This scheme is based on the CKKS distributed key, a fully homomorphic encryption algorithm.

[0118] Before using CKKS encryption, it is first necessary to define R. q =Zq [X] / (X N +1), a ring of residues of a polynomial. Simply put, R q It is a set of polynomials, where each element is a polynomial with coefficients in {0,1,…,q-1} and a degree not exceeding N.

[0119] The scheme described in this manual involves the following operations in the CKKS algorithm:

[0120] 1. KeyGen: Key generation algorithm, generates private key s←R q ;

[0121] 2. ECD: An encoding algorithm that takes a vector of length less than N as input and transforms it into R. q polynomials in;

[0122] 3. Enc: Encryption algorithm. Given a polynomial and a private key, it encrypts the polynomial into ciphertext. Specifically, for a polynomial m∈R... q The encryption algorithm first came from R q Randomly select polynomial a←R q , and noise polynomial e←R q (The coefficient of e is relatively small), and then calculate. And output;

[0123] 4. HomMulPt: Plaintext-Ciphertext Multiplication, takes a plaintext polynomial m′∈R as input. q And a secret calculate The new ciphertext encrypts m·m′;

[0124] 5. Dec: Decryption function, input ciphertext (c0, c1) ∈ R q Calculate c0,+c1·s;

[0125] 6. Dcd: Decoding function, takes a polynomial as input and outputs a vector (the inverse process of encoding function).

[0126] The reasoning process of the data side involves calculating matrix-vector multiplication. Matrix-vector multiplication can be achieved using coefficient encoding by calling the plaintext-ciphertext multiplication function HomMulPt once. For example, as shown in Figure 4, calculating two vectors... The inner product can be used to... These correspond sequentially to the coefficients of the polynomial, while the vector... By corresponding them in a certain reverse order, we can ensure that the constant term of the product of the two polynomials is the inner product of the two vectors.

[0127] Furthermore, it should be noted that this specification also includes improvements to CKKS. Specifically, during encryption, different ciphertexts share the same random polynomial a∈R. q Since distributed key encryption requires that the aggregated ciphertext of data encrypted with fragments can be decrypted with the aggregate key, CKKS can implement distributed key encryption by allowing different ciphertexts to share the same polynomial coefficients.

[0128] Encryption is performed by obfuscating the first linear layer (e.g., a fully connected layer). Regarding the parameters of the first and second linear layers, assuming there are two second linear layers and a total of 3 first linear layers, the encryption methods include:

[0129] The model encodes the parameter matrices of the three linear layers into three polynomials m0, m1, m2 ∈ R. q Then the model uses three fragment keys s0, s1, and s2 to encrypt the three polynomials respectively, resulting in three ciphertexts as shown in formula (3).

[0130] During the inference process, the data provider encodes its own sample vector (i.e., the target data) as a polynomial m′∈R. q Then the data provider calls HomMulPt to calculate the plaintext-ciphertext multiplication three times, and obtains the ciphertext output corresponding to the three linear layers, as shown in formula (4).

[0131] The data provider performs a homomorphic addition operation on the three ciphertext outputs to obtain the aggregated result of the three ciphertext outputs, (d,a′)=(c0+c1+c2,a′). The data provider uses the aggregated key for decryption, where the aggregated key s=s0+s1+s2. The decryption process is shown in formula (5). Dec(s,(d,a′))=d+a′·s =(c0+c1+c2)+a′·(s0+s1+s2) =m′·(b0+a·s0)+m ′ ·(b1+a·s1)+m′·(b2+a·s2) =m′·(m0+m1+m2)+m′·(e0+e1+e2) (5)

[0132] As can be seen, the data provider can directly decrypt m′·m, m=m0+m1+m2, where m′ is the sample vector and m is the sum of the parameter matrices of the three linear layers. These two polynomial multiplications already include the result of the inner product; therefore, the data provider can then call the decoding function to recover the result.

[0133] It should also be noted that the CKKS homomorphic encryption algorithm may contain noise, which can be eliminated by selecting appropriate parameters.

[0134] As for the correctness, it can be directly obtained through the above formula derivation.

[0135] This specification also provides a neural network model inference device, applied to the data side, as shown in Figure 5:

[0136] The receiving module 510 is used to receive an encrypted neural network model sent by the model provider; the encrypted neural network model includes a first ciphertext linear layer and N-1 second ciphertext linear layers, the first ciphertext linear layer and the N-1 second ciphertext linear layers are respectively encrypted using N fragmentation keys, the N fragmentation keys are generated based on a distributed encryption protocol that supports homomorphic operations; the first ciphertext linear layer is obtained by encrypting the first linear layer included in the neural network model, and the plaintext of the parameters of any second ciphertext linear layer is 0;

[0137] The first inference module 520 is used to infer the target data using the encrypted neural network model to obtain intermediate results of the first ciphertext linear layer and the second ciphertext linear layer to be input.

[0138] The second inference module 530 is used to input the intermediate results into the first ciphertext linear layer and N-1 second ciphertext linear layers respectively to obtain N ciphertext outputs;

[0139] The output module 540 is used to obtain the reasoning result of the target data by using the sum of N ciphertext outputs.

[0140] In an alternative implementation, the neural network model is also encrypted based on a second encryption algorithm.

[0141] In an optional embodiment, the receiving module 510 is further configured to receive an inference program for the encrypted neural network model obtained by obfuscation using a secure compilation method; wherein the inference of the target data to obtain the inference result is implemented based on the inference program.

[0142] In an optional implementation, the output module 540 is specifically used for the inference program to output the inference result of the target data based on the sum of the N ciphertext outputs; the inference program executes: determining the sum of the N ciphertext outputs; using the aggregation key corresponding to the N shard keys written in the inference program to decrypt the sum of the N ciphertext outputs to obtain the inference result of the target data.

[0143] In an optional embodiment, the neural network model is further encrypted based on a second encryption algorithm; the first inference module 520 is further configured to decrypt the encrypted neural network model using the key of the second encryption algorithm written in the inference program.

[0144] In one alternative implementation, the inference program invokes the unobfuscated Triton service framework during runtime; the inference program and the Triton service framework communicate via piped communication during runtime.

[0145] In one alternative implementation, the first linear layer is the last fully connected layer of the neural network model.

[0146] In one optional implementation, the distributed encryption protocol supporting homomorphic operations is a Paillier-based distributed key protocol; or, the distributed encryption protocol supporting homomorphic operations is a CKKS-based distributed key protocol, wherein the polynomial coefficients are the same when different fragment keys are encrypted.

[0147] As shown in Figure 6, this specification also provides a neural network model inference device, applied to a model provider with a neural network model, including:

[0148] The obfuscation module 610 is used to generate a second linear layer with N-1 parameters set to 0 corresponding to the first linear layer;

[0149] The encryption module 620 is used to obtain N fragment keys generated by the first encryption algorithm, and to encrypt the first linear layer and N-1 second linear layers respectively using the N fragment keys to obtain the first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations;

[0150] The sending module 630 is used to send an encrypted neural network model to the data party so that the data party can use the encrypted neural network model for inference.

[0151] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0152] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0153] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this specification does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0154] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.

[0155] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0156] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0157] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0159] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0160] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0161] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0162] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0163] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0164] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0165] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.

Claims

1. A neural network model inference method, applied to the data side, comprising: The receiver sends an encrypted neural network model; the encrypted neural network model includes a first ciphertext linear layer and N-1 second ciphertext linear layers, the first ciphertext linear layer and the N-1 second ciphertext linear layers are respectively encrypted using N fragmentation keys, the N fragmentation keys are generated based on a distributed encryption protocol that supports homomorphic operations; the first ciphertext linear layer is obtained by encrypting the first linear layer included in the neural network model, and the plaintext of the parameters of any second ciphertext linear layer is 0; The encrypted neural network model is used to reason about the target data to obtain intermediate results of the first ciphertext linear layer and the second ciphertext linear layer to be input. The intermediate results are input into the first ciphertext linear layer and N-1 second ciphertext linear layers respectively to obtain N ciphertext outputs; The inference result of the target data is obtained by summing the N ciphertext outputs.

2. The method according to claim 1, wherein the neural network model is further encrypted based on a second encryption algorithm.

3. The method according to claim 1, further comprising: An inference program that receives the encrypted neural network model obtained by obfuscation using a secure compilation method; The reasoning of the target data to obtain the reasoning result is implemented based on the reasoning program.

4. The method according to claim 3, wherein obtaining the inference result of the target data by summing the N ciphertext outputs comprises: The reasoning program outputs the reasoning result of the target data based on the sum of the outputs of N ciphertexts; The inference procedure executes as follows: Determine the sum of the N ciphertext outputs; Using the aggregation key corresponding to the N fragment keys written in the inference program, the sum of the N ciphertext outputs is decrypted to obtain the inference result of the target data.

5. The method according to claim 3, wherein the neural network model is further encrypted based on a second encryption algorithm; The inference procedure is also used to execute: The encrypted neural network model is decrypted using the key of the second encryption algorithm written in the inference program.

6. The method according to claim 3, wherein the inference program calls the unobfuscated Triton service framework during operation; the inference program and the Triton service framework communicate via pipe communication during operation.

7. The method according to claim 1, wherein the first linear layer is the last fully connected layer of the neural network model.

8. The method according to claim 1, wherein the distributed encryption protocol supporting homomorphic operations is a Paillier-based distributed key protocol; Alternatively, the distributed encryption protocol supporting homomorphic operations is a distributed key protocol based on CKKS, wherein... The polynomial coefficients are the same when different fragment keys are used for encryption.

9. A neural network model inference method, applied to a model provider possessing a neural network model, comprising: Generate a second linear layer with N-1 parameters set to 0, corresponding to the first linear layer; Obtain N fragment keys generated using the first encryption algorithm, and use the N fragment keys to encrypt the first linear layer and N-1 second linear layers respectively, to obtain the first ciphertext linear layer and N-1 second ciphertext linear layers; the first encryption algorithm is implemented based on a distributed encryption protocol that supports homomorphic operations; An encrypted neural network model is sent to the data party so that the data party can use the encrypted neural network model for inference.

10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-9.