Rapid binary neural network inference scheme based on three-party copy secret sharing

By adopting a fast binary neural network inference solution based on three-party replication secret sharing in deep learning as a service, the problem of how to perform efficient neural network inference without infringing on user privacy is solved, and the optimization of time and communication overhead while maintaining good accuracy is achieved.

CN120124739APending Publication Date: 2025-06-10BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510130510.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In deep learning as a service, how to conduct efficient neural network reasoning without infringing on user privacy, especially on devices with limited resources?

Method used

A fast binary neural network inference scheme based on three-party replication secret sharing is adopted, and the optimization of time and communication overhead while maintaining good accuracy is achieved by using binary neural networks and replication secret sharing technology.

Benefits of technology

Without leaking the original data, efficient neural network reasoning is achieved, which significantly improves computing efficiency and resource utilization, and is suitable for privacy-sensitive application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124739A_ABST
    Figure CN120124739A_ABST
Patent Text Reader

Abstract

The invention discloses a fast binary neural network reasoning scheme based on three-party copy secret sharing. Firstly, a binary neural network (BNN) is trained through linear scaling and network cutting, then input data and a model are encrypted and distributed to the other two parties, the three parties conduct encryption reasoning on the data based on copy secret sharing, and after reasoning is finished, a user decrypts a reasoning result and obtains final output. In the training and network pruning stage, a model owner firstly trains a binary neural network, and linear scaling and a model pruning strategy based on weight flipping frequency are applied in the training process to improve the reasoning precision of the neural network; in a data distribution and coding stage, input data and a model of a plaintext are converted into copied secrets to be shared and distributed to other two parties. BNN parameters are coded into arithmetic secret sharing, and coding bit widths can be flexibly selected for different layers; in a three-party safety reasoning stage, each party executes safety reasoning according to an algorithm customized for the height of each layer of the BNN.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of secure multi-party computing, and in particular to a fast binary neural network reasoning method based on three-party copy secret sharing. Background Art

[0002] Deep learning models are able to perform highly complex tasks by automatically learning complex features from large amounts of data. Deep learning can be widely used in many applications such as medical diagnosis, credit risk assessment, and facial recognition. Due to its superior performance, deep learning as a service has become a popular business model. In deep learning as a service, the service provider provides a trained neural network and the user calls the API for data analysis. However, this inference service requires the client to disclose its input to the service provider, which raises privacy issues. If the service provider discloses the neural network to the user in turn, it will not only cause business losses but also may violate the law. With the increasingly stringent data protection laws and regulations, how to use deep learning without infringing user privacy has become a pressing issue to be solved.

[0003] In this context, privacy-preserving deep learning has emerged. Its goal is to protect the privacy of data content while conducting deep learning, which involves various privacy protection technologies, such as homomorphic encryption, secure multi-party computation, differential privacy, etc. Homomorphic encryption technology allows data to be processed in an encrypted state, which means that neural networks can be trained and reasoned without decrypting the data. Differential privacy is a technology that ensures the protection of individual privacy when publishing aggregated information. The application of differential privacy in neural networks means that changes in algorithm output will not significantly depend on any single data point, thereby protecting the privacy of individuals in the data. Secure multi-party computation is a general cryptographic primitive that allows distributed participants to collaborate in computing arbitrary functions and output accurate calculation results without leaking the original input data of the participants. With the widespread application of distributed computing and cloud computing models, secure multi-party computation has increasingly become one of the research hotspots in the international cryptography community.

[0004] Techniques such as obfuscated circuits, secret sharing, and oblivious transfer are used to design secure multi-party computation protocols. Schemes based on homomorphic encryption have high communication efficiency but high computational overhead. Work based on obfuscated circuits requires only one round of constant interaction regardless of the depth of the circuit, but has high communication overhead and is expensive for arithmetic operations. Methods based on secret sharing provide efficient arithmetic operations and support nonlinear functions with less communication, but usually require an overhead proportional to the number of multiplications, which may result in significant runtime when evaluating circuits with higher depth.

[0005] While these technologies protect data privacy, they also bring new challenges, especially in terms of efficiency. Applying these privacy-preserving techniques in neural network inference often increases inference time and reduces overall performance. To address this problem, binary neural networks (BNNs) provide a promising direction. BNNs simplify the weights and activation functions in neural networks into binary forms containing only 0 and 1, significantly reducing the computational complexity and storage requirements of the model. This simplification makes BNNs faster and more efficient than traditional deep neural networks when performing inference tasks, especially on devices with limited resources.

[0006] By combining binary neural networks with the above privacy protection technologies, the efficiency of neural network reasoning can be improved while protecting the privacy of user data. For example, using secret sharing or homomorphic encryption technology to encrypt the weights and inputs of BNNs can perform efficient reasoning calculations without leaking the original data. This method is particularly suitable for application scenarios that need to process sensitive information, such as medical diagnosis and financial risk assessment, which not only ensures data privacy but also maintains the efficient operation of neural networks. In this way, the practicality of deep learning can be promoted in a wider range of application fields, especially in privacy-sensitive scenarios.

[0007] In order to solve the above problems, the present invention proposes a fast binary neural network inference scheme based on three-party replicated secret sharing, which sets the participants as three parties and uses a relatively advanced replicated secret sharing technology with low computational overhead. This scheme uses a binary neural network to achieve greater optimization in time and communication overhead while maintaining good accuracy. Summary of the invention

[0008] This paper proposes a fast binary neural network reasoning scheme based on three-party replicated secret sharing. This scheme sets the participants to three parties and uses a relatively advanced replicated secret sharing technology with low computational overhead as the underlying secure multi-party computing primitive. This paper uses a binary neural network to achieve significant optimization in time and communication overhead while maintaining good reasoning accuracy. The main contributions can be summarized as follows:

[0009] (1) The secure computation framework for binary neural network reasoning proposed in this invention can perform efficient computation using three parties in an honest majority setting. At a high level, the loop size of the binary neural network security evaluation can be flexibly selected according to the specific binary neural network architecture to improve efficiency.

[0010] (2) The present invention specifically optimizes the convolution, binary activation, batch normalization and maximum pooling operations in the binary neural network, and especially integrates the binary activation function into the batch normalization and maximum pooling operations to improve efficiency.

[0011] (3) The present invention implements network pruning on a binary neural network, and combines linear scaling and network pruning to achieve a balance between the inference accuracy and inference overhead of the binary neural network.

[0012] In order to achieve the above purpose, the technical solution adopted by the present invention is a fast binary neural network inference solution based on three-party copy secret sharing, and its system model has three types of entities: data owner, model owner, and auxiliary third party. Figure 1 shown.

[0013] The model owner first trains a binary neural network using linear scaling and network pruning. Then the data owner and the model owner encrypt the input data and the model respectively and distribute them to the other two parties. The three parties perform encrypted inference on the data based on replicated secret sharing. After the inference is completed, the user decrypts the inference result and obtains the final output.

[0014] This program has 3 steps, the specific steps are as follows.

[0015] (1) Training and network pruning phase: The model owner first trains a binary neural network and applies linear scaling and model pruning strategies based on weight flipping frequency during the training process to improve the neural network inference accuracy.

[0016] (2) Data distribution and encoding phase: The data owner and the model owner convert the plaintext input data and model into replicated secret shares and distribute them to the other two parties. The binary neural network parameters are encoded into arithmetic secret shares, and the encoding bit width can be flexibly selected for different layers.

[0017] (3) Three-party joint reasoning stage: The three parties perform joint reasoning without knowing anything about the data of the other two parties. The reasoning stage is divided into linear layer, binary activation layer, batch normalization layer and pooling layer.

[0018] a) Linear layer: The input layer scales the input matrix while keeping the weight matrix unscaled, removing the truncation operation while retaining enough fractional parts.

[0019] b) Binary Activation Layer: At the beginning of the binary activation layer calculation, the three parties hold the replicated secret share corresponding to the evaluation output of the previous layer of the network, and pass the participant S i The shared share of the replicated secret held by (i=1,2,3) is expressed as < <x>> i First, the replicated secret sharing is transformed into a two-party secret sharing (expressed as <x> i ), then extract the most significant bit (MSB), and finally call the three-party oblivious transfer module to convert the two-party secret sharing of the most significant bit into a replicated secret sharing form with binary activation output.

[0020] c) Batch Normalization Layer: While considering the batch normalization parameter to be negative, the ring conversion operation is removed and the binary activation calculation is integrated into this layer to improve efficiency.

[0021] d) Pooling layer: The max pooling operation will produce an output of 1 if all elements within the window are 1; conversely, if any element is -1, it will output -1.

[0022] This paper proposes a fast binary neural network reasoning scheme based on three-party copy secret sharing. By designing the submodules of binary neural network training and secure reasoning, it achieves a huge efficiency improvement while maintaining good accuracy. Compared with the existing secure neural network reasoning scheme, this paper has reached the forefront level in computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 System architecture diagram

[0024] Figure 2 Figure 2 is a pruning algorithm diagram based on weight flip frequency.

[0025] Figure 3 Algorithm diagram for extracting the highest security bit

[0026] Figure 4 Diagram of the three-party oblivious transmission algorithm

[0027] Figure 5 To generate the activation output algorithm graph

[0028] Figure 6 Comparison chart of calculation logic for the maximum pooling layer DETAILED DESCRIPTION

[0029] The model owner first trains a binary neural network using linear scaling and network pruning. Then the data owner and the model owner encrypt the input data and the model respectively and distribute them to the other two parties. The three parties perform encrypted inference on the data based on replicated secret sharing. After the inference is completed, the user decrypts the inference result and obtains the final output.

[0030] The specific implementation method is as follows.

[0031] (1) Training and network pruning phase: linear scaling and network pruning; Initially, the scaled BNN architecture is trained by uniformly adjusting the number of channels / neurons in all linear layers by a scaling factor u=2 before the training phase; Subsequently, the scaled model undergoes a weight pruning process based on the weight flipping frequency, in which the channels / neurons that contribute the least to the network inference accuracy are removed; The weight flipping frequency of a binary neural network represents the number of times the weight value changes from 0 / -1 to 1 within a specified training interval. The marginal increase in accuracy near the end of the training process is mainly attributed to the weight updates characterized by a high frequency of weight flipping. The effect of this group of weights with high flipping frequency on accuracy enhancement is negligible, so they are pruned; the pruning algorithm trains a BNN model from scratch (if there is a pre-trained model, the last training stage needs to be re-executed to record weight changes); then, the flipping frequency f of each weight is counted in the training interval stime to etime, and the insensitive weight ratio pT% that satisfies f≥2 is calculated layer by layer; the algorithm reduces the number of channels of the corresponding layer by pT%, compresses the model scale, and retrains the network at the new scale; the above process is iterated until the pT% of each layer is less than 0.5%, indicating that the model cannot be further compressed;

[0032] (2) Data distribution and encoding phase: The data owner and the model owner convert the plaintext input data matrix X and weight matrix W into replicated secret shares and distribute them to the other two parties. The data owner S 0 Holds the replicated secret share (X 0 , X 1 )、(W 0 , W 1 ), model owner S 1 Hold(X 1 , X 2 )、(W 1 , W 2 ), assisting third party S 2 Hold(X 2 , X 0 )、(W 2 , W 0 ), where X 0 +X 1 +X 2 =X,W 0 +W 1 +W 2 =W, for the convenience of representation, each participant is denoted as S i , the data held is recorded as (X i , X i+1 )、(W i , W i+1 ), where i∈{0, 1, 2}, the elements in the input data matrix X and the model weight matrix W are denoted as x and w respectively; in secure multi-party computing, all numbers need to be encoded as integers. For the input layer and output layer of BNN, a standard fixed-point encoding scheme with enhanced bit width, i.e., bit width k=32, is adopted to meet the accuracy requirements of real-valued data; the elements x, w∈±1, " and X in the intermediate layer parameter matrices X, W are of dimensions (h, n) and (n, k) respectively, and the resulting product WX must fall on [-n, n], taking k=log 2 (2n+2) to encode all numbers in this interval;

[0033] (3) Three-party secure reasoning stage: The three parties perform joint reasoning without knowing anything about the data of the other two parties. The reasoning stage is divided into a linear layer, a binary activation layer, a batch normalization layer, and a maximum pooling layer. The elements in the input data matrix X, the weight matrix W, and the output matrix z of each layer are represented as x, w, z, and z. i Indicates the output results of each calculation party;

[0034] a) Linear layer: Each participant S in the fully connected layer i Hold(W i , W i+1 ),satisfy and data sharing (X i , X i+1 ),satisfy Each participant S i According to formula (1), the secret sharing output Z of the fully connected layer is calculated i , where i∈{0, 1, 2}:

[0035] Z i =W i X i +W (i+1) X i +W i X (i+1) #(1)

[0036] If the current layer is not a hidden layer (i.e., the input layer or the output layer), each participant S i The b i Add to Z i , b represents the bias parameter, b i Represents the secret sharing share of the bias parameter b:

[0037] Z i =Z i +b i #(2)

[0038] Output the secret sharing of the fully connected layer calculation results; in the protocol, if the fixed points W and X have l W and l X bits of precision, then b should have l W +l X Bit accuracy;

[0039] b) Binary activation layer: After completing the reasoning of the fully connected layer or convolutional layer, each participant holds a part of the reasoning result. This layer calculates the activation output shown in formula (3), where x represents the element in the input matrix X; first, the copied secret share < <z>> i Transformed into two-party secret sharing <z> i , each participant S i First, a pseudo-random function is started to generate a random distribution value a for each participant. i ,satisfy Then the participant S 2 z 2 +a 2 Send to S 0 , because S 2 Sent 2 +a 2 is uniformly randomly distributed so no information is leaked. Now S 0 Hold m 0 =(z 0 +a 0 )+(z 2 +a 2 ), and S 1 hold

[0040] m 1 =z 1 +a 1 ; m=m 0 +m 1 =∑(z i )+∑(a i )=∑(z i )=z, thus realizing the conversion from copy secret sharing to two-party secret sharing; then extract the most significant bit MSB and use m 0 [i] and m 1 [i] represents m 0 and m 1 The i-th position of Where c is the carry bit of the l-1th bit, and the three parties use a 3-input AND gate to evaluate the parallel prefix adder circuit to calculate the carry bit c; call the randomness function and the three-party oblivious transfer to generate the output of the binary activation layer; in the three-party oblivious transfer algorithm, the receiver and the sender first independently generate a pair of cryptographic keys based on a pre-agreed asymmetric encryption scheme, and the sender encrypts the secret value with the receiver's public key. The sender calculates a blinded version of each encrypted message, and the sender transmits the blinded message to the auxiliary party. The receiver encrypts its selected bits with the auxiliary party's public key and sends it to the auxiliary party. After decryption, the auxiliary party identifies the receiver's choice and forwards the blinded version of the selected message to the receiver. The receiver removes the blinding mask from it and decrypts the message with its own private key; if the activation layer is followed by a maximum pooling layer, a pooling operation is performed after extracting the MSB before generating the activation output;

[0041]

[0042] c) Batch Normalization Layer: In traditional neural network architectures, the output of the batch normalization (BN) layer is determined by the following formula

[0043] z=γ·x+β#(4)

[0044] x, z represent the elements in the input matrix X and output matrix Z of the layer; γ and β are parameters optimized during the training process, and are fixed when the training and network pruning phase ends, that is, when the insensitive weight ratio pT% is less than 0.5%; in BNN, the BN layer is always followed by an activation layer. The BN layer and the activation layer are calculated together to obtain the output formula of the combined batch normalization and binary activation (BNBA) layer

[0045] z=Sign(γ·x+β)#(5)

[0046] Since γ·x requires complex and time-consuming ring transformation operations, the calculation formula is further converted to

[0047]

[0048] d) Maximum pooling layer: If MSB(m) = 0, the neuron should be activated with an activation value of 1, otherwise it should be deactivated with an inactivation value of -1. After extracting the most significant bit, the maximum pooling algorithm performs an AND operation on the elements in each sliding window, and then performs an activation operation on the result.

[0049] In the embodiments provided in this application, it should be understood that the disclosed method can be implemented in other ways without exceeding the spirit and scope of this application. The current embodiment is only an illustrative example and should not be used as a limitation. The specific content given should not limit the purpose of this application.

[0050] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.< / z> < / z> < / x> < / x>

Claims

1. A fast binary neural network inference scheme based on three-party copy secret sharing, characterized by: There are three types of entities: data owner, model owner, and auxiliary third party; the model owner first uses linear scaling and network pruning to train a binary neural network BNN, whose weights and activation values ​​are limited to ±1, and then the data owner and model owner encrypt the input data and model respectively and distribute them to the other two parties. The three parties perform encrypted reasoning on the data based on replicated secret sharing. After the reasoning is completed, the user decrypts the reasoning result and obtains the final output; the specific implementation method is as follows; (1) Training and network pruning phase: linear scaling and network pruning; Initially, the scaled BNN architecture is trained by uniformly adjusting the number of channels / neurons in all linear layers by a scaling factor u = 2 before the training phase; Subsequently, the scaled model undergoes a weight pruning process based on the weight flipping frequency, in which the channels / neurons that contribute the least to the network inference accuracy are removed; the weight flipping frequency of the binary neural network represents the number of times the weight value changes from 0 / -1 to 1 within a specified training interval. The marginal increase in accuracy near the end of the training process is mainly attributed to the weight update characterized by high-frequency weight flipping, while the group of weights with high flipping frequency has a negligible effect on the accuracy enhancement and is pruned; The pruning algorithm trains a BNN model from scratch; if there is a pre-trained model, the last training stage needs to be re-executed to record the weight changes; then, the flip frequency f of each weight is counted in the training interval stime to etime, and the insensitive weight ratio pT% that satisfies f≥2 is calculated layer by layer; the algorithm reduces the number of channels of the corresponding layer according to pT%, compresses the model size, and retrains the network at the new size; the above process is iterated until the pT% of each layer is less than 0.5%, indicating that the model cannot be further compressed, and the training and network pruning stages are completed; (2) Data distribution and encoding phase: The data owner and the model owner convert the plaintext input data matrix X and weight matrix W into replicated secret shares and distribute them to the other two parties. The data owner S0 holds the replicated secret shares (X0, X1) and (W0, W1) of the input data matrix X and the model weight matrix W, the model owner S1 holds (x1, X2) and (W1, W2), and the auxiliary third party S2 holds (X2, X0) and (W2, W0), where X0+X1+X2=X, W0+W1+W2=W. For the convenience of representation, each participant is denoted as S. i , the data held is recorded as (X i ,X i+1 )、(W i ,W i+1 ), where i∈{0,1,2}, the elements in the input data matrix X and the model weight matrix W are denoted as x and w respectively; in secure multi-party computation, all numbers need to be encoded as integers. For the input and output layers of BNN, a standard fixed-point encoding scheme with enhanced bit width, i.e., bit width k=32, is adopted to meet the accuracy requirements of real-valued data; the elements x,w∈±1 in the intermediate layer parameter matrices X,W, the dimensions of W and X are (h,n) and (n,k) respectively, and the obtained product WX must fall on [-n,n], and k=log2(2n+2) is taken to encode all numbers in this interval; (3) Three-party secure reasoning stage: The three parties perform joint reasoning without knowing anything about the data of the other two parties. The reasoning stage is divided into a linear layer, a binary activation layer, a batch normalization layer, and a maximum pooling layer. The elements in the input data matrix X, the weight matrix W, and the output matrix Z of each layer are represented as x, w, z, and Z i Indicates the output results of each calculation party; a) Linear layer: Each participant S in the fully connected layer i Hold(W i ,W i+1 ),satisfy and data sharing (X i ,X i+1 ),satisfy Each participant S i According to formula (1), the secret sharing output Z of the fully connected layer is calculated i , where i∈{0,1,2}; Z i =W i X i +W (i+1) X i +W i X (i+1) #(1) If the current layer is not a hidden layer, but an input layer or an output layer, then each participant S i The b i Add to Z i , b represents the bias parameter, b i Represents the secret sharing share of the bias parameter b: Z i =Z i +b i #(2) Output the replicated secret sharing share of the fully connected layer calculation result; the fixed-point matrices W and X of the input layer have l W and l X bits of precision, then b should have l W +l X Bit accuracy; b) Binary activation layer: After completing the reasoning of the fully connected layer or convolutional layer, each participant holds a part of the reasoning result. This layer calculates the activation output shown in formula (3), where x represents the element in the input matrix X; first, the copied secret share < <z>> i Transformed into two-party secret sharing <z> i , each participant S i First, start the pseudo-random function to generate a random distribution value a for each participant. i ,satisfy Then participant S2 sends z2+a2 to S0. Since z2+a2 sent by S2 is uniformly randomly distributed, no information will be leaked. Now S0 holds m0=(z0+a0)+(z2+a2), and S1 holds m1=z1+a1; m=m0+m1=∑(z i )+∑(a i )=∑(z i )=z, thus realizing the conversion from duplicate secret sharing to two-party secret sharing; then extract the most significant bit MSB, and use m0[i] and m1[i] to represent the i-th bit of m0 and m1 respectively; obviously, Where c is the carry bit of the l-1th bit. The three parties use the activation output generation algorithm to call the three-input AND gate to evaluate the parallel prefix adder circuit to calculate the carry bit c, and call the randomness function and the three-party oblivious transfer to generate the output of the binary activation layer. In the three-party oblivious transfer algorithm, the receiver and the sender first independently generate a pair of cryptographic keys based on a pre-agreed asymmetric encryption scheme. The sender encrypts the secret value with the receiver's public key. The sender calculates a blinded version of each encrypted message. The sender transmits the blinded message to the auxiliary party. The receiver encrypts its selected bit with the auxiliary party's public key and sends it to the auxiliary party. After decryption, the auxiliary party identifies the receiver's choice and forwards the blinded version of the selected message to the receiver. The receiver removes the blinding mask from it and decrypts it with its own private key to obtain the message.< / z> < / z> If the activation layer is followed by a maximum pooling layer, the pooling operation is performed before generating the activation output after extracting the MSB; c) Batch Normalization Layer: In traditional neural network architectures, the output of the batch normalization (BN) layer is determined by the following formula z=γ·x+β #(4) x, z represent the elements in the input matrix X and output matrix Z of the layer; γ and β are parameters optimized during the training process, and are fixed when the training and network pruning phase ends, that is, when the insensitive weight ratio pT% is less than 0.5%; in BNN, the BN layer is always followed by an activation layer. The BN layer and the activation layer are calculated together to obtain the output formula of the combined batch normalization and binary activation (BNBA) layer z=Sign(γ·x+β)#(5) Since γ·x requires complex and time-consuming ring transformation operations, the calculation formula is further converted to d) Maximum pooling layer: If MSB(m) = 0, the neuron should be activated with an activation value of 1, otherwise it should be deactivated with an inactivation value of -1. After extracting the most significant bit, the maximum pooling algorithm performs an AND operation on the elements in each sliding window, and then performs an activation operation on the result.