Novel privacy activation function design method based on secure multi-party computing

By designing a new privacy activation function that is suitable for secure multi-party computing, the problem of collaborative computing of new activation functions in neural networks is solved, data privacy protection and secure sharing are realized, and data security and calculation accuracy are improved.

CN120541876AActive Publication Date: 2025-08-26CENTRAL UNIVERSITY OF FINANCE AND ECONOMICS

Patent Information

Application Number
CN202510618613.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing secure multi-party computing methods are difficult to effectively implement collaborative computing of new activation functions, especially in neural networks, making it difficult for traditional methods to protect data privacy and realize the availability of data invisible.

Method used

A new type of privacy activation function based on secure multi-party computing is designed. By analyzing the activation function, the basic components adapted to secure multi-party computing are selected for transformation, so that the modified activation function holds addition shards of secret data at input and output, and adapts to the calculation framework of the neural network.

Benefits of technology

It realizes that all parties in the neural network share data for calculation without leaking sensitive information, improves data security and privacy protection, adapts to the computing needs of new activation functions, and avoids accuracy loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541876A_ABST
    Figure CN120541876A_ABST
Patent Text Reader

Abstract

The invention provides a novel privacy activation function design method based on secure multi-party computing, which comprises the following steps: analyzing an activation function in a neural network, and determining mathematical operation steps and algebraic properties of the activation function; selecting a corresponding basic security multi-party computing component according to the mathematical operation steps and algebraic properties; the activation function is transformed through the selected basic security multi-party calculation component, so that the transformed activation function adapts to a calculation framework required by security multi-party calculation, and calculation of secret input data is achieved; and keeping each participant to hold addition fragments of secret data when the modified activation function is input and output, so that the modified activation function is connected with other linear modules or non-linear modules in the neural network before and after. According to the method, the data privacy is protected, the availability of the data is ensured, and the development of data sharing and privacy calculation is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and information security technology, and in particular to a novel privacy activation function design method based on secure multi-party computing. Background Art

[0002] Neural networks (NNs) are a prominent example of artificial intelligence technology, now widely used across various industries for tasks such as classification, regression, and prediction. Using NNs typically requires inputting data in plain text. However, this data often contains sensitive personal information such as medical records and financial information, as well as internal company information such as financial information and source code. Therefore, protecting the privacy of data within neural network models while performing calculations is crucial for individuals and businesses.

[0003] Secure Multiparty Computation (MPC) is a cryptographic technology used to solve the problem of a group of mutually untrusting parties, each holding secret data, collaborating to compute a given function. Therefore, the introduction of this technology provides a mechanism for data holders to protect data privacy. Specifically, the holders share data secretly, splitting the data into multiple copies and storing them with different participants. Each copy is in the form of a random number. Using secure multiparty computation, these participants can then perform neural network calculations together without having to disclose their respective data shares. This approach allows all parties to share data and work together to complete calculations using a model without worrying about leaking sensitive information contained in the complete data, making the data "available but invisible."

[0004] Furthermore, activation functions, as key nonlinear components in neural network architectures, determine their ability to learn and represent complexity. However, implementing collaborative computation of nonlinear functions is a major challenge in secure multi-party computation. In particular, the recent emergence of complex new activation functions complicates linear approximation or splitting, making traditional secure multi-party computation methods difficult to directly apply.

[0005] In order to meet this challenge, it is urgent to develop and design privacy computing methods that are oriented towards the structural characteristics of new activation functions so that they can adapt to the computing framework required for secure multi-party computing. Summary of the Invention

[0006] In view of this, the present invention provides a novel privacy activation function design method based on secure multi-party computing to at least solve the above-mentioned technical problems, promote more extensive data sharing and privacy protection applications, and provide stronger support for the application of secure multi-party computing in the field of neural networks.

[0007] A first aspect of an embodiment of the present invention provides a novel privacy activation function design method based on secure multi-party computing, including: analyzing the activation function in a neural network to determine the mathematical operation steps and algebraic properties of the activation function; selecting the corresponding basic secure multi-party computing component based on the mathematical operation steps and algebraic properties; transforming the activation function through the selected basic secure multi-party computing component so that the transformed activation function adapts to the computing framework required for secure multi-party computing and realizes the calculation of secret input data; maintaining the modified activation function so that each participant holds the addition slice of the secret data at the input and output, so that the modified activation function can be connected to other linear modules or nonlinear modules in the neural network.

[0008] According to a second aspect of an embodiment of the present invention, a privacy neural network computing system is provided, including a data holder, a model holder, and multiple participants, characterized in that: the data holder is used to secretly share data, split it into multiple additive slices, and send the slices to each participant; the model holder is used to determine the neural network structure, train the neural network, obtain the parameters of each layer, and publish the parameters to the data holder; the multiple participants are used to receive the data slices sent by the data holder, and according to the neural network structure and parameters published by the model holder, use the new privacy activation function design method based on secure multi-party computing as described in the first aspect above to perform inference calculations of the neural network, including linear layer calculations and nonlinear layer calculations, and finally send the calculation results back to the data holder.

[0009] In summary, the solution of the present invention enables all parties to share data and jointly use models for calculations without worrying about leaking sensitive information contained in the complete data, achieving "available but invisible" data and improving data security. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 Schematic diagram of the privacy neural network calculation process of the present invention.

[0011] Figure 2 Schematic diagram of the traditional neural network calculation process.

[0012] Figure 3 Schematic diagram of the input layer, hidden layer, and output layer of the privacy neural network of the present invention.

[0013] Figure 4 It is a flow chart of the steps of the present invention. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0015] For ease of understanding, before describing the specific embodiments of the present invention in detail, the technical terms involved in the present invention are explained as follows:

[0016] 1. Representation of secret shared value:

[0017] For each secret value u that needs to be shared, n participants share the secret, using represents the state after secret u is added secret sharing, i Represents participant P i The addition shard held, [u] represents the state after the secret u is shared by the multiplication secret, [u] i Represents participant P i The multiplication shards held. At this point, there are:

[0018] u= 1+ 2+…+ i and u = [u]1 * [u]2 * ... * [u] i ,

[0019] i ± <v> i =<u±v> i ,c* i =<c*u> i

[0020]

[0021] 2. Element multiplication protocol in addition sharing state:

[0022] Assume that n participants want to jointly calculate the result z of x*y, but do not want to directly disclose their inputs x and y. By using the Beaver Triple protocol, they can implement the calculation of element multiplication in the addition shared state. Specifically, participant P i Holds the addition shard of x <x> i and the additive slice of y <y> i , do the following:

[0023] (1) Each participant P i Get a set of Beaver triples of shards ( i , i , <c> i ), where a and b are random numbers and satisfy c=a*b, then we have

[0024] (2) Each participant P i Calculated locally <e> i = <x> i -< / x> < / e> < / c> i and <f> i = <y> i - i , and send their calculation results to other participants, after which each participant can calculate and

[0025] (3) Final participant P1 setup <z> 1=e*f+f*< / z> < / y> < / f> 1+e* 1+ <c>1. Participant P i , i=2,3,...,n, set <z> i =f*< / z> < / c> i +e* i + <c> i , that is, the final calculation result is obtained, satisfying

[0026] 3. Element multiplication protocol in addition sharing state:

[0027] Assume that n participants want to convert the multiplication sharding of a secret value K into addition sharding. By using the following protocol proposed by Josef Pieprzyk, Hossein Ghodosi, Ron Steinfeld et al. in "Multi-Party Computation with Conversion of Secret Sharing" in 2011, they can implement the computation of addition sharding in the multiplication sharing state. Specifically, participant P j Multiplication shards holding K [m] j , that is, at this time there is [m]1*[m]2*...*[m] n =K, the value of K is not disclosed, where j = 1, 2, ..., n. In the calculation steps of this protocol, m is recorded i =[m] i .

[0028] The protocol requires the following auxiliary inputs: each participant P j Has n elements a 1,j ,....,a n,j , which satisfies in The above a i,j It is generated in the precompute as follows:

[0029] (1) Select independent uniform random numbers u1, u2, ..., u n-1 , and calculate

[0030] (2) For each i=1,...,n, select n-1 independent uniform random numbers {a i,j } i≠j , and calculate a i,i =u i *(Π j≠i a i,j ) -1 .

[0031] When the conversion protocol begins, each participant P j (j=1,...,n) sends v i,j =a i,j m j (i=1, .., n) to participant P i Afterwards, participant P i (i=1,...,n) calculation At this time i That is, the additive sharding of the secret value K, because as well as

[0032] 4. Security comparison protocol:

[0033] This protocol is used for secure multi-party computation of piecewise functions involved in this invention. It has different implementation forms, such as the truncation method proposed by Mohassel and Rindal et al. in their 2018 paper “ABY3: A Mixed Protocol Framework for Machine Learning”. Specifically, the protocol inputs a shared secret value <z>, where -2 k <<Z<<2 k ,return- ,in Protocol Analysis: If -2 k ≤Z<0, then so If 0≤Z<2 k ,but so In other words, the security comparison protocol can enter <z>Calculation output , where if Z < 0, then b = 1; if Z ≥ 0, then b = 0.

[0034] 5. ReLU activation function and Private ReLU:

[0035] The ReLU (Rectified Linear Unit) activation function is a classic nonlinear activation function. Its widespread application in deep learning began around 2010, primarily due to its introduction and research by Hinton et al. in their paper "Rectified Linear Units Improve Restricted Boltzmann Machines." Its definition is ReLU(x) = max(0, x). That is, when the input x ≤ 0, the output is 0; when x > 0, the output is x.

[0036] The ReLU activation function is useful for introducing nonlinearity, helping neural network models learn complex nonlinear relationships. The ReLU activation function is very simple to calculate, requiring only a threshold determination. This makes neural network training faster, alleviates the vanishing gradient problem, and is more effective at propagating gradients, making deep neural networks easier to train. Furthermore, the nature of the ReLU function makes neuron activations more sparse, which can help networks generalize and reduce overfitting.

[0037] The ReLU activation function is widely used in deep learning. It is particularly suitable for processing large-scale data and deep neural networks. It can accelerate the training process and improve model performance.

[0038] In 2020, Kumar et al. proposed the concept of Private ReLU in the paper "Cryptflow: Secure tensorflow inference":

[0039] Assume that n participants want to jointly calculate the results of a neural network that contains a nonlinear activation function ReLU, but do not want to disclose their respective inputs. By using the Private ReLU protocol shown below, they can implement ReLU function calculation with privacy function, and thus achieve complete privacy neural network calculation. Specifically, when the neural network reaches the ReLU function calculation step, participant P i Holds the addition slice of the function input x <x> i , at this time, satisfied <x> 1+ <x> 2+…+ <x> n =x, and x is a secret value for the participants, perform the following operations:

[0040] (1) Using the secure comparison protocol, input <x>Get the output , at this time n participants hold Shard 1, 2,..., n ;

[0041] (2) All participants use the Breaver Triple protocol to <x>and <1-b> as input, find<x*(1-b)> Shard, that is, participant P i get <z> i ,make <z> 1+ <z> 2+…+ <z> n =x*(1-b).

[0042] The specific implementation of the embodiment of the present invention is further described below with reference to the accompanying drawings of the embodiment of the present invention.

[0043] See also< / z> < / z> < / z> < / z> < / x> < / x> < / x> < / x> < / x> < / x> < / z> < / z> < / c> Figure 4 The present invention provides a novel privacy activation function design method based on secure multi-party computing, comprising:

[0044] S1. Analyze the activation function in the neural network and determine the mathematical operation steps and algebraic properties of the activation function;

[0045] S2. Selecting a corresponding basic secure multi-party computation component based on the mathematical operation steps and algebraic properties;

[0046] S3. Modify the activation function using the selected basic secure multi-party computing component, so that the modified activation function adapts to the computing framework required by secure multi-party computing and realizes the calculation of secret input data;

[0047] S4. Each participant holds the additive slice of the secret data at the input and output of the modified activation function, so that the modified activation function can be connected to other linear modules or nonlinear modules in the neural network.

[0048] Optionally, the basic secure multi-party computation component includes one or more of a representation of a secret shared value, an element multiplication protocol under an addition sharing state, an addition shard conversion protocol under a multiplication sharing state, and a secure comparison protocol.

[0049] Optionally, the activation function is one or more of Leaky ReLU, Elu, Celu, Hardsigmoid, Hardtanh, SiLU, GeLU, and Mish, and the corresponding modified activation functions are Private Leaky ReLU, Private Elu, Private Celu, Private Hardsigmoid, Private Hardtanh, Private SiLU, Private GeLU, and Private Mish, respectively.

[0050] Optionally, for the Leaky ReLU activation function, the specific steps for implementing the Private Leaky ReLU function are as follows: when the neural network reaches the Leaky ReLU function calculation step, the participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participants, perform the following operations: using the secure comparison protocol, input <x>Get the output , at this time each participant holds Shard 1, 2,..., n ; Use the Beaver Triple protocol twice to find <0.1x*b> and<x*(1-b)> Sharding, after which each participant calculates locally <z> i =<b*0.1x> i +<x*(1-b)> i .

[0051] Optionally, for the Elu activation function, the specific steps for implementing the Private Elu function are: when the neural network reaches the Elu function calculation step, the participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed: Participant P i Local computing get for e x Multiplication slice, record All participants P i Together, we use the multiplication sharding protocol to convert it into the addition sharding protocol. x ] i Convert to <e x > i , then all participants complete the linear calculation locally and get <v>=a*(e <x>< / x> -1); using the secure comparison protocol, input <x>Get the output Finally, we use two BeaverTriple protocols and addition to get <z> = * <v> +(1- )* <x>.

[0052] Optionally, for the SiLU activation function, the specific steps for implementing the Private SiLU function are as follows: when the neural network reaches the SiLU function calculation step, the participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, do the following: call <x> i The PrivateSigmoid protocol is input, and the participant P i get i ,make 1+ 2+…+ n =sigmoid(x); using the Breaver Triple protocol, find<x*a> Shard, that is, participant P i get <z> i ,make <z> 1+ <z> 2+…+ <z> n =x*sigmoid(x).

[0053] Optionally, the specific execution steps of the Private Sigmoid protocol include: each participant generates a random number as a multiplication slice of the random number ρ, that is, participant P i Hold [ρ] i , satisfying [ρ]1*[ρ2]*...*[ρ] n =ρ, and then use the multiplication sharding to convert it into the addition sharding protocol to calculate <ξ> = [u], that is, at this time ξ = <ξ>1+<ξ>2+…<ξ> n =ρ; Participant P i Local computing After that, all participants convert [k] i Convert to <φ> i , that is, we get e x *Additive slice of ρ <φ> i ; Each participant P i Local calculation <χ> i =<φ> i -<ξ> i , and obtain χ through the reconstruction algorithm, then χ=φ-ξ=e x *ρ-ρ; each participant P i Local computing The output is <y> i .

[0054] Optionally, when performing privacy activation function calculation, the participants include one or more of a cloud server, a cloud storage device, and a distributed node.

[0055] Another aspect of the present invention provides a privacy-preserving neural network computing system, including:

[0056] The data holder is used to share the data secretly, split it into multiple additive shards, and send the shards to each participant;

[0057] The model holder is responsible for determining the neural network structure, training the neural network, obtaining the parameters of each layer, and publishing the parameters to the data holder;

[0058] Multiple participants are used to receive data shards sent by the data holder, and perform neural network inference calculations based on the neural network structure and parameters published by the model holder, using the new privacy activation function design method based on secure multi-party computing as described in the first aspect above, including linear layer calculations and nonlinear layer calculations, and finally send the calculation results back to the data holder.

[0059] Specifically, the solution of the present invention is further described according to the following examples:

[0060] The present invention's novel privacy activation function design method based on secure multi-party computation addresses the different structures and computational steps of several representative new activation functions in the field of artificial intelligence, utilizing the fundamental secure multi-party computation components, including the five protocols listed above. Table 1 summarizes the correspondence between the improved activation functions presented in this invention and the original functions:

[0061] Table 1

[0062] Existing activation functions The present invention proposes Leaky ReLU Private Leaky ReLU Elu Private Elu Celu Private Celu Hardsigmoid Private Hardsigmoid Hardtanh Private Hardtanh SiLU Private SiLU GeLU Private GeLU Mish Private Mish

[0063] The innovations of this method can be summarized as follows: (1) This method can maintain the mathematical operation steps and algebraic properties of each activation function, thereby inheriting the design advantages of each activation function. At the same time, based on secure multi-party computing technology, it realizes the calculation of secret input data; (2) This method maintains that each participant holds the additive slice of the secret data when the function is input and output, so that it can connect other linear or nonlinear modules in the neural network before and after; (3) The new activation function contains nonlinear function components, such as logarithmic functions, exponential functions, and reciprocal functions, while the classic secure multi-party computing protocol can only complete addition and multiplication calculations. Therefore, the secure multi-party computing of these nonlinear functions is still a challenge. In recent years, some scholars have proposed the idea of ​​using approximate fitting to transform the secure computing problem of nonlinear functions into the secure computing problem of low-order polynomials, but this method will cause loss of function accuracy. Different from this, the method proposed in this invention does not include approximation in the computing steps, and is an accurate calculation without loss of accuracy.

[0064] The specific design and implementation methods of each function are as follows:

[0065] Leaky Relu activation function

[0066] The Leaky ReLU activation function is a variant of the rectified linear unit (ReLU) activation function. It was first proposed by Maas et al. in their 2013 paper titled "Rectifier Nonlinearities Improve Neural Network Acoustic Models." This function sets the negative region of the ReLU function to a gradient slope. When the input x≤0, the output is 0.1x; when x>0, the output is x. Its mathematical expression can be expressed as:

[0067] The Leaky ReLU activation function was designed primarily to address the issue of zero gradients in the negative regions of the ReLU function. This helps prevent neuron "death," a phenomenon where some neurons are never activated during training. Because it has a constant slope, its gradient never reaches zero in the negative regions, potentially making neural network training more stable. Compared to traditional activation functions like Sigmoid or Tanh, the Leaky ReLU activation function is linear in most regions, reducing the risk of vanishing gradients. It is also computationally simple and generally performs well in practice.

[0068] This paper proposes a private leaky ReLU function, a privacy-preserving Leaky ReLU function that inherits the design advantages of the leaky ReLU function while enabling computation on secret input data based on secure multi-party computation techniques. Furthermore, this method ensures that each participant holds an additive slice of the secret data at both the input and output of the function, allowing it to be connected to other linear or nonlinear modules in the neural network.

[0069] Private Leaky ReLU

[0070] When the neural network reaches the step of calculating the Leaky ReLU function, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participants, and the following operations are performed:

[0071] (1) Using the secure comparison protocol, input <x>Get the output , at this time each participant holds Shard 1, 2,..., n

[0072] (2) Use the Breaver Triple protocol twice to find <0.1x*b> and<x*(1-b)> Sharding, after which each participant calculates locally <z> i =<b*0.1x> i +<x*(1-b)> i

[0073] Elu activation function

[0074] The ELU (Exponential Linear Unit) activation function was first proposed by Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter in their 2015 paper "Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)". Its definition is as follows:

[0075]

[0076] The advantages of the Elu activation function include: the Elu function and its derivatives are smooth over the entire domain, which is conducive to gradient calculation and model optimization; there is also a gradient in the negative part, and there will be no zero gradient problem, which helps to improve the convergence speed of the network and the learning efficiency of the model; compared with some traditional activation functions, Elu can better retain information when processing negative inputs and prevent neuron death.

[0077] This paper proposes the following Private Elu function, namely an Elu function with privacy protection. It inherits the design advantages of the Elu function and, based on secure multi-party computation technology, implements computation on secret input data. In addition, this method ensures that each participant holds an additive slice of the secret data at both the input and output of the function, allowing it to be connected to other linear or nonlinear modules in the neural network.

[0078] Private Elu

[0079] When the neural network reaches the Elu function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0080] (1) Participant P i Local computing At this time there Right now, for e x Multiplication slice, record

[0081] (2) All participants P i Together, we use the multiplication sharding protocol to convert it into the addition sharding protocol. x ] i Convert to <e x > i , and then each local completes the linear calculation to obtain <v>=a*(e <x>< / x> -1);

[0082] (3) Using the secure comparison protocol, input <x>Get the output

[0083] (4) Finally, using two Breaver Triple protocols and addition, we get:

[0084] <z> = * <v> +(1- )* <x>

[0085] Celu activation function

[0086] The Celu activation function (Continuously Differentiable Exponential Linear Units) is often used in hidden layers of neural networks to enhance the network's expressive power by introducing nonlinear transformations. It was first proposed by Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter in their 2015 paper, "Fastand Accurate Deep Network Learning by Exponential Linear Units (ELUs)." CELU is a variant of ELU, but is smoother due to its exponential nature on the negative axis. Its formula is as follows:

[0087]

[0088] The Celu activation function's advantage over other activation functions lies in its smoothness and favorable properties, making it more stable during training. Furthermore, Celu's properties help alleviate the vanishing gradient problem and are particularly effective for networks with large numbers of negative inputs. In practice, Celu can improve model generalization, reduce overfitting, and improve convergence speed during training.

[0089] This paper proposes a privacy-preserving Private Celu function, inheriting the design advantages of the Celu function while leveraging secure multi-party computation to enable computation on secret input data. Furthermore, this method ensures that each participant possesses an additive slice of the secret data at both the input and output of the function, allowing it to be connected to other linear or nonlinear modules in the neural network.

[0090] Private Celu

[0091] When the neural network reaches the CeLU function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1= <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0092] (1) Participant P i Local computing in is a constant; P i Continue calculation At this time, there are So in fact, for Multiplication slice, record

[0093] (2) All participants use the multiplication sharding to convert to the addition sharding protocol for calculation <y>=[u], that is, at this time

[0094] (3) Then calculate <z> =a*( <y>-1), and <z>As input, the secure comparison protocol is called to obtain

[0095] (4) Then, using the Breaver Triple protocol, we can get

[0096] (5) The calculation steps of max{0, x} are the same as those of the Private Relu function, that is, participant P i get <z> i ,make <z> 1+ <z> 2+…+ <z> n =max{0,x}

[0097] (6) Participant P i The final result is obtained by adding the slices of the calculation results obtained in steps (4) and (5) locally, that is,

[0098] Hardsigmoid activation function

[0099]

[0100] The characteristic of this activation function is that the function outputs 0 when the input is less than -3, outputs 1 when it is greater than 3, and outputs according to the linear relationship slope*x+offset when the input is between -3 and 3. Here, slope is the coefficient (slope), x is the input value, and offset is the bias term (constant term).

[0101] The Hardsigmoid function is an approximate, computationally simple activation function commonly used in neural networks. It was first introduced and used by Krizhevsky et al. in the AlexNet model in 2013. The Hardsigmoid function is computationally simple, primarily based on linear operations and threshold truncation, making it suitable for resource-limited environments such as embedded systems, mobile devices, and edge computing. Furthermore, due to its simplicity, the Hardsigmoid function is generally more stable in numerical computations and less prone to exploding or vanishing gradients.

[0102] This paper proposes a private Hardsigmoid function, a privacy-preserving Hardsigmoid function that inherits the design advantages of the Hardsigmoid function while enabling computation on secret input data based on secure multi-party computation techniques. Furthermore, this method ensures that each participant holds an additive slice of the secret data at both the input and output of the function, allowing it to be connected to other linear or nonlinear modules in the neural network.

[0103] Private Hardsigmoid

[0104] When the neural network reaches the Hardsigmoid function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0105] (1) Calculation by all parties <y> = <x>+3 shards and <z> = <x>-3 shards;

[0106] (2) Using the secure comparison protocol, input <y>Get the output Shard 1+ 2,..., n ; At this time, when y≤0, that is, when x≤-3, b=1, on the contrary, when x>-3, b=0

[0107] (3) Using the security comparison protocol again, input <z>Get the output <c>Shard <c> 1, <c> 2,..., <c> n ; At this time, when z≤0, that is, when x≤3, c=1, on the contrary, when x>3, c=0;

[0108] (4) Using Breaver Triple protocol to obtain< / c> < / c> < / c> < / c> < / z> < / y> < / x> < / z> < / x> < / y> < / x> < / x> < / x> < / x> < / z> < / z> < / z> < / z> < / z> < / y> < / z> < / y> < / x> < / x> < / x> < / x> < / x> < / v> < / z> < / x> < / v> < / x> < / x> < / x> < / x> < / z> < / x> < / x> < / x> < / x> < / x> < / y> < / z> < / z> < / z> < / z> =(1- )* <c>, at this time, we have: when -3<x≤3, a=1, otherwise, a=0;

[0109] (5) Using one local multiplication, one local addition, one Breaver Triple protocol and one more local addition, we can get <z> = <c>+(slope* <x> +offst)*< / x> < / c> < / z> < / c> .

[0110] Hardtanh activation function

[0111] The Hardtanh activation function was first proposed by Yann LeCun in 2015. It is a simple activation function defined as:

[0112]

[0113] The Hardtanh function is a simple and intuitive activation function that requires no additional parameters and is easy to implement and calculate. Furthermore, its output range falls within a nonlinear range, effectively handling the control of the activation output range and helping neural network models learn complex data patterns. Furthermore, because Hardtanh is a hard saturation activation function, it can be used to clip gradients, helping to prevent exploding gradients and improving network stability and convergence.

[0114] This paper proposes a privacy-preserving Private Hardtanh function, inheriting the design advantages of the Hardtanh function while leveraging secure multi-party computation to enable computation on secret input data. Furthermore, this method ensures that each participant possesses additive slices of the secret data at both the input and output of the function, enabling connections to other linear or nonlinear modules in the neural network.

[0115] Private Hardtanh

[0116] When the neural network reaches the Hardtanh function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0117] (1) Calculation by all parties <y> = <x>-max shards and <z>=min- <x>Sharding;

[0118] (2) Using the secure comparison protocol, input <y>Get the output Shard 1, 2,..., n ; At this time, when y < 0, that is, when x < max, b = 1, and conversely when x ≥ max, b = 0;

[0119] (3) Use the secure comparison protocol again, input <z>Get the output <c>Shard <c> 1, <c> 2,..., <c> n ; At this time, when z < 0, that is, when x > min, c = 1, on the contrary, when x ≤ min, c = 0;

[0120] (4) Using Breaver Triple protocol to obtain< / c> < / c> < / c> < / c> < / z> < / y> < / x> < / z> < / x> < / y> < / x> < / x> < / x> < / x> = * <c>, at this time, we have: when min<x<max, a=1, otherwise, a=0;

[0121] (5) Using basic protocols such as Breaver Triple, calculate <z>=max*(1- )+min*(1- <c> )+ <x> *< / x> < / c> < / z> < / c> .

[0122] SiLU activation function

[0123] The SiLU (Sigmoid Linear Unit) activation function, also known as the Swish function, was first proposed by Belgian researchers Dan Hendrycks and Kevin Gimpel in a 2016 paper titled "Gaussian Error Linear Units (GELUs)". It is defined as follows:

[0124] SiLU(x)=x*sigmoid(x)

[0125] Among them, the sigmoid function is the S-type function (Logistic function) defined as:

[0126]

[0127] As a new type of continuously differentiable nonlinear activation function, the SiLU activation function combines the characteristics of the Sigmoid function and the linear function. It has excellent smoothness properties. Its gradient is continuous relative to the input and does not suffer from the vanishing gradient problem on the negative semi-axis like the ReLU function. This is beneficial for the optimization and training of neural networks. At the same time, the SiLU activation function is a parameterized function. The shape of the activation function can be controlled by adjusting the parameters, which helps to adapt to different types of data and task characteristics.

[0128] The present invention proposes the following Private SiLU function, that is, a SiLU function with privacy protection function, which inherits the design advantages of the SiLU function and realizes the calculation of secret input data based on secure multi-party computing technology. In addition, this method maintains that each participant holds the additive slice of the secret data during the input and output of the function, so that it can be connected to other linear or nonlinear modules in the neural network. The concept of using the secure multi-party computing principle to implement the privacy protection SiLU function is mentioned in the patent "A Privacy Protection Inference Method and System for Diffusion Model Sampling" proposed by Chen Xiaojun et al., but the core idea of ​​the method used by Chen Xiaojun et al. is to use Chebyshev polynomials to fit the exponential function e x , then decomposes the SiLU function into a combination of linear operations, and directly uses the secure multi-party computing basic protocol to complete the operation. This approximate calculation method will cause a loss of precision, while the method designed by the present invention is a precise calculation, without loss of precision in the execution steps.

[0129] Private SiLU

[0130] When the neural network reaches the SiLU function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0131] (1) Call <x> i The Private Sigmoid protocol is input, and the participant P i get< / x> < / x> < / x> < / x> < / x> i ,make 1+ 2+…+ n =sigmoid(x)

[0132] (2) Using the Breaver Triple protocol, find<x*a> Shard, that is, participant P i get <z> i ,make <z> 1+ <z> 2+…+ <z> n =x*sigmoid(x)

[0133] The Private Sigmoid protocol mentioned in step (1) was proposed in 2020 by Qi Feng, Debiao He, Zhe Liu, Huaqun Wang, and Kim-Kwang Raymond Choo in the paper “SecureNLP: A System for Multi-Party Privacy-Preserving Natural Language Processing”. Its input is the addition shard of x held by each participant, i.e., the participant P i hold <x> i ,satisfy <x> 1+ <x> 2+…+ <x> n =x, perform the following steps:

[0134] (1) Each participant generates a random number as a multiplication slice of the random number ρ, that is, participant P i Hold [ρ] i , satisfying [ρ]1*[ρ]2*...*[ρ] n =ρ, and then use the multiplication sharding to convert it into the addition sharding protocol to calculate <ξ> = [u], that is, at this time ξ = <ξ>1+<ξ>2+…+<ξ> n =ρ

[0135] (2) Participant P i Local computing After that, all participants convert [k] i Convert to <φ> i , that is, we get e x *Additive slice of ρ <φ> i

[0136] (3) Each participant P i Local calculation <χ> i =<φ> i -<ξ> i , and obtain χ through the reconstruction algorithm, then χ=φ-ξ=e x *ρ-ρ

[0137] (4) Each participant P i Local computing The output is <y> i

[0138] GeLU activation function

[0139] The GeLU (Gaussian Error Linear Units) activation function was published by Belgian researchers Dan Hendrycks and Kevin Gimpel in their 2016 paper Gaussian Error Linear Units (GELUs). It is defined as follows:

[0140]

[0141] Among them, the error function erf is a special mathematical function that calculates the integral value from negative infinity to x under the normal distribution curve, that is, the probability that the variable falls within a certain range. Its form is: If we use approximate calculations, we have:

[0142]

[0143] The activation function is a smooth function that does not suffer from the gradient saturation problem of functions like ReLU at large or small input values, which facilitates the stability of backpropagation and gradient descent. As a nonlinear activation function, GeLU enables neural networks to learn complex patterns and features and is widely used in some neural network architectures. GeLU has been widely used in various deep learning models, particularly in the field of natural language processing (NLP). For example, models such as BERT and GPT-2 use the GeLU activation function.

[0144] The present invention proposes the following Private GeLU function, that is, a GeLU function with privacy protection function, which inherits the design advantages of the GeLU function and, based on secure multi-party computing technology, realizes the calculation of secret input data. In addition, this method maintains that each participant holds the additive slice of the secret data during the input and output of the function, so that it can connect other linear or nonlinear modules in the neural network before and after. The patent "A Privacy Protection Entity Identification Tool Based on Secure Multi-Party Computing Technology" proposed by Li Mu et al. involves a secure multi-party computing method for the GeLU function, but there are two differences from the solution proposed in the present invention: one is that it sets the approximation of the GELU function to x*sigmoid (1.702x), which is different from the present invention; the other is that it uses the Newton-Raphson iterative method to optimize the polynomial approximation of the reciprocal, while the present invention does not use approximation in the calculation process, reducing the loss of precision.

[0145] Private GeLU

[0146] When using approximate calculation, when the neural network reaches the GeLU function calculation step, the participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0147] (1) Using the Breaver Triple protocol, find<x*x> Shard, that is, participant P i Get <x 2 > i ,make <x 2 >1+ <x 2 >2+…+ <x 2 > n =x*x=x 2

[0148] (2) Using the Breaver Triple protocol again, <x 2 *x> Shard, i.e. participant P i get <x 3 > i ,make <x 3 >1+ <x 3 >2+…+ <x 3 > n =x 2 *x=x 3

[0149] (3) Afterwards, each participant P i Compute locally

[0150] (4) Calling the Private Tanh protocol participant P i get i ,make:

[0151] 1+ 2+…+ n =Fish < / x> < / x> < / x> < / x> < / y> < / x> < / x> < / x> < / x> < / z> < / z> < / z> < / z> i

[0152] (5) Afterwards, each participant P i Compute locally <c> i =0.5*( i +1)

[0153] (6) Finally, using the Breaver Triple protocol, we can find<x*c> Shard, that is, participant P i get <z> i ,make:

[0154]

[0155] The Private Tanh protocol mentioned in step (4) is consistent with the aforementioned Private Sigmoid protocol, both of which were proposed in the paper "SecureNLP: A System for Multi-Party Privacy-Preserving Natural Language Processing". Its input is the addition slice of x held by each participant, that is, participant P i hold <x> i ,satisfy <x> 1+ <x> 2+…+ <x> n =x, perform the following steps:

[0156] (1) Each participant generates a random number as a multiplication slice of the random number ρ, that is, participant P i Hold [ρ] i , satisfying [ρ]1*[ρ]2*...*[ρ] n =ρ, and then use the multiplication sharding to convert it into the addition sharding protocol to calculate <ξ> = [u], that is, at this time ξ = <ξ>1+<ξ>2+…+<ξ> n =ρ

[0157] (2) Participant P i Local computing After that, all participants convert [k] i Convert to <φ> i , that is, we get e 2x *Additive slice of ρ <φ> i

[0158] (3) Each participant P i Local calculation <χ> i =<φ> i +<ξ> i , and obtain X through the reconstruction algorithm, then χ=φ+ξ=e x *ρ+ρ

[0159] (4) Each participant P i Local computing The output is <y> i

[0160] Mish activation function

[0161] The Mish activation function was first published by Indonesian researcher Diganta Misra in his 2019 paper "Mish: ASelf Regularized Non-Monotonic Neural Activation Function". It is defined as:

[0162] Mish(x)=x*tanh(ln(1+e x ))

[0163] Among them, the tanh(x) function is the hyperbolic tangent function, which is defined as follows:

[0164]

[0165] The Mish activation function combines smoothness and non-monotonicity, which makes it superior to traditional activation functions such as ReLU in some cases. At the same time, the Mish function has a certain degree of regularization, which helps to improve the generalization ability of the model and reduce the risk of overfitting.

[0166] The present invention proposes the following Private Mish function, i.e., a Mish function with privacy protection function, which inherits the design advantages of the Mish function and, based on secure multi-party computing technology, realizes the calculation of secret input data. In addition, this method maintains that each participant holds the additive slice of the secret data during the input and output of the function, so that it can connect to other linear or nonlinear modules in the neural network. Similar to the SiLU function, in the patent "A Privacy Protection Inference Method and System for Diffusion Model Sampling" proposed by Chen Xiaojun et al., the exponential function e is fitted using Chebyshev polynomials. x The method can also be used to calculate the Mish function, but this approximate calculation method will cause a loss of precision. The method designed in the present invention is an accurate calculation without loss of precision.

[0167] Private Mish

[0168] When the neural network reaches the Mish function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:

[0169] (1) Each participant generates a random number as a multiplication slice of the random number ρ, that is, participant P i Hold [ρ] i , satisfying [ρ]1*[ρ]2*...*[ρ] n =ρ, and then use the multiplication sharding to convert it into the addition sharding protocol to calculate <ξ> = [u], that is, at this time ξ = <ξ>1+<ξ>2+…+<ξ> n =ρ

[0170] (2) Participant P i Local computing After that, all participants convert [k] i Convert to <φ> i , that is, we get e x *Additive slice of ρ <φ> i

[0171] (3) Each participant P i Local calculation <χ> i =<φ> i +<ξ> i , and obtain χ through the reconstruction algorithm, then χ=φ+ξ=e x *ρ+ρ, and all parties can directly calculate ln(χ)

[0172] (4) Participant P i Compute ln([ρ] locally i ), at this time there is remember <y> i = ln([ρ] i ), then <y> i For participant P i The added shard of ln(ρ) held by each participant, at the same time, can calculate the added shard of ln(χ)-ln(ρ), which is recorded as <v> i ,at this time

[0173]

[0174] (5) <v> i As input, call the Private Tanh protocol participant P i get i ,make 1+ 2+…+ n =Fish <v> i

[0175] (6) Finally, using the Breaver Triple protocol, we can find<x*u> Shard, that is, participant P i get <z> i ,make <z> 1+ <z> 2+…+ <z> n = x * u = x * tanh(ln(1 + e x ))

[0176] Like< / z> < / z> < / z> < / z> < / v> < / v> < / v> < / y> < / y> < / x> < / x> < / x> < / x> < / y> < / x> < / x> < / x> < / x> < / z> < / c> Figure 1 As shown, the activation function described in the present invention is applied to the reasoning stage of the neural network, and is used by the data holder to complete the reasoning task (such as classification or prediction) while ensuring that the data held is invisible to the model holder (and other untrusted participants).

[0177] Specifically, the construction and use of neural networks correspond to the training phase and the inference phase respectively. The training phase is the process of the neural network inputting data and continuously adjusting the weights through the data and corresponding labels, that is, the process of generating a model; the inference phase is when the neural network model is trained (the weight parameters have been determined), the neural network processes the input data (for example, when a picture of a cat is input, the neural network classifies the picture as a cat), that is, the model starts working to perform tasks such as regression, prediction, and classification.

[0178] A trained neural network model has the following characteristics: Figure 2 The structure shown includes an input layer, an output layer, and a hidden layer (also called an implicit layer), wherein the hidden layer further includes a linear layer and a nonlinear activation function layer. In the inference stage, the parameters in the neural network model, such as the matrices W and B in the linear layer, have been determined in the training stage. By multiplying the data input to each linear layer with the weight matrix W of each layer and adding the bias term B, the linear layer can recombine the features of the input data and map it to a new representation space of different dimensions, and perform more complex classification and regression tasks in this space. At the same time, matrix operations can also accelerate the calculation process of the neural network through parallel computing, so that the neural network can process large-scale data sets faster; the nonlinear layer contains one or several types of activation functions, which are used to enhance the network's expressive power (fitting ability) and increase the network's stability and convergence speed, denoted as the F function, which can be selected as ReLU, Leaky ReLU, Elu, Celu, Hardsigmoid, Hardtanh, SiLU, GeLU, Mish and other functions. Before training, which activation function or functions are used in the model is already determined. In the inference stage, these activation functions are still determined. In addition, the calculation of inputting data into the function F in the form of a vector refers to using the function F to calculate each dimension of the input vector separately, that is, Z = F(Y) means {z1, z2, ..., z n }={F(y1), F(y2),…,F(y n The number and structure of linear and nonlinear layers in the hidden layer are customized according to the complexity and characteristics of the specific problem. Each layer receives input from the previous layer and passes the output to the next layer until the last layer is completed.

[0179] For the input and output layers, consider the image classification task of handwritten digit recognition, which is used to classify ten digits from 0 to 9. Specifically, given an image of a handwritten digit, the user is tasked with identifying the digit within it. During the inference phase, the data holder has an image, a matrix with dimensions of 28 (rows) × 28 (columns). This is flattened into a column vector of 28 × 28 = 784 rows and 1 column, and then fed into the neural network's input layer. After a series of hidden layer calculations, the results, in the form of vectors, are passed to the output layer, the final layer of the neural network. The output layer is designed based on the specific task, with the computational steps matching the output dimensions and requirements of the task. For example, for the handwritten digit recognition task, the output layer has 10 dimensions, with each dimension's output corresponding to a digit category. The output value represents the probability of that category, also calculated using an activation function.

[0180] See also Figure 3 In the present invention, the model holder first determines the neural network structure, including the number of layers, the dimension of each layer, the type of activation function used in the nonlinear layer, and other information, and then uses the training data to train the neural network, that is, to obtain the parameters W and B of each layer, and then publishes all W and B to the data holder.

[0181] After the inference phase begins, the data holder has a picture (still taking the handwritten digit recognition task as an example), he splits the picture into n additive slices and sends them to n participants P1, P2, ..., P n , for each participant P i , the slices it holds are still arranged in the same dimensional form as the input vector and are received by the input layer, ready to be passed to the hidden layer of the neural network. For example, for n = 2, the image held by the data holder is the following matrix (for ease of description, it is simplified to a 3 (number of rows) × 3 (number of columns) matrix): After flattening, it becomes X = [0.444 0.571 0.165 0.67 0.828 0.543 0.592 0.094 0.068]. Then, the data holder splits X into two shards:

[0182] <x> 1=[0.142 0.288 0.025 0.754 0.752 0.001 0.763 0.826 0.258]

[0183] <x>2 = [0.302 0.283 0.14 -0.084 0.075 0.542 -0.171 -0.733 -0.19]

[0184] And split <x> 1、 <x>2 are sent to participants P1 and P2 respectively. At this time, the shard <x>1 and <x>The sum of the data at each corresponding position in 2 is equal to the data at the corresponding position in X, while participants P1 and P2 can only see a shard in the form of a string of random numbers. <x>1 or <x>2. The original data X has not been leaked.

[0185] Next, we enter the calculation of the linear layer in the hidden layer. In this step, participants P1 and P2 complete the calculation alone without communication or interaction. Specifically, P1 calculates <y> 1=W* <x>1+B1, P2 calculation <y> 2=W* <x>2+B2, where B1 and B2 satisfy B1+B2=B and are pre-calculated by the data holder (e.g., B1=B, B2=0) and distributed to P1 and P2. <y>1 and <y>2 is still in vector form (but due to multiplication with the weight matrix W, <y>1 or <y>The dimension of 2 may be different from <x>1 or <x>2), and since matrix operations have linear properties, we have <y> 1+ <y> 2=W* <x>1+B1+W* <x>2+B2=W*X+B, let Y=W*X+B. That is, at this time, participants P1 and P2 respectively hold the addition slices of Y, and the sum Y obtained by adding them is consistent with the intermediate calculation result of the ordinary neural network inference stage without privacy protection strategy. However, in the present invention, each participant only holds the addition slice in the form of random numbers during the calculation process, and the original data (intermediate result) Y is not leaked.

[0186] Afterwards, the shards held by participants P1 and P2 <y>1 and <y>2 Enter the calculation of the nonlinear layer in the hidden layer. Specifically, for a neural network without privacy protection strategy, the purpose of the nonlinear layer is to calculate Z = F (Y), where F is a nonlinear activation function, and the calculation of Z is based on F for each coordinate position, that is, Z = {z1, z2, ..., z n }={F(y1), F(y2),..., F(y n )}. For the present invention, the purpose of this step is to have participants P1 and P2 hold <z>1 and <z>2. Satisfaction <z> 1+ <z>2 = Z. Specifically, each nonlinear activation function F needs to be replaced with the privacy-preserving activation function implementation proposed in this invention. That is, where ReLU, Leaky ReLU, Elu, Celu, Hardsigmoid, Hardtanh, SiLU, GeLU, and Mish appear in the original neural network, participants P1 and P2 apply the PrivateReLU, Private Leaky ReLU, Private Elu, Private Celu, Private Hardsigmoid, Private Hardtanh, Private SiLU, Private GeLU, and Private Mish function pairs described in this invention to the input slices. <y>1 and <y>The calculation is completed for each element in 2. The calculation steps are shown in the following example (only the representative Leaky ReLU and Elu are used as examples):

[0187] Example 1:

[0188] The computing parties are participants P1 and P2, who perform the Leaky ReLU function calculation while keeping the input data confidential. The following steps are performed according to the process:

[0189] (1) At the beginning of the system, participant P1 has <x>1 = -1.5, participant P2 has <x>2=0.5, at this time x= <x> 1+ <x>2=-1, x is unknown to both participants P1 and P2.

[0190] (2) Participant P1 calculates locally Participant P2 calculates locally At this time there is e x =[e x ]1*[e x ]2=u=0.367874。

[0191] (3) Using the multiplication sharding to convert to the addition sharding protocol (see Example 3), participant P1 converts the multiplication shard [e x ]1Convert to additive slice <e x >1=0.0315094, participant P2 will multiply the shards [e x ]2 Convert to additive slice:

[0192] <e x >2=0.336349

[0193] Then each participant completes the linear calculation locally, and participant P1 obtains Participant P2 receives:

[0194]

[0195] At this time there <v>=a-(e <x>< / x> -1)= <v> 1+ <v>2=-1.26428.

[0196] (4) Using the secure comparison protocol, since x>0, b=0, participant P1 now has 1=389129223699150112, participant P2 now has 2 = -389129223699150111, satisfied 1+ 2=1.

[0197] (5) Finally, using the Breaver Triple protocol twice, the final result can be obtained by participant P1:

[0198] <z> 1= 1* <v> 1+(1- 1)* <x>1=112507553679661.137237548828125

[0199] Participant P2 receives:

[0200] <z> 2= 2* <v> 2+(1- 2)* <x>2 = -112507553679662.401519775390625

[0201] The result satisfies z= <z> 1+ <z>2=-1.2642822265625, the initial x=-1 can also be obtained through Elu function operation to obtain z≈-1.264.

[0202] Example 2:

[0203] The calculation parties are participants P1 and P2. They calculate the Elu function when a=2 while keeping the input data confidential. The following operations are performed according to the process steps:

[0204] (1) At the beginning of the system, participant P1 has <x>1 = -1.5, participant P2 has <x>2=0.5, at this time x= <x> 1+ <x>2=-1, x is unknown to both participants P1 and P2.

[0205] (2) Participant P1 calculates locally Participant P2 calculates locally At this time there is e x =[e x ]1*[e x ]2=u=0.367874。

[0206] (3) Using the multiplication sharding to convert to the addition sharding protocol (see Example 3), participant P1 converts the multiplication shard [e x ]1Convert to additive slice <e x >1=0.0315094, participant P2 will multiply the shards [e x ]2 Convert to additive slice <e x >2=0.336349, and then each participant completes the linear calculation locally, and participant P1 obtains:

[0207]

[0208] Participant P2 receives:

[0209] At this time there are: <v>=a*(e <x>< / x> -1)= <v> 1+ <v>2=-1.26428.

[0210] (4) Using the secure comparison protocol, since x>0, b=0, participant P1 now has 1=389129223699150112, participant P2 now has 2 = -389129223699150111, satisfied 1+ 2=1.

[0211] (5) Finally, using the Breaver Triple protocol twice, the final result can be obtained by participant P1:

[0212] <z> 1= 1* <v> 1+(1- 1)* <x>1=112507553679661.137237548828125

[0213] Participant P2 receives:

[0214] <z> 2= 2* <v> 2+(1- 2)* <x>2 = -112507553679662.401519775390625

[0215] The result satisfies z= <z> 1+ <z>2=-1.2642822265625, the initial x=-1 can also be obtained through Elu function operation to obtain z≈-1.264.

[0216] Example 3:

[0217] This example shows the execution steps and results of the "element multiplication conversion protocol in the addition sharing state", which is used in step (3) of Example 2. The computing parties are participants P1 and P2.

[0218] (1) At the initialization of the system, participant P1 has [m]1 = 56.53496887348917, and participant P2 has [m]2 = 1.7688167516953, where [m]1 and [m]2 are both multiplication slices of K, that is, [m]1*[m]2 = K = 100, and the value of K is not disclosed.

[0219] (2) After the conversion protocol operation is completed, participant P1 has <m>1=14.201186127109288, participant P2 has <m>2 = 85.79881387289072, where <m>1 and <m>2 are all additive slices of K, that is, at this time there are <m> 1+ <m>2=K=100.

[0220] After multi-layer hidden layer operations, the data held by participants P1 and P2 are <z>1 and <z>2. This data is input into the neural network's output layer for computation. The linear and nonlinear operations involved in this computation follow the same principles as those for the hidden layer. This means that the two participants, P1 and P2, hold the computational results of this layer in the form of added slice vectors. P1 and P2 then send their respective computational results back to the data holder, who adds them together to obtain the desired neural network inference output. At this point, only the data holder knows the computational results; neither the participants nor the model holder can see the true value of the computational results.

[0221] The present invention has the following two application scenarios:

[0222] (1) Sensitive information protection: Using neural networks usually requires inputting data in plain text. However, this data often contains sensitive personal information such as medical records and financial information, as well as internal company information such as financial information and source code. In this case, using traditional data security methods, such as data desensitization or anonymization, will sacrifice some data accuracy, resulting in the inability to effectively utilize the data information and loss of availability. In this case, the data holder shares the data secretly, splits it into additive shards, and stores each shard at different participants. Participants include devices, systems, organizations, etc., which can be cloud servers, cloud storage devices, distributed nodes, etc. Then, using the privacy neural network model, through secure multi-party computing, these participants can jointly perform neural network calculations without having to disclose the data shards they hold, thereby achieving data privacy protection while completing the computing task.

[0223] (2) Privacy computing: Traditional machine learning technologies require data from decentralized data sources (such as edge devices, different institutions, etc.) to be aggregated and centralized in data centers or machines. The privacy-preserving neural network computing involved in the present invention can facilitate the model's utilization of decentralized data sources without centralizing or disclosing the original data. Specifically, each decentralized data source is regarded as a participant, and the data it holds is used as a secret shard. The neural network structure and parameters required for the computing task are published to each participant. The participants complete the computing task of the neural network through local computing and collaborative computing, and the data shards of the computing results are combined to obtain the required result of the task.

[0224] The method of the present invention enables all parties to share data and jointly use the model for calculations without worrying about leaking sensitive information contained in the complete data, making the data "available but invisible." Specific application scenarios include: In the financial field, the model described in the present invention can be used for data sharing and protection during cross-organizational collaboration, solving the problem of data silos, and helping banks, insurance companies and other institutions perform risk assessments, credit approvals and other operations while protecting user privacy; In the medical field, the model described in the present invention can be used for the analysis and processing of patient data, protecting patient privacy, and conducting disease prediction, drug development and other research without disclosing patient personal information; In the field of government services, the model described in the present invention can be used for big data analysis, providing public services while protecting personal privacy, and collaboratively scheduling multi-party sharing without leaking original data.

[0225] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.< / z> < / z> < / m> < / m> < / m> < / m> < / m> < / m> < / z> < / z> < / x> < / v> < / z> < / x> < / v> < / z> < / v> < / v> < / v> < / x> < / x> < / x> < / x> < / z> < / z> < / x> < / v> < / z> < / x> < / v> < / z> < / v> < / v> < / v> < / x> < / x> < / x> < / x> < / y> < / y> < / z> < / z> < / z> < / z> < / y> < / y> < / x> < / x> < / y> < / y> < / x> < / x> < / y> < / y> < / y> < / y> < / x> < / y> < / x> < / y> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / v> < / z> < / x> < / v> < / x> < / x> < / x> < / x> < / z> < / x> < / x> < / x> < / x> < / x> < / y> < / x> < / v>

Claims

1. A novel privacy activation function design method based on secure multi-party computation, characterized by: include: Analyze the activation function in the neural network and determine the mathematical operation steps and algebraic properties of the activation function; Selecting corresponding basic secure multi-party computation components based on the mathematical operation steps and algebraic properties; The activation function is modified by the selected basic secure multi-party computing components to adapt the modified activation function to the computing framework required by secure multi-party computing and realize the calculation of secret input data; Each participant holds the additive slice of the secret data at the input and output of the modified activation function, so that the modified activation function can be connected to other linear modules or nonlinear modules in the neural network.

2. The method according to claim 1, characterized in that The basic secure multi-party computing component includes one or more of a representation of a secret shared value, an element multiplication protocol under an addition sharing state, an addition shard conversion protocol under a multiplication sharing state, and a secure comparison protocol.

3. The method according to claim 1, characterized in that The activation function is one or more of Leaky ReLU, Elu, Celu, Hardsigmoid, Hardtanh, SiLU, GeLU, and Mish, and the corresponding modified activation functions are Private Leaky ReLU, Private Elu, Private Celu, Private Hardsigmoid, Private Hardtanh, Private SiLU, Private GeLU, and Private Mish, respectively.

4. The method according to claim 3, characterized in that For the Leaky ReLU activation function, the specific steps to implement the PrivateLeaky ReLU function are: When the neural network reaches the step of calculating the Leaky ReLU function, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:< / x> < / x> < / x> < / x> Using the secure comparison protocol, input <x>Get the output , at this time each participant holds Shard 1, 2,..., n ; < / x> The Beaver Triple protocol was used twice to find <0.1x*b> and<x*(1-b)> Sharding, after which each participant calculates locally <z> i =<b*0.1x> i +<x*(1-b)> i 。< / z> 5. The method according to claim 3, characterized in that For the Elu activation function, the specific steps to implement the Private Elu function are: When the neural network reaches the Elu function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:< / x> < / x> < / x> < / x> Participant P i Local Computing <x> i ,get e <x> i for e x Multiplication slice, record < / x> < / x> All participants P i Together, we use the multiplication sharding protocol to convert it into the addition sharding protocol. x ] i Convert to <e x > i , then all participants complete the linear calculation locally and get <v>=a*(e <x>< / x> -1);< / v> Using the secure comparison protocol, input <x>Get the output ; < / x> Finally, using two Beaver Triple protocols and addition, we get <z> = * <v> +(1- )* <x> 。< / x> < / v> < / z> 6. The method according to claim 3, characterized in that For the SiLU activation function, the specific steps to implement the Private SiLU function are: When the neural network reaches the SiLU function calculation step, participant P i Holds the addition slice of the function input x <x> i ,Right now <x> 1+ <x> 2+…+ <x> n = x, where x is a secret value for the participant, and the following operations are performed:< / x> < / x> < / x> < / x> Call with <x> i The Private Sigmoid protocol is input, and the participant P i get i ,make 1+ 2+…+ n =sigmoid(x); < / x> Using the Breaver Triple protocol,<x*a> Shard, that is, participant P i get <z> i ,make <z> 1+ <z> 2+…+ <z> n =x*sigmoid(x)。< / z> < / z> < / z> < / z> 7. The method according to any one of claims 1 to 6, characterized in that The specific execution steps of the Private Sigmoid protocol include: Each participant generates a random number as a multiplication slice of the random number ρ, that is, participant P i Hold [ρ] i , satisfying [ρ]1*[ρ]2*...*[ρ] n =ρ, and then use the multiplication sharding to convert it into the addition sharding protocol to calculate <ξ> = [u], that is, at this time ξ = <ξ>1+<ξ>2+…+<ξ> n =ρ; Participant P i Local computing After that, all participants convert [k] i Convert to <φ> i , that is, we get e x *Additive slice of ρ <φ> i ; Each participant P i Local calculation <χ> i =<φ> i -<ξ> i , and obtain χ through the reconstruction algorithm, then χ=φ-ξ=e x *ρ-ρ; Each participant P i Local computing The output is <y> i 。< / y> 8. The method according to any one of claims 1 to 7, characterized in that When calculating the privacy activation function, the participants include one or more cloud servers, cloud storage devices, and distributed nodes.

9. A privacy-preserving neural network computing system, comprising a data holder, a model holder, and multiple participants, characterized in that: The data holder is used to share the data secretly, split it into multiple additive shards, and send the shards to each participant; The model holder is responsible for determining the neural network structure, training the neural network, obtaining the parameters of each layer, and publishing the parameters to the data holder; Multiple participants are used to receive data shards sent by the data holder, and perform neural network inference calculations based on the neural network structure and parameters published by the model holder using the novel privacy activation function design method based on secure multi-party computing as described in any one of claims 1-8, including linear layer calculations and nonlinear layer calculations, and finally send the calculation results back to the data holder.

Citation Information

Patent Citations

  • Privacy protection neural network reasoning method based on three-party server

    CN117744798A

  • Nonlinear activation function security calculation method for neural network

    CN118133329A

  • Systems and methods for providing a multi-party computation system for neural networks

    US20230049860A1

Cited By

  • Welding process tracing method based on multi-source data fusion

    CN120725545A