Unmanned aerial vehicle data protection method based on verifiable privacy protection joint learning
By adopting a joint learning method based on verifiable privacy protection in the drone cluster, blinding the gradient and Lagrangian interpolation processing are carried out, ciphertext is generated and aggregated, the problem of shared gradient preservation and sensitive information and aggregation of the aggregation server is solved, and the security protection of privacy data and correct aggregation of gradients is achieved.
Patent Information
- Application Number
- CN202411934223.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, shared gradients still retain sensitive information from the training set, and malicious aggregation servers may return fake aggregation gradients, causing participants to be unable to effectively verify that their gradients are correctly aggregated.
Using a method based on verifiable privacy protection joint learning, the random integer sequence and seed sequence sent by the private key generator are obtained, blinded processing and Lagrangian interpolation processing are performed, ciphertext is generated and aggregated, and finally the correctness of the aggregation gradient is verified through decryption and interpolation functions.
It effectively protects the privacy data of the drone cluster, prevents the leakage of sensitive information, and ensures the correct aggregation of gradients through verification mechanisms, and prevents the forgery of the aggregation server.
Smart Images

Figure CN120012148A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Internet of Things, and in particular relates to a data protection method for a drone based on verifiable privacy-preserving joint learning. Background Art
[0002] As the intelligence level of UAV systems continues to improve, big data and big models are widely used in the fields of autonomous vehicles, crop inspection drones, medical robots, etc. to achieve intelligent collaborative control of large-scale UAV systems. However, the popularity of UAV systems and the increase in data size have brought security risks, including cyber attacks, malware injection, and data privacy. These security challenges pose a potential threat to the collaborative control of UAV systems and may even endanger public safety. Therefore, in order to cope with challenges such as software and hardware security, data privacy, and artificial intelligence security, it is crucial to take comprehensive and multi-level security measures.
[0003] In recent years, deep learning has made great progress thanks to its powerful analytical ability for big data. However, the traditional method of collecting data first and then learning it in a centralized manner on the terminal is not suitable for scenarios that are sensitive to the training set. Federated learning has received widespread attention. Federated learning only trains models by sharing the model gradients of participants without accessing the original training set. Therefore, it is applied to ensure the security of privacy data in drone clusters. However, existing research shows that shared gradients still retain sensitive information of the training set. In order to solve the problem of sensitive training sets being exposed through gradients, most researchers encrypt the gradients uploaded by participants in federated learning. Malicious aggregation servers may return forged aggregate gradients to prevent participants from training models or mislead participants to train incorrect models, and participants have no effective means to verify such forgery.
[0004] Therefore, providing a method to ensure the privacy data security of drone clusters has become an urgent problem to be solved. Summary of the invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a data protection method for a drone based on verifiable privacy-preserving joint learning.
[0006] The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0007] The present invention provides a data protection method for a drone based on verifiable privacy-preserving joint learning, the data protection method for the drone comprising:
[0008] Obtain m+1 positive integers, a first random integer sequence, a second random integer sequence, a first seed sequence, a second seed sequence, a third seed sequence, and a pseudo-random number sent by a private key generator, wherein the pseudo-random numbers of all participants are the same;
[0009] Based on the privacy training data set of the i-th participant, a loss function of the neural network model is obtained to determine the private gradient of the i-th participant's finite domain according to the private gradient of the i-th participant obtained by the loss function;
[0010] Based on the pseudo-random numbers of the first seed sequence and the second seed sequence, blinding the private gradient of the finite field of the i-th participant to obtain the blinded private gradient of the i-th participant;
[0011] Determine interpolation points based on m-1 preset parameters selected from the private gradient of the ith participant after blinding processing, the first random integer sequence, and pseudo-random numbers of the third seed sequence, and perform Lagrangian interpolation processing using a Lagrangian interpolation method to obtain a first interpolation function;
[0012] Based on the first interpolation function, generating a ciphertext according to the second random integer sequence;
[0013] Obtaining an aggregated ciphertext obtained by aggregating the ciphertexts of all participants;
[0014] A second interpolation function is determined according to a decrypted result obtained by decrypting the aggregated ciphertext and the second random integer sequence, and a decoding result is obtained based on the second interpolation function.
[0015] Optionally, obtaining m+1 positive integers, a first random integer sequence, a second random integer sequence, a first seed sequence, a second seed sequence, and a third seed sequence sent by a private key generator includes:
[0016] Obtain m+1 positive integers sent by the private key generator, where each of the m+1 positive integers is relatively prime to each other;
[0017] Obtaining a first random integer sequence and a second random integer sequence sent by the private key generator, wherein the first random integer sequence and the second random integer sequence have no intersection, and the first random integer sequence and the second random integer sequence each include m integers;
[0018] Obtain a first seed sequence, a second seed sequence, a third seed sequence and a pseudo-random number sent by the private key generator.
[0019] Optionally, based on the privacy training data set of the i-th participant, a loss function of the neural network model is obtained, so as to determine the private gradient of the finite field of the i-th participant according to the private gradient of the i-th participant obtained by the loss function, including:
[0020] Based on the neural network model, the loss function is obtained according to the privacy training data set of the i-th participant, and the loss function is expressed as:
[0021]
[0022] Among them, L f (D i ,M) is the loss function, M is the parameter of the neural network model, D i is the private training dataset of the ith participant, D i ={ <x j ,y j >||j=1,2,…,T}, T is the size of the private training dataset of the i-th participant, x j is the jth data in the privacy training dataset of the i-th participant, y j is the label of the jth data, f(x j ,M) is about x j and M is a univariate n-degree polynomial;
[0023] Based on the back propagation algorithm, the private gradient of the ith participant is obtained according to the loss function, and the private gradient of the ith participant is expressed as:
[0024]
[0025] Among them, w i is the private gradient of the ith participant, ▽L f is the loss function L f The derivative of *i D i A random subset of ;
[0026] The private gradient of the i-th participant is converted from the real number field to the finite field to obtain the private gradient of the i-th participant in the finite field.
[0027] Optionally, converting the private gradient of the i-th participant from the real number field to the finite field to obtain the private gradient of the i-th participant in the finite field includes:
[0028] Amplifying the private gradient of the i-th participant to obtain an amplified private gradient;
[0029] Rounding the amplified private gradient to obtain a rounded private gradient;
[0030] The rounded private gradient is converted into a finite field through finite field mapping to obtain the private gradient of the finite field of the i-th participant. The private gradient of the finite field of the i-th participant is expressed as:
[0031]
[0032] Among them, w i is the private gradient of the finite field of the i-th participant, l w is a preset integer, (·) round is the rounding operation, and ψ(·) is a mapping from the real number field to the finite field.
[0033] Optionally, based on the pseudo-random numbers of the first seed sequence and the second seed sequence, blinding the private gradient of the finite field of the i-th participant to obtain the blinded private gradient of the i-th participant includes:
[0034] Obtaining a first parameter according to the pseudo-random numbers of the first seed sequence and the second seed sequence;
[0035] The private gradient of the finite field of the i-th participant is blinded based on the first parameter to obtain the private gradient of the i-th participant after blinding. The private gradient of the i-th participant after blinding is expressed as:
[0036]
[0037] in, is the private gradient of the ith participant after blinding, w i is the private gradient of the finite field of the i-th participant, is the t-th seed subsequence.
[0038] Optionally, interpolation points are determined based on m-1 preset parameters selected from the private gradient of the i-th participant after blind processing, the first random integer sequence, and pseudo-random numbers of the third seed sequence, so as to perform Lagrangian interpolation processing using a Lagrangian interpolation method to obtain a first interpolation function, including:
[0039] Randomly generate m-1 preset parameters, the preset parameters satisfy in, is the private gradient of the ith participant after blinding, v i,j The jth preset parameter randomly generated for the i-th participant;
[0040] Determine the interpolation point by using the first random integer sequence as the abscissa of the interpolation point and the second parameter generated according to the pseudo-random number of the third seed sequence and the preset parameter as the ordinate of the interpolation point;
[0041] Based on the interpolation points, Lagrangian interpolation processing is performed using a Lagrangian interpolation method to obtain a first interpolation function.
[0042] Optionally, generating ciphertext according to the second random integer sequence based on the first interpolation function includes:
[0043] Based on the first interpolation function, obtaining a difference result according to the second random integer sequence;
[0044] The CRT function is executed on the difference result to obtain the ciphertext.
[0045] The aggregated ciphertext is obtained by adding the ciphertexts of all participants.
[0046] Optionally, determining a second interpolation function according to a decrypted result obtained by decrypting the aggregated ciphertext and the second random integer sequence, and obtaining a decoding result based on the second interpolation function includes:
[0047] Decrypting the aggregated ciphertext by modular operation to obtain a decrypted result;
[0048] Performing Lagrange interpolation on a data set consisting of the second random integer sequence and the decrypted result as interpolation points to obtain a second interpolation function;
[0049] Confirm whether the aggregated ciphertext is correct according to the interpolation result obtained by the mth element in the first random integer sequence and the second interpolation function, and if it is wrong, end; if it is correct, obtain the aggregate value of the original gradient according to the first random integer sequence and the second interpolation function;
[0050] Decoding is performed using the aggregated value of the original gradient to obtain a decoding result.
[0051] Optionally, obtaining an aggregate value of the original gradient according to the first random integer sequence and the second interpolation function includes:
[0052] Based on the second interpolation function, obtaining a calculation result according to each element in the first random integer sequence;
[0053] All the calculation results are added together to obtain the aggregate value of the original gradient.
[0054] Optionally, decoding is performed using the aggregate value of the original gradient to obtain a decoding result, including:
[0055] The aggregate value of the original gradient is converted into a floating point number to obtain a decoding result.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The data protection method for a drone provided by the present invention first obtains a loss function of a neural network model based on a privacy training data set of an i-th participant, so as to determine a private gradient of a finite field of the i-th participant based on the private gradient of the i-th participant obtained according to the loss function, and then blinds the private gradient of the finite field of the i-th participant based on pseudo-random numbers of a first seed sequence and a second seed sequence to obtain the private gradient of the i-th participant after blinding, and then determines an interpolation point based on m-1 preset parameters selected from the private gradient of the i-th participant after blinding, a first random integer sequence, and pseudo-random numbers of a third seed sequence, so as to perform Lagrangian interpolation processing using a Lagrangian interpolation method to obtain a first interpolation function, and then generates a ciphertext based on the first interpolation function and a second random integer sequence, and obtains an aggregated ciphertext obtained by aggregating the ciphertexts of all participants, and finally determines a second interpolation function based on a decrypted result obtained by decrypting the aggregated ciphertext and the second random integer sequence, and obtains a decoding result based on the second interpolation function. The present invention uses blinding technology to protect privacy gradients, solves the problem of exposing sensitive information in plaintext transmission of gradients, and uses Lagrange interpolation to set interpolation points, and uses the function value of the special point of the function restored by quadratic interpolation to verify the correctness of the aggregated gradient. Participants can perform verification operations on their own terminals to effectively verify whether their gradients are aggregated and prevent forgery by the aggregation server. The drone data protection method provided by the present invention can not only ensure the security of the privacy data of the drone cluster, but also realize the global model training of artificial intelligence applications.
[0058] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flowchart of a data protection method for a drone based on verifiable privacy-preserving joint learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0061] Embodiment 1
[0062] See also Figure 1 , Figure 1 1 is a flow chart of a method for protecting data of a drone based on verifiable privacy-preserving joint learning provided by an embodiment of the present invention. An embodiment of the present invention provides a method for protecting data of a drone based on verifiable privacy-preserving joint learning. The method for protecting data of a drone includes:
[0063] Step 1: Obtain m+1 positive integers, a first random integer sequence, a second random integer sequence, a first seed sequence, a second seed sequence, a third seed sequence, and a pseudo-random number sent by a private key generator (PKG), wherein the pseudo-random numbers of all participants are the same.
[0064] Specifically, there are N participants (participants are users, and users are drone users), and P i (i∈N) represents the i-th participant, where the set N={1,2,…,n}(n≥2). In the initialization phase, the PKG (private key generator) needs to generate and distribute m+1 positive integers, the first random integer sequence, the second random integer sequence, the first seed sequence, the second seed sequence, the third seed sequence and the PRG (pseudorandom number generator). These parameters and the PRG are kept secret from the aggregation server.
[0065] In an optional embodiment, step 1 may specifically include:
[0066] Step 1.1, obtain m+1 positive integers sent by the private key generator, and the m+1 positive integers are mutually prime.
[0067] Specifically, according to the convolutional neural network model architecture agreed upon by all participants (the convolutional neural network model is agreed upon by all participants who own drones, and this embodiment does not limit its specific form), PKG generates a learning rate η and initializes the parameter M of the neural network model. The security parameter m is set according to the needs of the participants. m is adjustable, just like training a neural network model. The larger the value of m, the higher the security, and the smaller the value of m, the lower the security. For example, m≥2. In addition, all participants will receive m+1 positive integers (i.e. g0, g1, g2, ..., g m ), these positive integers are mutually prime, that is, their greatest common factor is gcd(g i ,g j )=1,(i≠j), i=0,1,2,...,m,j=0,1,2,...,m,they act as the modulus value of modulo, so the m+1 positive integers are large enough to avoid overflow errors.
[0068] Step 1.2: Obtain a first random integer sequence and a second random integer sequence sent by a private key generator, wherein the first random integer sequence and the second random integer sequence have no intersection, and each of the first random integer sequence and the second random integer sequence includes m integers.
[0069] Specifically, PKG will distribute two random integer sequences to each participant, namely the first random integer sequence and the second random integer sequence. The first random integer sequence is denoted as {a i |i=1,2…,m},a i is the i-th element in the first random integer sequence, and the second random integer sequence is denoted by {b i |i=1,2…,m},b i is the i-th element in the second random integer sequence. The two random integer sequences have no intersection. They serve as the abscissa points of interpolation during Lagrange interpolation. The lengths of the two random integer sequences are related to the above-mentioned security parameter m.
[0070] Step 1.3: Obtain the first seed sequence, the second seed sequence, and the third seed sequence sent by the private key generator.
[0071] Specifically, PKG sends three seed sequences to each participant, and PRG(·) is a pseudo-random number. The three seed sequences are the first seed sequence, the second seed sequence, and the third seed sequence. The i-th participant P i The first subsequence received is is the jth element in the sequence of the ith participant, and the second seed sequence is is the element in the i-th sequence of the j-th participant, and the third seed sequence is {α i |i∈N},α i is a set sequence, specifically a set of real numbers from 0 to 1. PRG satisfies additive homomorphism, PRG(·+*)=(PRG(·)+PRG(*))(modg), PRG(·+*) is a random function that satisfies additive and multiplicative homomorphism, PRG(*) is a pseudo-random function that satisfies multiplicative homomorphism, PRG(·) is a universal pseudo-random function, and modg is the modulo operation of g.
[0072] Step 2: Based on the privacy training data set of the i-th participant, obtain the loss function of the neural network model to determine the private gradient of the i-th participant's finite field according to the private gradient of the i-th participant obtained by the loss function.
[0073] In an optional embodiment, step 2 may specifically include:
[0074] Step 2.1: Based on the neural network model, obtain the loss function according to the privacy training data set of the i-th participant.
[0075] Specifically, the training of the neural network model can be expressed as a function, where x is the input and M is the parameter of the neural network model. Assume that the i-th participant P iHold your own private training dataset D i ={ <x j ,y j >|j=1,2,…,T},D i is the privacy training dataset of the ith participant. The data in the privacy training dataset are the relevant element data of the drone owned by the participant, specifically the coordinate data collected by the drone. T is the size of the privacy training dataset of the ith participant. j is the jth data in the privacy training data set of the i-th participant, is the input of the neural network model, y j is the label of the jth data.
[0076] After the parameters of the neural network model are established, there will be a loss function. The loss function represents the degree of error of the neural network model. That is, the less accurate the neural network model is, the greater the value of the loss function is. The loss function can be defined as follows:
[0077]
[0078] Among them, L f (D i ,M) is the loss function, f(x j ,M) is about x j and M is a univariate nth-degree polynomial.
[0079] Step 2.2: Based on the back-propagation algorithm, the private gradient of the i-th participant is obtained according to the loss function.
[0080] Specifically, the goal of training a neural network model is to find the gradient to update the parameters M of the neural network model, thereby minimizing the value of the loss function. At this stage, participants execute the backpropagation algorithm to calculate private gradients. The private gradient of the i-th participant is expressed as:
[0081]
[0082] in, is the private gradient of the ith participant, is the loss function L f The derivative of *i D i A random subset of i A set of randomly selected elements.
[0083] Step 2.3: Convert the private gradient of the i-th participant from the real number field to the finite field to obtain the private gradient of the i-th participant in the finite field.
[0084] Specifically, after calculating the private gradient, in order to perform encryption operations, each participant needs to Convert from the field of real numbers to the field of finite numbers.
[0085] Step 2.31: Amplify the private gradient of the i-th participant to obtain an amplified private gradient.
[0086] Step 2.32: round the amplified private gradient to obtain a rounded private gradient.
[0087] Step 2.33: Convert the rounded private gradient to a finite field through the mapping of the finite field to obtain the private gradient of the finite field of the i-th participant. The private gradient of the finite field of the i-th participant is expressed as:
[0088]
[0089] Among them, w i is the private gradient of the finite field of the i-th participant, l w is a preset integer, l w is a larger integer, l w Control the accuracy of calculation, l w The larger the value, the more accurate the calculation, but there may be overflow and high occupancy problems, so a suitable value should be selected, such as 10 5 , (·) round is the rounding operation, and ψ(·) is a mapping from the real number field to the finite field.
[0090] Specifically, the rounding operation function is:
[0091]
[0092] in, is the largest integer less than or equal to x1.
[0093] The function ψ(·) is a function from the domain of integers To a finite field The mapping is defined as follows:
[0094]
[0095] Where q is a generator of a finite field. Note that the domain of x2 should be At the same time, in order to avoid an error in the calculation process, the prime number q must be large enough and satisfy q>>l w In finite fields, there are not only specified modular operations and decoding, but also encryption and decryption of gradients.
[0096] Step 3: Based on the pseudo-random numbers of the first seed sequence and the second seed sequence, the private gradient of the finite field of the i-th participant is blinded to obtain the blinded private gradient of the i-th participant.
[0097] In an optional embodiment, step 3 may specifically include:
[0098] Step 3.1: Obtain a first parameter according to the pseudo-random number of the first seed sequence.
[0099] Specifically, in the training phase of the neural network model, only the encryption algorithm is included. Encrypting the gradient before uploading is the key to ensure gradient aggregation and verifiability, so the encryption algorithm will be introduced next. Obviously, w i The value of is the same as the parameter M of the neural network model. In the tth round of loop iteration, the i-th participant P i Generate the first parameter using PRG(·) and the first seed sequence and the second seed sequence Right now Indicates that in round t A random number, is the tth element of the first seed sequence, is the tth element of the second seed sequence, n is the sequence dimension, which is a positive integer. The value of the first parameter is also related to the gradient w i Same. Define the gradient w i The value of |w i |, the training of PRG and neural network models can be performed in parallel.
[0100] Step 3.2: Based on the first parameter, the private gradient of the finite field of the i-th participant is blinded to obtain the private gradient of the i-th participant after blinding. The private gradient of the i-th participant after blinding is expressed as:
[0101]
[0102] in, is the private gradient of the i-th participant after blinding.
[0103] Step 4: Determine the interpolation points based on the m-1 preset parameters selected based on the private gradient of the i-th participant after blind processing, the first random integer sequence, and the pseudo-random numbers of the third seed sequence, and perform Lagrangian interpolation processing using the Lagrangian interpolation method to obtain a first interpolation function.
[0104] In an optional embodiment, step 4 may specifically include:
[0105] Step 4.1: Randomly generate m-1 preset parameters, the preset parameters satisfy Among them, v i,jis the jth preset parameter for the ith participant.
[0106] Specifically, the i-th participant P i Split The i-th participant P i Randomly generate m-1 preset parameters, and {v i,j |j=1,2,…,m-1},v i,j The jth preset parameter is randomly generated for the i-th participant, and the preset parameter satisfies
[0107] Step 4.2: Determine the interpolation point by using the first random integer sequence as the abscissa of the interpolation point and the second parameter generated according to the pseudo-random number of the third seed sequence and the preset parameter as the ordinate of the interpolation point.
[0108] Specifically, the second parameter is generated according to the pseudo-random number of the third seed sequence, that is, According to the tth round α i The random number gets the second parameter And take the i-th element in the first random integer sequence as the horizontal coordinate, and the second parameter and v i,j As the ordinate of the interpolation point, the interpolation point is specifically expressed as: (a1,v i,1 ),(a2,v i,2 ),…,(a m-1 ,v i,m-1 ),(a m ,A i ).
[0109] Step 4.3: Based on the interpolation points, Lagrangian interpolation processing is performed using the Lagrangian interpolation method to obtain a first interpolation function.
[0110] Specifically, the first interpolation function is expressed as:
[0111]
[0112] Among them, F i (x) is the first interpolation function, F i (x) specifications and w i Similarly, each element is a (m-1)-degree polynomial.
[0113] Step 5: Generate ciphertext based on the first interpolation function and the second random integer sequence.
[0114] In an optional embodiment, step 5 may specifically include:
[0115] Step 5.1: Based on the first interpolation function, obtain a difference result according to the second random integer sequence.
[0116] Specifically, the second random integer sequence (b j (j=1,2,…,m)) is substituted into the first interpolation function to obtain the difference result, which is recorded as (F i (b1),F i (b2),…,F i (b m )).
[0117] Step 5.2: Execute the CRT function on the difference result to obtain the ciphertext.
[0118] Specifically, in order to reduce space overhead and improve security, the Chinese Remainder Theorem (CRT) is used to encapsulate the results. That is, the i-th participant P i The CRT function is performed on the difference result, corresponding to the modulus value g j , (j=1,2,…,m), and get the packed data, i.e. ciphertext, which is expressed as:
[0119] [[w i ]]=CRT[F i (b1),F i (b2),…,F i (b m )]
[0120] Among them, [[w i ]] is the ciphertext, which is the i The final result of encryption. To avoid Overflow error in, these m positive integers should satisfy g i >n×g and g>n×l w , (i=1,2,…,m). In order to reduce the computational overhead, the parameters H of the Lagrangian interpolation i G i and It can be pre-calculated by the participants or sent directly by the PKG. Finally, the participants upload the ciphertext to the aggregation server.
[0121] Step 6: Obtain the aggregated ciphertext obtained by aggregating the ciphertexts of all participants.
[0122] Specifically, the aggregated ciphertext obtained by adding the ciphertexts of all participants is obtained.
[0123] In this embodiment, after receiving the ciphertexts uploaded by all participants, the aggregation server performs aggregation. The aggregation process is to add up all the received ciphertexts, that is, In fact, for two packed data If their corresponding packing modulus values are g1, g2, ..., g m , then the following equation holds:
[0124]
[0125] Among them, ≡ is identical. The above equation reveals that CRT satisfies the addition homomorphism, so the following equation is derived:
[0126]
[0127] Among them, [[w]] is the ciphertext after aggregation, Each element of F(x) is also a polynomial of degree m-1. It is worth noting that the aggregation server cannot obtain the expression of the function F(x) because the integer sequence {b i |i=1,2…,m} is kept secret from the aggregation server, which only knows the value of F(x) at some unknown points. After that, the aggregation server distributes the aggregated ciphertext [[w]] to each participant.
[0128] Step 7: Determine a second interpolation function according to the decrypted result obtained by decrypting the aggregated ciphertext and the second random integer sequence, and obtain a decoding result based on the second interpolation function.
[0129] Step 7.1: Decrypt the aggregated ciphertext through modular operation to obtain the decrypted result.
[0130] Specifically, each participant decrypts the received ciphertext [[w]] through modular operation, that is, F(b i )←[[w]]mod g i ,(i=1,2,…,m), decrypted to get [F i (b1),F i (b2),…,F i (b m )], recorded as the decrypted result.
[0131] Step 7.2: Use the data set consisting of the second random integer sequence and the decrypted result as interpolation points to perform Lagrange interpolation to obtain a second interpolation function.
[0132] Specifically, the decrypted result and the second random integer sequence {b i |i=1,2…,m} to form a data set, and then use this data set as the interpolation point to perform Lagrange interpolation using the Lagrange interpolation method to calculate the second interpolation function, which is expressed as:
[0133]
[0134] in,
[0135] Step 7.3: confirm whether the aggregated ciphertext is correct based on the mth element in the first random integer sequence and the interpolation result obtained by the second interpolation function. If it is wrong, end; if it is correct, obtain the aggregate value of the original gradient based on the first random integer sequence and the second interpolation function.
[0136] Specifically, each participant sets x = a m (a m is the mth element in the first random integer sequence) into the second interpolation function and calculates F(a m ), F(a m ) is recorded as the interpolation result. Then determine F(a m )=A t If the equation is not true, the result should be regarded as a forged result and the participants should end the agreement. m )=A t If true, it means that the aggregated ciphertext is correct, where Then the aggregate value of the original gradient is obtained according to the first random integer sequence and the second interpolation function. At the same time, it is noted that the size of F(x) is the same as the parameter M of the neural network model, and each element in it is a variable function, so the process is actually the verification of multiple equations.
[0137] In a specific embodiment, obtaining the aggregate value of the original gradient according to the first random integer sequence and the second interpolation function includes:
[0138] S1. Based on the second interpolation function, obtain a calculation result according to each element in the first random integer sequence.
[0139] Specifically, if the aggregated ciphertext is verified to be correct, each participant will replace the first random integer sequence {a i |i=1,2…,m-1} is substituted into the calculated second interpolation function to obtain F(a1), F(a2),…, F(a m-1 ) and record it as the calculation result.
[0140] S2. Add all the calculation results to get the aggregate value of the original gradient.
[0141] Specifically, the calculation formula of the aggregate value of the original gradient is expressed as:
[0142] Step 7.4: Decode using the aggregated value of the original gradient to obtain the decoded result.
[0143] Specifically, because the aggregate value of the original gradient is an integer, each participant needs to convert it to a floating point number to get the decoded result and then update the neural network model. This process can be implemented by the following function:
[0144]
[0145] Among them, ψ -1 (x) is from a finite field To the real number field 's mapping.
[0146] At the end of this phase, the participant locally updates the parameters M of the neural network model: After that, the next round of training will be carried out, repeating the model training stage, aggregation stage and the process of this stage until the termination conditions of the training rounds required by the participants are met, and the final trained model is obtained. The final trained model can be applied to the drone cluster system scenario, making its data available but invisible to protect the privacy data of the drone.
[0147] In order to illustrate the safety of the method provided by the present invention, a safety analysis will be performed next.
[0148] (1) Correctness analysis
[0149] Theorem 1: In the framework of the present invention, if each entity executes the protocol honestly, participants can obtain the correct aggregate gradients to update the model.
[0150] Proof: If each entity executes the protocol honestly, When this is true, participants can obtain the correct aggregate gradient. Next, we will prove that this formula is true, according to Algorithm 1:
[0151]
[0152] And because:
[0153]
[0154] From the above two equations, we can conclude that Therefore, if each entity executes the protocol honestly, the participants can obtain the correct aggregate gradients to update the model.
[0155] (2) Privacy Analysis
[0156] Theorem 2: In the framework of this method, each participant’s private gradient w i and the parameters M of the neural network model will not be leaked to the aggregation server.
[0157] Proof: Protecting the parameters M of the neural network model requires ensuring that the aggregate value w is not leaked to the aggregation server in each round. If the aggregation server attempts to calculate the private gradient w i and aggregate value w, then it needs to get F i (a j ) and F(a j ), where j = 1, 2, ..., m-1. However, the integer sequence {b i |i=1,2…,m} and m positive integers {g1,g2,…,g m} is also kept secret from the aggregation server, so the aggregation server cannot obtain the information from the [[w]] and [[w i ]]Recover the interpolation function F i (x) and F(x), and the integer sequence {a i |i=1,2…,m-1} is also kept secret from the aggregation server. In the example, the probability that the aggregation server obtains two integer sequences is in is the number of combinations. The probability of obtaining m positive integers is where len is g i The length of is, for example, 32 bits. Therefore, and The larger the value of i The more secure the aggregate gradient w is, the greater the value of g is. The value of g is always large, such as 2 to the power of 25. In summary, the private gradient w i and the parameters M of the neural network model will not be leaked to the aggregation server.
[0158] But the above theorem is based on the aggregation server's independent ownership of [[w]] and [[w i ]], but more often some participants collude with the aggregation server to obtain the private gradients of other participants. In this case, can this method still protect the user's private gradient? In order to solve this problem, this method proposes Theorem 3.
[0159] Theorem 3: In the framework of this method, if the aggregation server colludes with k (k ≤ n-2) participants, the private gradients of other participants will not be leaked.
[0160] Proof: Without loss of generality, assume that participants P1, P2, ..., and P K (k≤n-2) colludes with the aggregation server. Note that the seed sequence and It is kept secret from the colluders. From the data uploaded by other participants, they can only get:
[0161]
[0162] Private gradients are The colluder cannot obtain the private gradients of other participants due to blinding. Therefore, if the aggregation server colludes with k (k ≤ n-2) participants, the gradients of others will not be threatened with privacy.
[0163] (3) Verifiability analysis
[0164] Theorem 4: In the framework of this method, each participant can independently verify the correctness of the aggregated result, and the verification mechanism can detect falsified results with overwhelming probability.
[0165] Proof: In round t, if the participant receives the correct aggregate result [[w]], then obviously, F(x) satisfies:
[0166]
[0167] Each participant can independently check Is it true?
[0168] But if the aggregation server forges the aggregation result and it is accepted by the participants, without loss of generality, assume that the aggregation server tampers with the aggregation result to be:
[0169] [F(b1)+Δx1,F(b2)+Δx2,…,F(b m )+Δx m ]
[0170] Where Δx i is F(b i ) is modified. If the aggregation server attempts to successfully circumvent the verification mechanism so that the tampering is not detected by the participants, it needs to ensure that:
[0171]
[0172] Among them, F * (x) is the interpolation function calculated based on the forged result, l i (x) It has been given in and It can be deduced that is equivalent to:
[0173]
[0174] Therefore, the order of the aggregated servers should be If it is established, the aggregation result can be tampered without being discovered by the participants. However, it should be noted that l i(a m ) is the same as {b i |i=1,2…,m} and a m associated, they are all kept secret from the aggregation server. The probability that the aggregation server can obtain these parameters is (g>>m). Therefore, l i (a m ) is kept secret from the aggregation server. m} is also kept secret from the aggregation server, which means that the aggregation server cannot obtain [F(b1),F(b2),…,F(b m )]. Therefore, it is impossible for the aggregation server to forge a result to make Established.
[0175] The present invention uses blinding technology to protect privacy gradients, solving the problem of exposing sensitive information in plaintext transmission gradients. As long as no more than n-2 of the n participants collude with the aggregation server, the encrypted gradients of other participants can be guaranteed not to be decrypted. On the other hand, Lagrange interpolation is used to set the interpolation points, and the special point function values of the function restored by quadratic interpolation are used to verify the correctness of the aggregated gradients. Participants can perform verification operations on their own terminals to effectively verify whether their own gradients are aggregated, prevent forgery by the aggregation server, and the verification overhead of the scheme remains unchanged, that is, it is independent of the number of participants.
[0176] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0177] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification.
[0178] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the drawings and the disclosure. In the specification, the word "comprising" does not exclude other components or steps, and "one" or "a" does not exclude multiple situations. Certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0179] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A data protection method for drones based on verifiable privacy-preserving federated learning, characterized in that: The data protection method of the drone includes: Obtain m+1 positive integers, a first random integer sequence, a second random integer sequence, a first seed sequence, a second seed sequence, a third seed sequence, and a pseudo-random number sent by a private key generator, wherein the pseudo-random numbers of all participants are the same; Based on the privacy training data set of the i-th participant, a loss function of the neural network model is obtained to determine the private gradient of the i-th participant's finite domain according to the private gradient of the i-th participant obtained by the loss function; Based on the pseudo-random numbers of the first seed sequence and the second seed sequence, blinding the private gradient of the finite field of the i-th participant to obtain the blinded private gradient of the i-th participant; Determine interpolation points based on m-1 preset parameters selected from the private gradient of the ith participant after blinding processing, the first random integer sequence, and pseudo-random numbers of the third seed sequence, and perform Lagrangian interpolation processing using a Lagrangian interpolation method to obtain a first interpolation function; Based on the first interpolation function, generating a ciphertext according to the second random integer sequence; Obtaining an aggregated ciphertext obtained by aggregating the ciphertexts of all participants; A second interpolation function is determined according to a decrypted result obtained by decrypting the aggregated ciphertext and the second random integer sequence, and a decoding result is obtained based on the second interpolation function.
2. The data protection method of a drone according to claim 1, characterized in that: Obtaining m+1 positive integers, a first random integer sequence, a second random integer sequence, a first seed sequence, a second seed sequence, and a third seed sequence sent by a private key generator, including: Obtain m+1 positive integers sent by the private key generator, where each of the m+1 positive integers is relatively prime to each other; Obtaining a first random integer sequence and a second random integer sequence sent by the private key generator, wherein the first random integer sequence and the second random integer sequence have no intersection, and the first random integer sequence and the second random integer sequence each include m integers; Obtain a first seed sequence, a second seed sequence, a third seed sequence and a pseudo-random number sent by the private key generator.
3. The data protection method of a drone according to claim 1, characterized in that: Based on the privacy training data set of the i-th participant, a loss function of the neural network model is obtained, so as to determine the private gradient of the i-th participant's finite field according to the private gradient of the i-th participant obtained by the loss function, including: Based on the neural network model, the loss function is obtained according to the privacy training data set of the i-th participant, and the loss function is expressed as: Among them, L f (D i ,M) is the loss function, M is the parameter of the neural network model, D i is the private training dataset of the ith participant, D i ={ <x j ,y j >|j=1,2,…,T}, T is the size of the private training dataset of the i-th participant, x j is the jth data in the privacy training dataset of the i-th participant, y j is the label of the jth data, f(x j ,M) is about x j and M is a univariate n-degree polynomial; Based on the back propagation algorithm, the private gradient of the ith participant is obtained according to the loss function, and the private gradient of the ith participant is expressed as: in, is the private gradient of the ith participant, is the loss function L f The derivative of *i D i A random subset of ; The private gradient of the i-th participant is converted from the real number field to the finite field to obtain the private gradient of the i-th participant in the finite field.
4. The data protection method of a drone according to claim 3, characterized in that: The private gradient of the i-th participant is converted from the real number field to the finite field to obtain the private gradient of the i-th participant in the finite field, including: Amplifying the private gradient of the i-th participant to obtain an amplified private gradient; Rounding the amplified private gradient to obtain a rounded private gradient; The rounded private gradient is converted into a finite field through finite field mapping to obtain the private gradient of the finite field of the i-th participant. The private gradient of the finite field of the i-th participant is expressed as: Among them, w i is the private gradient of the finite field of the i-th participant, l w is a preset integer, (·) round is the rounding operation, and ψ(·) is a mapping from the real number field to the finite field.
5. The data protection method of a drone according to claim 1, characterized in that: Based on the pseudo-random numbers of the first seed sequence and the second seed sequence, blinding the private gradient of the finite field of the i-th participant to obtain the blinded private gradient of the i-th participant includes: Obtaining a first parameter according to the pseudo-random numbers of the first seed sequence and the second seed sequence; The private gradient of the finite field of the i-th participant is blinded based on the first parameter to obtain the private gradient of the i-th participant after blinding. The private gradient of the i-th participant after blinding is expressed as: in, is the private gradient of the ith participant after blinding, w i is the private gradient of the finite field of the i-th participant, is the t-th seed subsequence.
6. The drone data protection method according to claim 1, characterized in that: The interpolation points are determined based on the m-1 preset parameters selected from the private gradient of the i-th participant after blind processing, the first random integer sequence, and the pseudo-random numbers of the third seed sequence, so as to perform Lagrangian interpolation processing using the Lagrangian interpolation method to obtain a first interpolation function, including: Randomly generate m-1 preset parameters, the preset parameters satisfy in, is the private gradient of the ith participant after blinding, v i,j The jth preset parameter randomly generated for the i-th participant; Determine the interpolation point by using the first random integer sequence as the abscissa of the interpolation point and the second parameter generated according to the pseudo-random number of the third seed sequence and the preset parameter as the ordinate of the interpolation point; Based on the interpolation points, Lagrangian interpolation processing is performed using a Lagrangian interpolation method to obtain a first interpolation function.
7. The data protection method of a drone according to claim 1, characterized in that: Generating ciphertext according to the second random integer sequence based on the first interpolation function includes: Based on the first interpolation function, obtaining a difference result according to the second random integer sequence; The CRT function is executed on the difference result to obtain the ciphertext. The aggregated ciphertext is obtained by adding the ciphertexts of all participants.
8. The drone data protection method according to claim 1, characterized in that: Determining a second interpolation function according to a decrypted result obtained by decrypting the aggregated ciphertext and the second random integer sequence, and obtaining a decoding result based on the second interpolation function, including: Decrypting the aggregated ciphertext by modular operation to obtain a decrypted result; Performing Lagrange interpolation on a data set consisting of the second random integer sequence and the decrypted result as interpolation points to obtain a second interpolation function; Confirm whether the aggregated ciphertext is correct according to the interpolation result obtained by the mth element in the first random integer sequence and the second interpolation function, and if it is wrong, end; if it is correct, obtain the aggregate value of the original gradient according to the first random integer sequence and the second interpolation function; Decoding is performed using the aggregated value of the original gradient to obtain a decoding result.
9. The drone data protection method according to claim 8, characterized in that: Obtaining an aggregate value of the original gradient according to the first random integer sequence and the second interpolation function, including: Based on the second interpolation function, obtaining a calculation result according to each element in the first random integer sequence; All the calculation results are added together to obtain the aggregate value of the original gradient.
10. The drone data protection method according to claim 8, characterized in that: Decoding is performed using the aggregated value of the original gradient to obtain a decoding result, including: The aggregate value of the original gradient is converted into a floating point number to obtain a decoding result.