Privacy-Preserving Nonlinear Federated Support Vector Machine Training Method and System Based on Homomorphic Encryption
By adopting homomorphic encryption technology in federated learning SVM, the original data is mapped to high-dimensional space and joint training is carried out using the privacy-protected federated SVM algorithm, the problem of privacy leakage and training is solved, and efficient privacy protection and model accuracy are achieved.
Patent Information
- Application Number
- CN202210758385.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The existing federated learning SVM research has the risk of privacy leakage and low training efficiency, especially in the horizontal federated learning scenario, it is difficult to directly solve nonlinear federated SVM using nuclear skills.
The nonlinear federated support vector machine training method based on homomorphic encryption is adopted to obtain group keys and random seeds through the key negotiation protocol, map the original data to the same high-dimensional space, and use the privacy-protected federated SVM algorithm for joint training of the model.
It realizes the improvement of training speed while ensuring privacy, avoids homomorphic multiplication with high time overhead, ensures losslessness of model accuracy, and has deep application value.
Smart Images

Figure CN115392487B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning for privacy protection, and particularly relates to a privacy protection nonlinear federated support vector machine training method and system based on homomorphic encryption. Background Art
[0002] With the rapid development of big data and artificial intelligence, data security issues such as privacy protection and data security have attracted increasing attention. How to perform joint mining of multi-source data while effectively protecting the privacy data of users has become an urgent problem to be solved. As a privacy protection technology based on the theory of public key cryptography and ensuring the security of ciphertext through mathematical problems, homomorphic encryption provides an effective solution for solving the privacy problem of joint mining of multi-source data.
[0003] Currently, the research on privacy protection SVM mainly focuses on four directions, namely: secure multi-party computation, data perturbation, secure outsourcing computation, and federated learning. The research on SVM based on secure multi-party computation mostly realizes joint modeling of multiple participants through secure multi-party computation protocols, which can ensure the privacy of the training data of each participant, but has the defects of large time and communication overhead; SVM based on data perturbation refers to ensuring data privacy by injecting noise into the Gram matrix, but it will reduce the accuracy of the model; SVM based on secure outsourcing computation refers to leveraging the powerful computing power of the cloud server, encrypting the original data and uploading it to the cloud, and the cloud performs model training and prediction in the ciphertext domain, but its overall computational complexity is high, and local data needs to be encrypted and uploaded to the cloud server. Federated learning, on the other hand, can establish a global model only by exchanging model parameters between participants while ensuring that the data does not leave the local. Compared with other methods, federated learning has more excellent training efficiency and a wider application background. Therefore, in recent years, the research on privacy protection federated learning based on homomorphic encryption has gradually gained popularity.
[0004] However, there are also many challenges in the current existing research on federated learning SVM. First, some research directly applies the federated SVM algorithm to corresponding scenarios without using any privacy protection means, which poses a potential risk of privacy leakage; second, the mainstream horizontal federated learning solves problems through gradient descent, and there is no relevant research on solving problems through the SMO algorithm. The reason is that the globally optimal coefficients calculated for the dual problem are different from the locally optimal coefficients calculated for local data. Since each party has a data subset, it is difficult to generate globally optimal coefficients by solving the dual problem on the data subset. This means that in the scenario of horizontal federated learning, it will be difficult to directly solve the nonlinear federated SVM using the kernel trick. Summary of the Invention
[0005] To this end, the present invention provides a privacy-preserving non-linear federated support vector machine training method and system based on homomorphic encryption, which realizes privacy-preserving federated SVM joint training, can balance privacy protection and training speed of each participating party, and is convenient for practical scenario applications.
[0006] According to the design scheme provided by the present invention, a privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption is provided, which includes the following contents:
[0007] Each participating party uses a key agreement protocol to obtain a group key and a random seed for generating an original data mapping function;
[0008] Each participating party uses the random seed to generate a mapping function, and uses the mapping function to map the local data set of each participating party as the original data to the same high-dimensional space to obtain the high-dimensional data corresponding to each participating party;
[0009] Each participating party uses the high-dimensional data as training samples, uses a privacy-preserving federated SVM algorithm to train local model parameters, and the server uses homomorphic encryption operations to aggregate the local model parameters, and obtains the final model through joint training.
[0010] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, each participating party adopts the Burmester-Desmedt protocol as the key agreement protocol to obtain a public-private key pair; and broadcasts its own public key to other participating parties; the group key is obtained through the intermediate parameters of each participating party participating in the key agreement.
[0011] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, the group key is used as the input of a pseudo-random number generator, and the pseudo-random number generator is used to obtain a random seed for generating an original data mapping function.
[0012] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, in the generation of the public-private key pair using the Burmester-Desmedt protocol, the selected group is the 2048-bit multiplicative cyclic group in RFC 3526.
[0013] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, the pseudo-random number generator adopts the ChaCha pseudo-random number generator.
[0014] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, for the obtained random seed, the original data of each participating party is mapped to the same high-dimensional space in combination with the random Fourier feature algorithm.
[0015] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, in the random Fourier feature algorithm, a random number generator is set according to a random seed; and a mapping function is constructed by the random number generator; the original data of the local data sets of each party is mapped by using the mapping function to obtain the corresponding high-dimensional data after mapping.
[0016] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, in the joint training of the model by using the privacy-preserving federated SVM algorithm, the server generates initial model parameters and broadcasts the initial model parameters to each party; each party performs iterative training on the local model. In the iterative training, first, each party uses the local training samples to train the model to obtain local model parameters, homomorphically encrypts the local model parameters and uploads them to the server. The server uses homomorphic encryption operations to aggregate the local model parameters of each party to obtain the ciphertext of the global model parameters, and sends the ciphertext of the global model parameters to each party. Each party decrypts according to the received ciphertext of the global model parameters to obtain the local model parameters, and enters the next round of model training until the preset maximum number of iterative rounds is satisfied.
[0017] As the privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption of the present invention, further, in the joint training, the contribution degree of each party is set according to the sample size of the local data set of each party. In the iterative training, each party uses the local training samples to train the model parameters of this round, and uses the received global model parameters, the local model parameters of this round of training and the contribution degree of the party to update the global model parameters received in the current round, and uses the updated global model parameters as the local model parameters of the party, homomorphically encrypts the local model parameters and sends them to the server.
[0018] Further, the present invention also provides a privacy-preserving non-linear federated support vector machine training system based on homomorphic encryption, including: a key negotiation module, a sample construction module and a joint training module, wherein,
[0019] The key negotiation module is used for each party to obtain a group key and a random seed for generating an original data mapping function by using a key negotiation protocol;
[0020] The sample construction module is used for each party to generate a mapping function by using the random seed, and use the mapping function to map the local data sets of each party as the original data to the same high-dimensional space to obtain the corresponding high-dimensional data of each party;
[0021] The joint training module is used for each participant to use high-dimensional data as training samples, train local model parameters using a privacy-preserving federated SVM algorithm, and aggregate the local model parameters by the server using homomorphic encryption operations, and obtain the final model through joint training.
[0022] Advantages of the present invention:
[0023] In the present invention, multiple participants jointly negotiate to generate a group key and a random seed, use the random Fourier feature algorithm to map the original training data to the same high-dimensional space, and use a privacy-preserving federated SVM training algorithm for joint training of the model, so as to take into account both privacy and training speed; furthermore, in the training stage, CKKS homomorphic encryption is used to ensure the privacy of the model parameters of each participant and the contribution of the participant (the size of the local dataset) during the training process, and the use of homomorphic multiplication with high time overhead can be avoided to alleviate the high time overhead brought by the introduction of the cryptosystem. Through experimental verification, compared with other federated SVM training methods, the solution of this case can obtain lossless model accuracy while ensuring privacy, and has deep application value. Description of the drawings
[0024] Figure 1 It is a schematic diagram of the privacy-preserving non-linear federated support vector machine training process based on homomorphic encryption in the embodiment;
[0025] Figure 2 It is a schematic diagram of CKKS homomorphic encryption used in the embodiment;
[0026] Figure 3 It is a schematic diagram of a generated dataset Circle adopted in the embodiment;
[0027] Figure 4 It is a schematic diagram of a generated dataset Moon adopted in the embodiment;
[0028] Figure 5 It is a performance schematic diagram on 4 datasets adopted in the embodiment. Detailed implementation manners
[0029] To make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and technical solutions.
[0030] In the embodiment of the present invention, referring to Figure 1 as shown, a privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption is provided, including the following contents:
[0031] S101. Each participant uses a key negotiation protocol to obtain a group key and a random seed for generating an original data mapping function;
[0032] S102. Each participant uses the random seed to generate a mapping function, and uses the mapping function to map the local data sets of the participants as the original data into the same high-dimensional space to obtain the high-dimensional data corresponding to each participant.
[0033] S103. Each participant uses the high-dimensional data as training samples, uses the privacy-preserving federated SVM algorithm to train the local model parameters, and the server aggregates the local model parameters through homomorphic encryption operations, and obtains the final model through joint training.
[0034] In the solution of this case, multiple participants jointly negotiate to generate a group key and a random seed, use the random Fourier feature algorithm to map the original training data into the same high-dimensional space, and use the privacy-preserving federated SVM training algorithm for joint training of the model, so as to take into account both privacy and training speed, and realize the joint training of privacy-preserving federated SVM.
[0035] Further, each participant adopts the Burmester-Desmedt protocol as the key negotiation protocol to securely negotiate the random seed, and each participant obtains the same random seed. The participant broadcasts its own public key to other participants; the group key is obtained through the intermediate parameters of each participant participating in the key negotiation.
[0036] The algorithm process of securely negotiating the random seed using the Burmester-Desmedt protocol can be designed as the following content:
[0037] Step 1.1: Generate the public and private keys of the Burmester-Desmedt protocol locally at each participant
[0038] Step 1.2: Broadcast the public key to other participants
[0039] Step 1.3: Each participant calculates X i , where X i = Z i+1 / Z i ,
[0040] Step 1.4: Broadcast the intermediate result X i to other participants
[0041] Step 1.5: Calculate the group key K i , where
[0042] Step 1.6: Input the negotiated group key into the PRNG to obtain the random seed seed
[0043] Among them, the group selected by the Burmester-Desmedt protocol can be the 2048-bit multiplicative cyclic group in RFC 3526; in step 1.5, the group key length can be 2048; in step 1.6, the selected PRNG can be the cryptographically secure ChaCha pseudorandom number generator.
[0044] Furthermore, in the embodiments of this case, for the obtained random seed, the original data of each participant is mapped to the same high-dimensional space by combining the random Fourier feature algorithm. In the random Fourier feature algorithm, the random number generator is set according to the random seed; and the mapping function is constructed through the random number generator; the original data of the local data set of each participant is mapped by using the mapping function to obtain the corresponding high-dimensional data after mapping.
[0045] Using the negotiated random seed, data mapping is performed by combining the random Fourier feature algorithm. The specific algorithm can be designed to include the following steps:
[0046] Step 2.1: Use the seed negotiated in step 1 to set and obtain the random number generator random, and set the Gamma value γ of the Gaussian kernel
[0047] Step 2.2: Generate a standard normal distribution that can generate a fixed sequence of random numbers according to the random seed through the random number generator random and a uniform distribution
[0048] Step 2.3: Sample ω from the distributions p(ω), Uniform(0, 2π) i ,b i ,
[0049] Step 2.5: Obtain the mapping function where
[0050] Step 2.6: Map each piece of data in the data set D i to obtain the mapped data set X i
[0051] In the embodiments of this case, further, in the joint training of the model using the privacy-preserving federated SVM algorithm, the server generates initial model parameters and broadcasts the initial model parameters to each participating party; each participating party performs iterative training on the local model. In the iterative training, first, each participating party uses the local training samples to train the model to obtain local model parameters, homomorphically encrypts the local model parameters, and uploads them to the server. The server uses homomorphic encryption operations to aggregate the local model parameters of each participating party to obtain the ciphertext of the global model parameters, and sends the ciphertext of the global model parameters to each participating party. Each participating party decrypts the received ciphertext of the global model parameters to obtain the local model parameters and enters the next round of model training until the preset maximum number of iterative rounds is satisfied. In the joint training, the sample size of each participating party's local dataset is the contribution degree of the participating party in this round of training, denoted as n. After the client training is completed, the local model parameters obtained in this round of training need to be multiplied by n, and at the same time, n is placed in the last dimension of the current ciphertext vector, as shown in step 3.1.8, so as to ensure the privacy of the model parameters of each participating party and the contribution of the participating party (the size of the local dataset) during the training process; in the iterative training, each participating party uses the local training samples to train the model parameters of this round, and uses the received global model parameters, the local model parameters of this round of training, and the contribution degree of the participating party to update the global model parameters received in the current round, and uses the updated global model parameters as the local model parameters of the participating party, homomorphically encrypts the local model parameters, and sends them to the server.
[0052] After the original data is homomorphically encrypted by Homomorphic Encryption (HE), specific operations are performed on the ciphertext to obtain the ciphertext calculation result; after homomorphic decryption, the obtained plaintext is equivalent to the data result obtained by directly performing the same calculation on the original plaintext data. See Figure 2 As shown, in the embodiments of this case, in the joint training, CKKS homomorphic encryption can be used to ensure the privacy of the model parameters of each participating party and the contribution degree of the participating party during the training process; and further, by optimizing the calculation process, the use of homomorphic multiplication with high time overhead can be avoided, and the high time overhead brought by introducing the cryptosystem can be alleviated.
[0053] Using the high-dimensional data obtained after mapping as the training data, in the solution of this case, the algorithm for training the model based on homomorphic encryption and using the privacy-preserving federated SVM algorithm can include two parts: the client and the server. The specific training processes of the two can be designed as the following steps:
[0054] Client:
[0055] Step 3.1.1: Initialize the system parameters: learning rate, ciphertext vector length, number of communication rounds
[0056] Step 3.1.2: If it is the first interaction between the client and the server, receive the initial model parameter ω from the server 0
[0057] Step 3.1.3: Otherwise, receive the encrypted model parameter of the t-th round of iteration from the server
[0058] Step 3.1.4: Decrypt the received encrypted model parameter
[0059] Step 3.1.5: Calculate the global model parameter
[0060] Step 3.1.6: Use the local dataset D i Train the parameter g updated in this round b
[0061] Step 3.1.7: Update the model parameter ω t+1 = ω t - ηg b
[0062] Step 3.1.8: Hide the contribution ω t+1 = nω t+1 ,
[0063] Step 3.1.9: Encrypt the vector ω t+1 to get c i = Enc pk (ω t+1 ) and send it to the server
[0064] Step 3.1.10: Until the communication round number is reached
[0065] Server:
[0066] Step 3.2.1: Initialize the system parameters: the length of the encrypted vector, the communication round number
[0067] Step 3.2.2: Receive the encrypted model parameters c from all participants i
[0068] Step 3.2.3: Aggregate the global model parameters using homomorphic addition C = Add(C, c i )
[0069] Step 3.2.4: Send the aggregated ciphertext C to each participant
[0070] Furthermore, based on the above method, an embodiment of the present invention further provides a privacy-preserving non-linear federated support vector machine training system based on homomorphic encryption, including: a key negotiation module, a sample construction module, and a joint training module, where
[0071] The key negotiation module is used for each participant to obtain a group key and a random seed for generating an original data mapping function by using a key negotiation protocol;
[0072] The sample construction module is used for each participant to generate a mapping function by using the random seed, and use the mapping function to map the local data set of each participant as the original data to the same high-dimensional space to obtain the high-dimensional data corresponding to each participant;
[0073] The joint training module is used for each participant to use the high-dimensional data as training samples, train local model parameters by using a privacy-preserving federated SVM algorithm, and aggregate the local model parameters by the server using homomorphic encryption operations, and obtain the final model through joint training.
[0074] To verify the effectiveness of the solution of this case, the following further explains with experimental data:
[0075] Two real datasets Ring and BCD, and two generated datasets Moon and Circle are used for simulation experiments. The generated datasets Moon and Circle used in the simulation are as Figure 3 , shown in Figure 4. Figure 5 reflects the change of the model accuracy of the four datasets with the increase of the iteration number in the implementation algorithm of this case and the FedAvg algorithm without privacy protection. It can be seen from Figure 5 the results in that the PPNLFedSVM algorithm disclosed in this case can protect the privacy of the participant data without loss of model accuracy.
[0076] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the present invention.
[0077] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, which are used to illustrate the technical solutions of the present invention rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily conceive of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption, characterized in that, It includes the following content: Each participating party uses a key negotiation protocol to obtain a group key and a random seed for generating an original data mapping function. Among them, each participating party adopts the Burmester-Desmedt protocol as the key negotiation protocol to obtain a public-private key pair; and broadcasts its own public key to other participating parties; the group key is obtained through the intermediate parameters of each participating party's participation in key negotiation, and the group key K i is expressed as X i is the intermediate parameter, and X i =Z i+1 / Z i , Each participating party uses a random seed to generate a mapping function, and uses the mapping function to map the local data sets of each participating party as the original data into the same high-dimensional space to obtain the corresponding high-dimensional data of each participating party; Each participating party uses the high-dimensional data as training samples, uses the privacy-preserving federated SVM algorithm to train the local model parameters, and the server aggregates the local model parameters through homomorphic encryption operations, and obtains the final model through joint training.
2. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 1, wherein The group key is used as the input of the pseudo-random number generator, and the pseudo-random number generator is used to obtain the random seed for generating the original data mapping function.
3. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 1, wherein, In the generation of the public-private key pair using the Burmester-Desmedt protocol, the selected group is the 2048-bit multiplicative cyclic group in RFC 3526.
4. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 2, wherein The pseudo-random number generator adopts the ChaCha pseudo-random number generator.
5. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 1, characterized in that For the obtained random seed, the original data of each participating party is mapped into the same high-dimensional space in combination with the random Fourier feature algorithm.
6. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 5, characterized in that, In the random Fourier feature algorithm, the random number generator is set according to the random seed; and the mapping function is constructed through the random number generator; the original data of the local data set of each participating party is mapped by the mapping function to obtain the corresponding mapped high-dimensional data.
7. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 1, wherein In the joint training of the model using the privacy-preserving federated SVM algorithm, the server generates the initial model parameters and broadcasts the initial model parameters to each participating party; each participating party performs iterative training on the local model. In the iterative training, first, each participating party uses the local training samples to train the model to obtain the local model parameters, homomorphically encrypts the local model parameters and uploads them to the server. The server uses homomorphic encryption operations to aggregate the local model parameters of each participating party to obtain the ciphertext of the global model parameters, and sends the ciphertext of the global model parameters to each participating party. Each participating party decrypts according to the received ciphertext of the global model parameters to obtain the local model parameters, and enters the next round of model training until the preset maximum number of iterative rounds is satisfied.
8. The privacy-preserving non-linear federated support vector machine training method based on homomorphic encryption according to claim 7, characterized in that, In joint training, the contribution degree of each participating party is set according to the sample size of the local data set of each participating party. In iterative training, each participating party uses the local training samples to train the model parameters of this round, and uses the received global model parameters, the local model parameters of this round of training, and the contribution degree of the participating party to update the global model parameters received in the current round, and uses the updated global model parameters as the local model parameters of the participating party, homomorphically encrypts the local model parameters and sends them to the server.
9. A privacy-preserving non-linear federated support vector machine training system based on homomorphic encryption, characterized in that, It includes: a key negotiation module, a sample construction module, and a joint training module, where The key negotiation module is used for each participating party to obtain the group key and the random seed for generating the original data mapping function by using the key negotiation protocol. Among them, each participating party adopts the Burmester-Desmedt protocol as the key negotiation protocol to obtain the public-private key pair, broadcasts its own public key to other participating parties, and obtains the group key through the intermediate parameters of each participating party's participation in the key negotiation. The group key K i is expressed as X i is the intermediate result, and X i =Z i+1 / Z i , The sample construction module is used for each participating party to use a random seed to generate a mapping function, and use the mapping function to map the local data sets of each participating party as the original data into the same high-dimensional space to obtain the corresponding high-dimensional data of each participating party; The joint training module is used for each participating party to use the high-dimensional data as training samples, use the privacy-preserving federated SVM algorithm to train the local model parameters, and the server aggregates the local model parameters through homomorphic encryption operations, and obtains the final model through joint training.
Citation Information
Patent Citations
Linear SVM model training algorithm for privacy protection based on vector homomorphic encryption
CN108521326A
Federal learning privacy protection method based on homomorphic encryption
CN113434873A