Systematic privacy protection federated learning method and system based on multi-receiver encryption and differential privacy
By combining multi-receiver encryption and differential privacy technologies in federated learning, the shortcomings of traditional federated learning in data privacy and model security are solved, and efficient and secure data privacy protection is achieved, suitable for large-scale distributed environments.
Patent Information
- Application Number
- CN202510215761.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional federated learning methods pose potential risks in data privacy and model security, especially when data transmission is susceptible to eavesdropping and tampering attacks, while at the same time, high computing and communication overhead, and limited flexibility and scalability.
A systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy is adopted. By adding differential privacy noise in the local model update stage and using multi-receiver encryption technology in the global model distribution stage, we ensure that only authorized clients can decrypt model parameters.
It significantly improves the security of data privacy, reduces the risk of attacks during data transmission, reduces communication overhead and computing resource consumption, supports dynamic user access and efficient identity management, and is suitable for complex scenarios of large-scale and multi-users.
Smart Images

Figure CN120146223A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to, but is not limited to, the technical field of federated learning, and particularly relates to a systematic privacy protection federated learning method and system based on multi-receiver encryption and differential privacy. Background Art
[0002] With the development of artificial intelligence and big data technologies, federated learning, as an emerging distributed machine learning framework, can achieve collaborative modeling without directly accessing user data, effectively protecting data privacy. However, in practical applications, federated learning still faces many challenges, among which data privacy and model security issues are particularly prominent. Traditional federated learning relies on a centralized aggregation server, and this architecture has potential risks such as sensitive data leakage and malicious attacks. Therefore, how to effectively protect user privacy during the federated learning process while improving the system robustness has become a key issue of concern in the academic and industrial communities.
[0003] In recent years, privacy protection methods based on encryption technologies have gradually attracted wide attention. Currently, commonly used encryption methods include homomorphic encryption (HE), partial encryption, and secure multi-party computation (SMPC). These methods can achieve data privacy protection to a certain extent. Homomorphic encryption allows computations to be directly performed on encrypted data, and the computation result is consistent with the result of the plaintext data operation after decryption. However, homomorphic encryption has high computational overhead and low efficiency, and is particularly unsuitable for large-scale real-time tasks in a distributed environment. Partial encryption balances privacy protection and computational efficiency by encrypting some fields of sensitive data, but its protection strength is insufficient and it is vulnerable to side-channel attacks or differential attacks. Secure multi-party computation achieves privacy protection by splitting and distributing data processing, but requires a high degree of synchronization among participants, has high communication overhead, and is difficult to adapt to dynamic environments.
[0004] Traditional encryption methods show relatively insufficient in the following aspects:
[0005] Higher computation and communication overhead. Commonly used homomorphic encryption and SMPC methods have high computational complexity. Especially when dealing with large-scale data or high-frequency model updates, it will significantly increase the computational and communication costs.
[0006] Limited flexibility and scalability. Key management of homomorphic encryption and partial encryption is complex in a dynamic environment and does not support flexible user access. SMPC has high requirements for participant synchronization and is difficult to adapt to complex distributed scenarios of multiple users.
[0007] Insufficient privacy protection and performance balance. The high-intensity privacy protection of homomorphic encryption comes at the cost of performance, while partial encryption improves computational efficiency but has insufficient protection strength.
[0008] It is difficult to apply in actual projects. Common encryption methods require complex system integration and professional knowledge support in actual applications, increasing the deployment difficulty and cost. Summary of the Invention
[0009] In view of the problems existing in the prior art, the present invention provides a systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy.
[0010] The present invention is implemented as follows. A systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy, the method comprising:
[0011] S1: Initialize and distribute global model parameters
[0012] The server generates the initial global model parameters W 0 , and distributes them to all legitimate clients through multi-receiver encryption technology;
[0013] Assign a unique identity ID to each client i , and generate the corresponding encryption key SK using the master key MSK and the public parameter PK i ;
[0014] S2: Local model training on the client side
[0015] The client decrypts the global model and trains the local model based on the local private dataset D i to generate updated parameters W t+1 ; The training process is based on Stochastic Gradient Descent (SGD) or its variants, and the formula is as follows:
[0016]
[0017] where W t is the current model parameter, η is the learning rate, L(·) is the loss function, is the gradient of the loss function;
[0018] S3: Add privacy perturbation to the client model parameters
[0019] The client adds differential privacy noise ξ to the locally updated parameters to prevent potential reverse attacks. The noise satisfies the Laplace distribution, and the formula is as follows:
[0020]
[0021] where Δ is the sensitivity and ∈ is the privacy budget. By dynamically adjusting ∈ and Δ, the privacy protection and model performance can be flexibly balanced in different scenarios;
[0022] S4: The client uploads the perturbed parameters
[0023] The client uploads the perturbed parameter ΔW' i to the server;
[0024] S5: Global aggregation module
[0025] The server receives the perturbed parameters from all clients and calculates the global model update parameter:
[0026]
[0027] where S is the set of selected clients;
[0028] S6: The server uses multi - recipient encryption to distribute the global model
[0029] The server uses identity - based multi - recipient encryption technology (MR - IBE) to encrypt the updated global model parameter W t+1 to generate the ciphertext C G , ensuring that only authorized recipients can decrypt it; the encryption process is as follows:
[0030] C G = Encrypt(PK, {ID i}, W t+1 )
[0031] Distribute the ciphertext C G to the authorized clients, and the clients use the private key to decrypt and update the local model. This ciphertext supports simultaneous distribution to multiple recipients, avoiding the redundant operation of generating separate ciphertexts for each recipient in traditional encryption methods, significantly reducing the communication overhead; at the same time, ensuring the secure distribution of the global model; the clients decrypt and update the local model parameters through the private key.
[0032] Another object of the present invention is to provide a systematic privacy - protected federated learning device based on multi - recipient encryption and differential privacy for the above - mentioned systematic privacy - protected federated learning method based on multi - recipient encryption and differential privacy. This device specifically includes:
[0033] A server, responsible for the initialization, parameter aggregation, and distribution of the global model;
[0034] Client devices, responsible for local model training, parameter decryption, and noise addition.
[0035] Furthermore, the server includes:
[0036] A calculation unit: used to perform encryption, screening, and aggregation operations;
[0037] A storage unit: stores the global model parameters and client identity information;
[0038] Communication module: Performs encrypted data transmission with the client.
[0039] Furthermore, the client device includes:
[0040] Computing unit: Used to perform local model training and decryption operations.
[0041] Storage unit: Stores the local dataset and model parameters.
[0042] Communication module: Performs data transmission with the server.
[0043] Another object of the present invention is to provide a computer device, which includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy.
[0044] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy.
[0045] Another object of the present invention is to provide an information data processing terminal, which is used to implement the systematic privacy protection federated learning system based on multi-receiver encryption and differential privacy.
[0046] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are:
[0047] First, by combining multi-receiver encryption (MRE) and differential privacy (DP) technologies, the present invention proposes a systematic privacy-preserving federated learning method, significantly enhancing the security, efficiency, and applicability of existing technologies. First of all, the present invention adopts a dual privacy protection mechanism. During the local update phase, random noise is added through differential privacy technology to prevent data leakage; during the global distribution phase, multi-receiver encryption technology is used to ensure that model parameters can only be decrypted by authorized clients, effectively resisting attacks during the transmission process. Specifically, during the local update phase, the local model update of the client adds noise through the differential privacy mechanism, so that even if an attacker obtains the model update, the original data cannot be recovered from it. This technology effectively avoids the risk of data leakage and solves the limitation of only relying on encryption to protect privacy in traditional technologies. At the same time, during the global distribution phase, the present invention encrypts the model parameters through multi-receiver encryption technology, and only authorized clients can decrypt the corresponding model updates. This encryption scheme greatly improves the security of data transmission, solves the problems of eavesdropping and tampering attacks encountered during the transmission process, and ensures the comprehensive protection of data privacy during the federated learning process. There are two major defects in existing technologies in terms of privacy protection: on the one hand, the model parameters in the data transmission phase will be maliciously intercepted or tampered with, resulting in privacy leakage or model damage; on the other hand, although encryption technology plays a certain role in protecting the data transmission process, traditional encryption mechanisms often have problems such as excessive computational and communication overheads and the risk of information leakage during the transmission of ciphertext. This dual protection significantly enhances the anti-attack ability of the system and solves the above two privacy defects, and is particularly suitable for scenarios with high privacy requirements such as medical and financial fields.
[0048] Secondly, the present invention optimizes the ciphertext generation and distribution process through multi-receiver encryption technology, solving the problem of excessive communication overhead and computational resource consumption in traditional encryption methods. Multi-receiver encryption allows the same ciphertext to be distributed to multiple receivers simultaneously, without the need to generate a separate ciphertext for each receiver, greatly reducing the communication bandwidth requirements and computational costs, and is particularly suitable for large-scale distributed environments. In addition, the present invention supports dynamic adjustment of the differential privacy budget, flexibly balancing the privacy protection intensity and model performance according to specific scenarios, further enhancing the applicability and practicality of the method. At the same time, the modular design and dynamic scalability of the present invention enable the system to be quickly deployed in distributed environments of different scales, support dynamic user access and efficient identity management, and adapt to complex scenarios of large scale and multiple users.
[0049] Finally, the present invention is implemented based on a standard cryptographic library, with low hardware requirements and low implementation costs, and has broad application prospects. Its characteristics of high efficiency, security, and easy extensibility make the present invention have significant innovation and practicality in the field of privacy-preserving federated learning, and can provide reliable solutions for multiple fields such as medical, financial, and Internet of Things.
[0050] Second, by introducing a privacy protection scheme that combines multi - recipient encryption and differential privacy, the present invention significantly enhances the security of data privacy. This method can effectively prevent data leakage and malicious attacks, especially suitable for sensitive data scenarios such as finance and healthcare, and also complies with strict global data protection regulations such as the GDPR. This not only reduces the legal risks faced by enterprises due to privacy issues but also significantly improves their brand reputation and market competitiveness.
[0051] Based on traditional encryption schemes, the present invention optimizes multi - recipient encryption technology, significantly reducing the communication overhead and computational resource consumption in large - scale distributed federated learning. By reducing bandwidth requirements and computational complexity, enterprises can deploy privacy protection systems at a lower cost, especially showing significant economic advantages in multi - user and large - scale scenarios.
[0052] Although there has been some development in federated learning privacy protection technology, most of them are currently limited to the single application of differential privacy or traditional encryption, and the protection ability against malicious clients is weak. The present invention first introduces multi - recipient encryption technology into federated learning, achieving a balance among privacy, security, efficiency, and resource consumption through a dual privacy protection mechanism. This breakthrough technical solution fills the technical gap in the field of privacy - protected federated learning at home and abroad.
[0053] The technical solution of the present invention takes into account the balance between privacy protection and model performance, providing a new solution idea for the large - scale application of federated learning in sensitive data scenarios. Its efficient privacy protection mechanism and excellent scalability can meet the actual needs of fields such as finance, healthcare, and intelligent manufacturing, bringing direct application value and economic benefits to enterprises and users.
[0054] Traditional technologies generally believe that there is an irreconcilable contradiction between privacy protection and model performance. However, the present invention proposes a comprehensive solution by combining differential privacy and multi - recipient encryption technology. By dynamically adjusting the differential privacy budget, the present invention successfully achieves a win - win situation between privacy protection and model training performance, breaking through the limitations of traditional thinking and the inherent biases of privacy protection technology.
[0055] The innovative technical solution of the present invention not only provides a solution path for current privacy - protected federated learning but also lays a foundation for the design and optimization of future large - scale distributed machine learning systems. The balanced advantages of its systematic privacy protection method in terms of security, efficiency, and scalability will become the core driving force for the application of federated learning in sensitive data environments, providing important reference and guidance for subsequent technological development. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1It is the flowchart of the systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy provided by the embodiments of the present invention;
[0057] Figure 2 It is the architecture diagram of the systematic privacy protection federated learning system based on multi-receiver encryption and differential privacy provided by the embodiments of the present invention;
[0058] Figure 3 It is the flowchart of applying multi-receiver encryption to federated learning provided by the embodiments of the present invention;
[0059] Figure 4 It is the performance comparison diagram after adding noise of this system and three common federated learning algorithms under the Fashion-MINIST dataset provided by the embodiments of the present invention;
[0060] Figure 5 It is the performance comparison diagram after adding noise of this system and three common federated learning algorithms under the MINIST dataset provided by the embodiments of the present invention;
[0061] Figure 6 It is the model performance table under the same privacy budget provided by the embodiments of the present invention. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.
[0063] As Figure 1 shown, the embodiments of the present invention provide a systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy, and the method includes:
[0064] S1: Initialize and distribute global model parameters
[0065] The server generates the initial global model parameters W 0 , and distributes them to all legitimate clients through multi-receiver encryption technology;
[0066] Assign a unique identity ID i to each client, and generate the corresponding encryption key SK i using the master secret key MSK and the public parameters PK;
[0067] S2: Local model training on the client side
[0068] The client trains the model based on the local private dataset D i , and generates updated parameters W t+1 ; The training process is based on stochastic gradient descent (SGD) or its variants, and the formula is as follows:
[0069]
[0070] where W t is the current model parameter, η is the learning rate, and L(·) is the loss function, and ∇L(·) is the gradient of the loss function;
[0071] S3: Adding Privacy Perturbation to the Client Model Parameters
[0072] The client adds differential privacy noise ξ to the locally updated parameters to prevent potential reverse attacks. The noise follows the Laplace distribution, and the formula is as follows:
[0073]
[0074] where Δ is the sensitivity and ∈ is the privacy budget. By dynamically adjusting ∈ and Δ, the privacy protection and model performance can be flexibly balanced in different scenarios;
[0075] S4: Uploading the Perturbed Parameters by the Client
[0076] The client uploads the perturbed parameter ΔW' i to the server;
[0077] S5: Global Aggregation Module
[0078] The server receives the perturbed parameters from all clients and calculates the global model update parameter:
[0079]
[0080] where S is the set of filtered clients;
[0081] S6: The Server Uses Multi-Recipient Encryption to Distribute the Global Model
[0082] The server uses identity-based multi-recipient encryption technology (MR-IBE) to encrypt the updated global model parameter W t+1 for multi-recipient encryption, generating the ciphertext C G , ensuring that only authorized recipients can decrypt it; the encryption process is as follows:
[0083] C G = Encrypt(PK, {ID i}, W t+1 )
[0084] The ciphertext C GDistribute to authorized clients, and the client uses the private key to decrypt and update the local model. The ciphertext supports simultaneous distribution to multiple recipients, avoiding the redundant operation of generating ciphertext for each recipient separately in traditional encryption methods, significantly reducing communication overhead; while ensuring the secure distribution of the global model; the client decrypts and updates the local model parameters through the private key.
[0085] The server first generates the global model initial parameters W 0 , and uses multi-receiver encryption technology to 0 Afterwards, a unique ID is assigned to each legitimate client. i , and use the master key MSK and public parameter PK to generate the corresponding encryption key SK i During this process, the server strictly manages the initial model parameters and client keys to ensure the correct generation and distribution of encryption keys, providing security for subsequent data transmission.
[0086] Each client is based on its local private dataset D i Perform model training and use stochastic gradient descent (SGD) or its variants to update model parameters. Specifically, each client updates the model parameters W t Based on the loss function L(W t ;D i ) Then according to the formula Calculate the updated model parameters. This process involves the extraction and numerical calculation of gradient information in the signal data to ensure the accuracy and efficiency of the training process.
[0087] After the client completes the local model update, in order to prevent reverse attacks, each client updates the data ΔW in the gradient i Add Laplace noise ξ to make the updated parameter ΔW i ′=ΔW i +ξ. The noise ξ follows the Laplace(0,Δ / ε) distribution, where Δ is the sensitivity and ε is the privacy budget. By dynamically adjusting the noise parameters, the client can achieve privacy protection of sensitive information during signal data processing while ensuring model training performance.
[0088] Each client will process the perturbation parameter ΔW after differential privacy processing i 'Uploaded to the server through a secure communication channel. At this stage, the signal data transmission is encrypted and the network protocol is used to ensure the integrity and confidentiality of the transmission process. When receiving the signal data, the server performs a preliminary check on the uploaded data to ensure that the data transmitted by each client meets the predetermined format and integrity requirements.
[0089] The server performs global aggregation on all client perturbation parameters received, and obtains the global model update parameter ΔW through mathematical operations such as averaging. This aggregation process integrates the local model information of each client at the signal data level, ensuring a certain degree of robustness and consistency of the data during the fusion process. The aggregated global model parameters are updated subsequently to form a new global model W t+1 。
[0090] After the global model update is completed, the server uses identity-based multi-receiver encryption technology to encrypt the updated model parameters to generate ciphertext. This ciphertext supports simultaneous distribution to multiple authorized clients, effectively reducing the redundant calculation of generating ciphertext for each client separately. After receiving the ciphertext, the client uses its respective private key SK i to decrypt and obtain the updated global model parameters, and uses them for subsequent local model updates, thus completing the entire process of secure distribution of the global model and signal data processing. As Figure 2 shown, an embodiment of the present invention provides a systematic privacy protection federated learning device based on multi-receiver encryption and differential privacy for the above-mentioned systematic privacy protection federated learning method based on multi-receiver encryption and differential privacy. The device specifically includes:
[0091] A server, responsible for initializing, aggregating, and distributing the global model;
[0092] Client devices, responsible for local model training, parameter decryption, and noise addition.
[0093] The server includes:
[0094] A calculation unit: used to perform encryption, screening, and aggregation operations;
[0095] A storage unit: stores global model parameters and client identity information;
[0096] A communication module: performs encrypted data transmission with the client.
[0097] The client device includes:
[0098] A calculation unit: used to perform local model training and decryption operations.
[0099] A storage unit: stores local datasets and model parameters.
[0100] A communication module: performs data transmission with the server.
[0101] In this embodiment, identity-based multi-receiver encryption technology is adopted to encrypt the parameter updates transmitted by the server in the federated learning system. This technology allows multiple clients to use their respective identity keys to decrypt the encrypted global model sent by the server simultaneously, thus achieving secure data transmission and storage in a multi-party participation data sharing environment.
[0102] To protect the privacy of local client data and models, the present invention adds Laplace noise during the federated learning process to construct a differential privacy mechanism. This mechanism embeds noise in the client model updates to ensure that sensitive information of individual clients is not leaked during the parameter aggregation process, while maintaining the effectiveness and accuracy of the overall model.
[0103] The server aggregates the differentially private processed model update parameters uploaded by each client. The aggregation process performs predetermined mathematical operations on the server to integrate the scattered local model information into a unified global model. This aggregation method achieves efficient model fusion and parameter update on the premise of ensuring data privacy protection.
[0104] The global model after global aggregation is encrypted again through multi-receiver encryption technology and distributed to each participating client. The implementation of the entire system includes computer devices, computer-readable storage media, and information data processing terminals. After storing the corresponding computer programs, these devices can all execute the steps of the federated learning method of the present invention to achieve systematic privacy protection federated learning with multi-party collaboration.
[0105] Experiments were conducted on the examples of the present invention. The following is the experimental part: The data of the present invention uses the public datasets Fashion-MINIST and MINIST datasets, both of which are real data. The MNIST task involves classifying handwritten digits (0-9), while the Fashion-MNIST task focuses on classifying various fashion items. The MNIST dataset contains 50,000 training samples and 10,000 test samples, each sample has 784 features, and there are a total of 10 categories. The number of training samples and test samples in the Fashion-MNIST dataset is the same as that of MNIST, which are 50,000 and 10,000 respectively, each sample also has 784 features, and there are also 10 categories.
[0106] From Figure 4 and Figure 5 it can be seen that the method of the present invention can almost be comparable to the most commonly used FedAvg, FedSGD, and FedAdam in terms of performance while strictly improving the privacy protection performance. The average accuracy after convergence is no more than 1% compared with FedAvg, FedSGD, and FedAdam. This fully illustrates that the method of the present invention has the advantage of improving privacy protection performance while maintaining high performance. FromFigure 6 It can be seen that after comparing the effects of noise addition of different algorithms under the same differential privacy budget (∈ = 2), the method of the present invention is significantly superior to other algorithms, such as FedAvg, FedAdam, and FedSGD, on the MNIST and Fashion-MNIST datasets. On the MNIST dataset, the method of the present invention achieved a test accuracy of 91% and an F1 score of 90%, exceeding 83% and 82% of FedAvg, 88% of FedAdam, and 83% of FedSGD. The recall rate of the present invention also reached 91%, showing a significant improvement. For the more challenging Fashion-MNIST dataset, the method of the present invention performed strongly, with a test accuracy of 77%, an F1 Score of 77%, and a recall rate of 75%. In contrast, the accuracy of FedAvg was 72% (F1 Score of 0.71), while the accuracies of FedAdam and FedSGD were even lower, only 66%. As Figure 6 shown, these results indicate that under the same differential privacy budget, the present invention maintains superior performance even after adding noise, effectively demonstrating its advantages in preserving accuracy and privacy.
[0107] Example 1: Privacy-Preserving Federated Learning in Medical Data Sharing
[0108] In a medical data sharing project, multiple medical institutions (such as hospitals and clinics) cooperate in federated learning to train a model for predicting the disease risk of patients. Since it involves sensitive information of patients, each institution does not want to directly share the data. The following is the specific implementation process:
[0109] 1) Global model initialization and distribution:
[0110] The central server generates the initial parameters of the global model and distributes the encrypted model parameters to each medical institution using identity-based multi-receiver encryption technology.
[0111] 2) Local model training:
[0112] Each medical institution trains the model based on its local patient dataset (such as electronic medical records) and calculates the updated parameters using the stochastic gradient descent algorithm.
[0113] 3) Privacy perturbation addition:
[0114] After each training, each institution adds noise that satisfies differential privacy (Laplace distribution noise) to the updated model parameters to ensure the privacy of patient data.
[0115] 4) Upload parameters and global aggregation:
[0116] Each medical institution uploads the perturbed model parameters to the central server. The server receives and filters the legitimate client uploads, and updates the global parameters through weighted averaging.
[0117] 5) Global parameter distribution:
[0118] The updated global model parameters are encrypted again using multi-receiver encryption technology and distributed to all medical institutions. After decryption, the institutions update their local models.
[0119] This method ensures that medical data participates in model training without leaving the local area, while protecting patient privacy and data security, and complies with data protection regulations such as the GDPR.
[0120] Example 2: Joint fraud detection system for financial institutions
[0121] Multiple financial institutions (such as banks and payment companies) need to cooperate in training a machine learning model to identify potential financial fraud. However, the data of each institution is highly sensitive (such as transaction records, customer identity information). The following is the specific implementation process:
[0122] 1) Global model initialization and distribution:
[0123] The central server initializes a global model and distributes the model parameters to all participating financial institutions through multi-receiver encryption technology.
[0124] 2) Local data training:
[0125] Each institution uses its local financial transaction dataset (such as payment records, transfer records) to train the model and calculates the updated model parameters.
[0126] 3) Differential privacy perturbation:
[0127] Each financial institution adds differential privacy noise to its updated model parameters. The noise parameter is dynamically adjusted according to the privacy budget and data sensitivity to ensure that transaction data privacy is not leaked.
[0128] 4) Upload and aggregation of perturbed parameters:
[0129] Each institution uploads the model parameters processed with privacy perturbation to the central server. The server screens and aggregates the upload parameters of all participating institutions and calculates the new global model parameters.
[0130] 5) Global model distribution and update:
[0131] The central server encrypts the updated global model parameters using multi-receiver encryption and distributes the encrypted model parameters to each financial institution. After decryption, each institution updates its local model.
[0132] Through this method, multiple financial institutions can jointly train an efficient fraud detection model while protecting user privacy, which not only improves the model performance but also ensures data security and compliance.
[0133] It should be noted that the embodiments of the present invention can be implemented through hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.
[0134] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A systematic privacy-preserving federated learning method based on multi-receiver encryption and differential privacy, characterized in that: The method comprises the following steps: Initialize global model parameters and distribute them to clients through multi-receiver encryption technology; The client trains the model based on the local private dataset and generates local update parameters; The client adds privacy perturbation noise that satisfies the Laplace distribution to the local update parameters; The client uploads the updated parameters with privacy perturbation to the server; The server aggregates global parameters based on the updated parameters received from the client; The server performs multi-recipient encryption on the aggregated global parameters and distributes them to the client.
2. The method for systematic privacy-preserving federated learning based on multi-receiver encryption and differential privacy as claimed in claim 1, characterized in that: In the step of initializing the global model parameters, the server assigns a unique identity to each client and generates a corresponding encryption key.
3. The method for systematic privacy-preserving federated learning based on multi-receiver encryption and differential privacy as claimed in claim 1, characterized in that: The privacy disturbance noise satisfies the Laplace distribution, and the generation of the noise is based on preset parameters of sensitivity and privacy budget.
4. The method for systematic privacy-preserving federated learning based on multi-receiver encryption and differential privacy as claimed in claim 1, characterized in that: The step of aggregating the global parameters uses the parameters of the client set for weighted averaging.
5. The method for systematic privacy-preserving federated learning based on multi-receiver encryption and differential privacy as claimed in claim 1, characterized in that: The multi-recipient encryption technology is an identity-based multi-recipient encryption scheme, and the ciphertext supports being distributed to multiple clients simultaneously.
6. The method for systematic privacy-preserving federated learning based on multi-receiver encryption and differential privacy as claimed in claim 1, characterized in that: The method further includes: S1: Initialize global model parameters and distribute The server generates the global model initial parameter W0 and distributes it to all legitimate clients through multi-receiver encryption technology; Assign a unique ID to each client i , and use the master key MSK and public parameters PK to generate the corresponding encryption key SK i ; S2: Client local model training The client is based on the local private dataset D i Train the model and generate updated parameters W t+1 ; The training process is based on stochastic gradient descent (SGD) or its variants, and the formula is as follows: Among them, W t is the current model parameter, η is the learning rate, L(·) is the loss function, is the gradient of the loss function; S3: Adding privacy perturbations to client model parameters The client adds differential privacy noise ξ to the local update parameters to prevent potential reverse attacks. The noise satisfies the Laplace distribution, and the formula is as follows: Where Δ is the sensitivity and ∈ is the privacy budget. By dynamically adjusting ∈ and Δ, we can flexibly balance privacy protection and model performance in different scenarios. S4: The client uploads the perturbed parameters The client will perturb the parameter ΔW' i Upload to the server; S5: Global aggregation module The server receives the perturbation parameters of all clients And calculate the global model update parameters: Among them, S is the filtered client set; S6: Server distributes the global model using multi-recipient encryption The server uses multi-receiver identity-based encryption (MR-IBE) to encrypt the updated global model parameters W t+1 Perform multi-receiver encryption to generate ciphertext C G , ensuring that only authorized recipients can decrypt; the encryption process is as follows: C G =Encrypt(PK,{ID i },W t+1 ) The ciphertext C G Distribute to authorized clients, and the client uses the private key to decrypt and update the local model. The ciphertext supports simultaneous distribution to multiple recipients, avoiding the redundant operation of generating ciphertext for each recipient separately in traditional encryption methods, significantly reducing communication overhead; while ensuring the secure distribution of the global model; the client decrypts and updates the local model parameters through the private key.
7. A device based on the systematic privacy-preserving federated learning method based on multi-receiver encryption and differential privacy as claimed in any one of claims 1 to 6, characterized in that: The device includes: A server, having a computing unit, a storage unit, and a communication module, for performing model initialization, parameter aggregation, and encrypted distribution; A client device with a computing unit, a storage unit, and a communication module for performing local model training, parameter decryption, and privacy perturbation addition.
8. The systemic privacy-preserving federated learning device based on multi-receiver encryption and differential privacy as claimed in claim 7, characterized in that: The server comprises: Computational unit: used to perform encryption, filtering and aggregation operations; Storage unit: stores global model parameters and client identity information; Communication module: Encrypted data transmission with the client.
9. The systemic privacy-preserving federated learning device based on multi-receiver encryption and differential privacy as claimed in claim 8, characterized in that: The client device comprises: Compute unit: used to perform local model training and decryption operations. Storage unit: stores local datasets and model parameters. Communication module: data transmission with the server.
10. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the systematic privacy-preserving federated learning method based on multi-receiver encryption and differential privacy as claimed in claim 1.
Citation Information
Patent Citations
Federal learning privacy protection method based on homomorphic encryption
CN113434873A
Privacy protection federated learning method, system and device in distributed environment and medium
CN116383864A
Federal learning-oriented privacy protection method
CN117294469A
Federal learning privacy protection method based on differential privacy and homomorphic encryption
CN117421762A
Federal learning method based on local random differential privacy
CN117874816A
Cited By
Federal learning-based medical data privacy protection analysis method
CN120705908A