Construction method of trusted federated hierarchical model based on adaptive mutual learning
Patent Information
- Application Number
- CN202310676066.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-08
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-06-08
AI Technical Summary
[0010]本发明的目的就是提供一种基于自适应相互学习的可信联邦分层模型的构建方法,以解决联邦学习下的异构性与可信性问题
[0036] This invention proposes a local client-side layered mechanism. It sets up an outer loop model and a local model locally, abandoning the global model construction process of traditional frameworks. The outer loop model replaces the local model for parameter passing and secure aggregation, and learns from the local model instead of overriding it. This local client-side layered mechanism ensures that the real parameters of the local model are retained locally, while the outer loop model learns from the local model to handle data transmission, preventing the real parameters from being exposed to insecure channels and thus avoiding inference attacks targeting privacy datasets. Furthermore, the mutual learning setup makes model construction more flexible and can more effectively address data and model heterogeneity.
Smart Images

Figure CN116861992B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a computer system based on a biological model, and more specifically, to a method for constructing a trusted federated hierarchical model based on adaptive mutual learning. Background Technology
[0002] With the advent of the information age, machine learning technology has also ushered in new opportunities and development. However, because user terminal devices store a large amount of personal information and sensitive data, which may include sensitive information such as personal preferences, behaviors, and biometrics, third parties may encounter privacy risks during data collection, modeling, and publication. To avoid these threats, it is necessary to take measures to protect personal information and ensure the fairness and transparency of machine learning; therefore, how to securely handle data has become a critical issue.
[0003] As a novel distributed machine learning technology, federated learning effectively alleviates the aforementioned problems. Unlike traditional machine learning processes, federated learning requires user data to be stored locally, avoids privacy leaks by passing gradients or model parameters, and uses parameter aggregation to build a global model. This allows for privacy and security while maintaining model accuracy.
[0004] Because federated learning focuses on integrating distributed clients to obtain a higher-quality global model, it cannot effectively capture the individual characteristics of participants. Furthermore, in complex federated application scenarios, sample databases are typically collected and integrated from different users, terminals, and systems. Therefore, while the data from each client participating in training are independently distributed, they do not adhere to the same sampling method, leading to data heterogeneity. Moreover, each client has different application requirements for building its local model, necessitating personalized differences between the global and local models. Additionally, within the federated framework, client devices differ in computing power and communication speed, and their computational load tolerance varies. Therefore, the construction of local models requires comprehensive consideration, i.e., the model heterogeneity problem.
[0005] The existence of heterogeneity can lead to a decline in the global model's inference or classification capabilities. Current solutions for heterogeneity issues mostly aim to achieve faster model convergence by modifying the framework process, hoping to update the convergence direction between heterogeneous models to make model aggregation more efficient. These approaches include sharing common datasets, excluding clients with abnormal convergence, and constructing new loss functions. Research on improving the objective function attempts to make the parameter optimization direction of each client as close as possible to the optimization direction of the global parameters, thereby accelerating convergence and approaching the effect of centralized data training. Other works address performance degradation by improving aggregation methods. While most solutions are indeed effective in heterogeneous environments, showing significant improvements in model accuracy and aggregation efficiency, most only address data heterogeneity, neglecting the limitations imposed by client-specific heterogeneity, i.e., model heterogeneity. Therefore, a key challenge in addressing heterogeneity is how to collaboratively solve both data and model heterogeneity issues.
[0006] Through continuous improvement and development of federated learning, researchers have found that the credibility within this framework is extremely fragile, stemming from two aspects: privacy and robustness. Credibility requires federated learning systems to ensure model quality and reliability while protecting data privacy and defending against malicious attacks.
[0007] Regarding privacy, privacy-preserving computation within the framework is not entirely secure. Malicious third parties can completely recover the training data through a few gradient steps. Therefore, the security of the framework alone cannot guarantee the privacy of sensitive data; it needs to be combined with other privacy protection techniques (such as differential privacy, homomorphic encryption, and multi-party secure computation) to defend against privacy attacks. Most existing solutions are optimizations and improvements of the above three approaches to achieve privacy protection. Among them, mainstream privacy protection schemes mainly rely on adding noise perturbations or data encryption. This has achieved good results in terms of privacy security. However, each privacy protection technique has its advantages and disadvantages, and should be used selectively according to specific situations.
[0008] Regarding robustness, machine learning algorithms in real-world applications may be subject to various targeted or untargeted malicious attacks, leading to reduced model accuracy or incorrect classification predictions, thus affecting the overall model construction. Therefore, defense against robust attacks primarily relies on general defenses that do not make specific attack assumptions and detection of malicious clients. General defenses are simpler and more convenient, but their effectiveness is weaker than malicious client detection.
[0009] In summary, the key to achieving credibility in a federated framework lies in balancing privacy and robust defense. Effective detection of malicious attacks while ensuring secure aggregation has become a primary objective that urgently needs to be achieved. Summary of the Invention
[0010] The purpose of this invention is to provide a method for constructing a trusted federated hierarchical model based on adaptive mutual learning, so as to solve the problems of heterogeneity and trustworthiness under federated learning.
[0011] This invention is implemented as follows: a method for constructing a trusted federated hierarchical model based on adaptive mutual learning, comprising the following steps:
[0012] S1. Constructing a hierarchical model based on mutual learning: The client locally initializes two models of different specifications, including a local model and a similar model. The local model is a complex model set up to ensure data accuracy. The similar model is a lightweight model built using mutual learning, which serves as an outer loop model to replace the local model in the insecure channel under the federated framework for transmission. The parameters of the outer loop model replace the true gradients or model parameters in the parameter aggregation operation on the server. After aggregation, the outer loop model returns to the local machine and performs mutual learning with the local model to continue the next round of iteration. Due to the mutual interaction and sharing of models, during model training, the local model can learn different features and representations from the outer loop model, improving the robustness and generalization performance of the model. Therefore, it can effectively solve the heterogeneity problem and improve the overall performance of the model.
[0013] S2. Adaptive Temperature Adjustment: During the unidirectional knowledge distillation process between the local model and the outer loop model, different temperature coefficients (τ1, τ2) are used for the correct and incorrect categories of the output probability after local model training to smooth the output probability and provide better guidance. The relationship between temperature coefficients τ1 and τ2 is: τ2 = [τ1-3, τ1-2]. During the mutual learning process between the local model and the outer loop model, the local model initializes temperature coefficient τ1 and applies it to the correct category; temperature coefficient τ2 applies to the incorrect category. Mutual learning is optimized because large differences in model size can lead to incorrect knowledge transfer, and the same temperature parameter cannot solve this problem. Therefore, different temperature coefficients are used for the correct and incorrect categories. Temperature coefficient τ1 represents the effectiveness of the local model as a complex model in guiding the outer loop model, and also represents the degree to which the local client participates in global aggregation.
[0014] S3. After locally pruning the model parameters of the outer loop model, Weak DP is added to the outer loop model, along with Gaussian noise with a variance of 0.01–0.03, to perturb the sensitive data. During this process, the client not only applies norm boundaries but also obfuscates the actual parameter data.
[0015] Furthermore, the mutual learning in step S1 involves multiple models learning collaboratively, each with its own dataset and model structure. They then exchange parameters and features, providing each other with information and feedback to obtain a model with higher accuracy and generalization.
[0016] Furthermore, the specific operation method of step S1 is as follows:
[0017] S1-1 Client initializes local model L i and outer loop model O i The size and number of parameters of the outer loop model are much smaller than those of the local model; the local models of each client may be the same or different, while the size and configuration of the outer loop model should be consistent.
[0018] S1-2 Mutual learning is defined as the optimization objective of a multi-model collaborative learning problem. If there are multiple deep learning models M1, M2, M3...M... i Each deep learning model corresponds to a loss function: L1, L2, L3...L i The goal of deep learning is to minimize the weighted combination of these loss functions, and the parameters of each deep learning model must be updated:
[0019]
[0020] Where α1, α2, α3...α i The weights of each model are determined using cross-validation.
[0021] S1-3 The local model and the outer loop model learn from each other, and the parameters WO of the completed outer loop model are... i It is sent to the server.
[0022] The S1-4 server received the model parameter sequence sent by each client: {WO1,WO2,WO3...WO...} i The model is then federated and averaged to obtain the mean, which is used as the global model and sent to each client for interaction and sharing of the outer loop model.
[0023] Furthermore, the specific operation method of step S2 is as follows:
[0024] Before each round of mutual learning begins, the client first sets the temperature coefficient τ1 for the correct category.
[0025] S2-2 Adaptive temperature control randomly selects the temperature coefficient τ2 within a limited range based on the relationship between temperature coefficient τ1 and temperature coefficient τ2, τ2=[τ1-3,τ1-2].
[0026] S2-3 Constructing the loss function KL:
[0027]
[0028] Where, θ i Let i represent model i, C represent output class, P represent output probability, and KL represent the divergence of the loss function KL.
[0029] Substituting the temperature coefficients τ1 and τ2 into the loss function KL, we obtain:
[0030]
[0031] Where y represents the correct category.
[0032] The S2-4 external loop model sets the temperature coefficient of the local model to 1, minimizes the loss function, and continues until the external loop model and the local model converge.
[0033] Furthermore, the specific operation method of step S3 is as follows:
[0034] S3-1 sets a fixed positive constant as the norm α, with the norm boundary being |-α,α|; uses this norm α as the upper bound of the parameter vector; and prunes the model parameters uploaded from the outer loop model according to the norm boundary |-α,α| to ensure that the size of the model parameters does not exceed the norm α.
[0035] S3-2 adds Weak DP to the outer loop model after parameter trimming, along with Gaussian noise with a variance of approximately 0.01–0.03, to perturb sensitive data. The amount of noise can be adjusted appropriately without significantly affecting the accuracy of the outer loop model. A suitable amount of Gaussian noise can effectively obfuscate the original data and improve the model's defense against backdoor attacks.
[0036] This invention proposes a local client-side layered mechanism. It sets up an outer loop model and a local model locally, abandoning the global model construction process of traditional frameworks. The outer loop model replaces the local model for parameter passing and secure aggregation, and learns from the local model instead of overriding it. This local client-side layered mechanism ensures that the real parameters of the local model are retained locally, while the outer loop model learns from the local model to handle data transmission, preventing the real parameters from being exposed to insecure channels and thus avoiding inference attacks targeting privacy datasets. Furthermore, the mutual learning setup makes model construction more flexible and can more effectively address data and model heterogeneity.
[0037] This invention extends and optimizes mutual learning. Addressing data heterogeneity, it sets up mutual learning between the outer loop model and the local model to achieve personalization. The use of mutual learning allows the construction of the local model to be tailored to both the local dataset and the feature knowledge of other client models. Furthermore, considering model heterogeneity, this invention differentiates model sizes. First, the size of the outer loop model is set smaller than the local models on each client. This reduces communication overhead for model parameters, and smaller models are easier to detect malicious attacks, facilitating the implementation of protective measures during data transmission. Second, the size of the local models can vary between clients, allowing users to design model sizes according to their needs, making the model structure more flexible and diverse. The accuracy of mutual learning between different models is guaranteed by an adaptive temperature adjustment algorithm, making model construction more effective and stable.
[0038] To address the credibility issue of the framework, this invention employs differential privacy technology to protect the parameters of the outer loop model. From a privacy perspective, differential privacy has a rigorous mathematical definition to prove its security. Furthermore, the relatively small size of the outer loop model facilitates the inclusion of noise. From a robustness perspective, the addition of noise prevents certain targeted attacks from succeeding. Taking backdoor attacks as an example, the addition of noise effectively prevents backdoor attacks, making the model more robust, and the model's accuracy reaches the expected level.
[0039] This invention's trusted federated hierarchical model exists on the client side as an outer loop model and a local model. The outer loop model performs data transmission on behalf of the local model, and the outer loop model and the local model communicate using an optimized mutual learning algorithm. This optimized mutual learning algorithm supports interaction between models of different sizes and can exchange more accurate features and knowledge, improving the training performance of both models. The outer loop model performs Weak DP injection between uploaded model parameters, providing better protection for the parameter data. This invention's trusted federated hierarchical model effectively solves the heterogeneity and trust issues within a federated framework. Attached Figure Description
[0040] Figure 1 This is a flowchart comparing the processes of knowledge distillation and mutual learning.
[0041] Figure 2 This is a schematic diagram of the structure of the trusted federated hierarchical model constructed in this invention. Detailed Implementation
[0042] The invention will now be described in further detail with reference to the accompanying drawings.
[0043] like Figure 2As shown, the method for constructing the trusted federated hierarchical model of the present invention includes the following steps:
[0044] S1. Constructing a hierarchical model based on mutual learning: The client initializes two models of different specifications locally, including a local model and a similar model. The local model is a complex model set to ensure data accuracy, while the similar model is a lightweight model built using mutual learning, used as the outer loop model to replace the local model's transmission in the insecure channel under the federated framework. Mutual learning involves multiple models collaboratively learning, each with its own dataset and model structure. By exchanging parameters and features, they provide information and feedback to each other, resulting in a model with stronger accuracy and generalization. The parameters of the outer loop model replace the true gradients or model parameters in the parameter aggregation operation on the server. After aggregation, the outer loop model returns to the local machine to learn from the local model again, thus continuing the next round of iteration. The specific operation is as follows:
[0045] S1-1 Client initializes local model L i and outer loop model O i outer loop model L i The size and number of parameters of the local model are much smaller than those of the local model. The local model L between each client... i The constructions can be the same or different, but the outer loop model O i The scale and configuration need to be consistent.
[0046] S1-2 Mutual learning is defined as the optimization objective of a multi-model collaborative learning problem, assuming there are multiple deep learning models M1, M2, M3...M i Each deep learning model corresponds to a loss function: L1, L2, L3...L i The goal of deep learning is to minimize the weighted combination of these loss functions, and the parameters of each deep learning model must be updated:
[0047]
[0048] Where α1, α2, α3...α i These are the weights of each model, which can be determined through methods such as cross-validation.
[0049] S1-3 Local Model L i With the outer loop model O i Through mutual learning, the outer loop model parameters WO are completed. i It is sent to the server.
[0050] The S1-4 server received the model parameter sequence sent by each client: {WO1,WO2,WO3...WO...} iThe data is then federated and averaged to obtain the mean, which is used as the global model before being sent to each client.
[0051] Due to the interaction and sharing among models, each model can learn different features and representations from other models during training, improving the robustness and generalization performance of the models. Therefore, it effectively solves the heterogeneity problem and improves the overall performance of the models.
[0052] S2, Adaptive Temperature Regulation: In the local model L i Outer loop model O i During the process of one-way knowledge distillation, for the local model L i After training, different temperature coefficients (τ1, τ2) are used to smooth the output probabilities for the correct and incorrect classes. The relationship between temperature coefficients τ1 and τ2 is: τ2 = [τ1-3, τ1-2]. In the local model L... i With the outer loop model O i During the mutual learning process, the local model L i Initialize the temperature coefficient τ1 and apply it to the correct category; apply the temperature coefficient τ2 to the incorrect category.
[0053] like Figure 1 As shown, mutual learning is similar to knowledge distillation, with a temperature parameter τ used to adjust the guidance effect. This invention optimizes mutual learning because large differences in model size can lead to incorrect knowledge transfer, and the same temperature parameter cannot solve this problem. Therefore, different temperature coefficients are used for correct and incorrect categories. The temperature coefficient τ1 represents the effectiveness of the local model, as a complex model, in guiding the outer loop model, and also represents the degree to which the local client participates in global aggregation.
[0054] The specific operation method of step S2 is as follows:
[0055] Before each round of mutual learning begins, the client first sets the temperature coefficient τ1 for the correct category.
[0056] S2-2 Adaptive temperature control randomly selects the temperature coefficient τ2 within a limited range based on the relationship between temperature coefficient τ1 and temperature coefficient τ2, τ2=[τ1-3,τ1-2].
[0057] S2-3 Constructing the loss function KL:
[0058]
[0059] Where, θ i Let i represent model i, C represent output class, P represent output probability, and KL represent the divergence of the loss function KL.
[0060] Substituting the temperature coefficients τ1 and τ2 into the loss function KL, we obtain:
[0061]
[0062] Where y represents the correct category.
[0063] S2-4 External Circulation Model O i For the local model L i The temperature coefficient is set to 1, and the loss function is minimized until the outer loop model O. i With local model L i convergence.
[0064] S3, External Circulation Model O i After locally trimming the model parameters, in the outer loop model O i Weak dynamics (WeDP) is added, meaning the client not only applies the norm boundary but also adds Gaussian noise with a variance of approximately 0.01–0.03 to perturb sensitive data. The addition of differential privacy requires minimal assumptions, performs well in a federated environment, provides excellent defense against model theft attacks, and maintains model accuracy. Its specific operation is as follows:
[0065] S3-1 sets a fixed positive constant as the norm α, with the norm boundary being |-α,α|; uses this norm α as the upper bound of the parameter vector; and prunes the model parameters uploaded from the outer loop model according to the norm boundary |-α,α| to ensure that the size of the model parameters does not exceed the norm α.
[0066] S3-2 adds Weak Dynamic Dispersion (WDP) to the outer loop model after parameter trimming, along with Gaussian noise with a variance of approximately 0.01–0.03, to perturb sensitive data. Appropriate amounts of Gaussian noise can effectively obfuscate the original data and improve the model's defense against backdoor attacks. Compared to strictly defined differential privacy, Weak Dynamic Dispersion introduces less noise and has less impact on model accuracy. Using appropriate Gaussian noise can perturb sensitive data, thereby improving the model's defense against backdoor attacks.
Claims
1. A method for constructing a trustworthy federated hierarchical model based on adaptive mutual learning, characterized in that, Includes the following steps: S1. Construct a hierarchical model based on mutual learning: The client initializes two models of different specifications locally, including a local model and a similar model; the local model is a complex model set up to ensure data accuracy; the similar model is a lightweight model built using mutual learning, which serves as an outer loop model to replace the local model in the insecure channel under the federated framework for transmission; the parameters of the outer loop model replace the true gradient or model parameters in the parameter aggregation operation performed on the server. After aggregation, the outer loop model returns to the local machine and learns from the local model again to continue the next round of looping. S2, Adaptive Temperature Adjustment: During the unidirectional knowledge distillation process of the local model and the outer loop model, different temperature coefficients are selected for the correct and incorrect categories of the output probability after the local model is trained. , To smooth the output probability; temperature coefficient With temperature coefficient The relationship is: During the mutual learning process between the local model and the external circulation model, the local model improves its understanding of the temperature coefficient. Initialize and apply to the correct category; temperature coefficient Apply to error categories; S3. After locally pruning the model parameters of the outer loop model, Weak DP is added to the outer loop model, along with Gaussian noise with a variance of 0.01-0.03, to perturb the sensitive data. The specific operation method for step S1 is as follows: S1-1 Client initializes local model and outer loop model The size and number of parameters of the outer loop model are much smaller than those of the local model; the local models of each client may be the same or different, while the size and configuration of the outer loop model should be consistent. S1-2 mutual learning is defined as the optimization objective of a multi-model collaborative learning problem, where there are multiple deep learning models. Each deep learning model corresponds to a loss function: The goal of deep learning is to minimize the weighted combination of these loss functions, and the parameters of each deep learning model must be updated: , in, These are the weights of each model, determined through cross-validation. S1-3 The local model and the outer loop model learn from each other, and the parameters of the completed outer loop model are... It was sent to the server; The S1-4 server received the model parameter sequences sent by each client: { The model is then federated and averaged to obtain the mean, which is used as the global model and sent to each client for model interaction and sharing. The specific operation method for step S2 is as follows: Before each round of mutual learning begins, the client first sets the temperature coefficient for the correct category in S2-1. ; S2-2 Adaptive Temperature Control Based on Temperature Coefficient With temperature coefficient Relationship Randomly select the temperature coefficient within the specified range. ; S2-3 Constructing the loss function KL: , in, Representative model , Represents the output category. Represents the output probability. Calculate the divergence of the loss function KL; Temperature coefficient and temperature coefficient Substituting the loss function KL, we get: , in, Correct category; The S2-4 external loop model sets the temperature coefficient of the local model to 1, minimizes the loss function, and continues until the external loop model and the local model converge.
2. The method for constructing a trusted federated hierarchical model based on adaptive mutual learning according to claim 1, characterized in that, The mutual learning mentioned in step S1 involves multiple models learning collaboratively, each with its own dataset and model structure. They then exchange parameters and features, providing each other with information and feedback to obtain a model with higher accuracy and generalization.
3. The method for constructing a trusted federated hierarchical model based on adaptive mutual learning according to claim 1, characterized in that, The specific operation method for step S3 is as follows: S3-1 sets a fixed positive constant as the norm. The norm boundary is: ; with this norm As an upper bound for the parameter vector; according to the norm boundary The model parameters uploaded from the outer loop model are pruned to ensure that the size of the model parameters does not exceed the norm. ; S3-2 adds Weak DP to the outer loop model after trimming the model parameters, and adds Gaussian noise with a variance of 0.01-0.03 to perturb the sensitive data.
Citation Information
Patent Citations
Federal target detection method and system based on knowledge distillation
CN114863092A
Client member reasoning attack method based on federated distillation learning framework
CN116187469A
Federal learning method and system for classification prediction of connection data of Internet of Vehicles terminal
CN116227631A