Personalized federal learning method based on margin constraint and credibility aggregation
By adopting technical means of feature distillation alignment, margin constraint cosine normalization, dispersion regularization and trust index aggregation in federated learning, the problems of insufficient accuracy and convergence speed, difficulty in identifying malicious clients, and neglecting fairness of small-scale clients in federated learning, the model is efficient, robust and fair.
Patent Information
- Application Number
- CN202510350086.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
AI Technical Summary
The existing federated learning technology is difficult to balance accuracy and efficient convergence in a highly heterogeneous environment, and it lacks the ability to identify and defend malicious clients, ignoring the fairness of small-scale or specially distributed clients.
A personalized federated learning method based on margin constraints and trust aggregation is adopted, and the generalization ability, robustness and fairness of the model are enhanced through technical means such as feature distillation alignment, margin constraint cosine normalization, dispersion regularization and trust index aggregation.
The model is stable convergence and high precision in a high heterogeneous environment, significantly improving the defense ability against malicious attacks, and ensuring performance fairness among multiple clients.
Smart Images

Figure CN120181264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a personalized federated learning method based on margin constraint and confidence aggregation. Background Art
[0002] In today's large-scale and multi-scenario data era, due to its privacy advantage of keeping data local, federated learning has become an important research direction in fields such as healthcare, Internet of Things, and financial security. Federated learning realizes collaborative modeling without sharing raw data by locally training on multiple clients respectively and then uploading the model updates to the server for aggregation. With the help of the distributed collaborative training mechanism, federated learning can be widely applied in sensitive scenarios such as healthcare, Internet of Things, and financial risk control, avoiding the security and privacy threats caused by centralized storage of massive private data on a central server.
[0003] However, federated learning faces multiple challenges in large-scale real-world environments. The primary difficulty is the non-independent and identically distributed problem, that is, the data held by each client may have significant differences in terms of class composition, data volume scale, or feature distribution, resulting in the difficulty for the aggregated global model to be simultaneously adapted to all clients. If the traditional average aggregation method is adopted, phenomena such as slow model convergence and unstable test accuracy may occur due to gradient direction conflicts or local overfitting. In addition, the data imbalance between different clients is likely to be amplified in a non-independent and identically distributed environment, leading to a significant decrease or even failure of the model accuracy on some clients. Besides statistical heterogeneity, malicious client attacks also pose a severe threat to the federated learning system. Common attack methods include label poisoning attacks, that is, tampering with or flipping some labels in the local training data to guide the global model to learn wrong classification decisions; and random update interference, where attackers upload random gradients or malicious parameters to cause the server aggregation direction to deviate from the normal convergence path. These attacks not only lead to a decline in the generalization performance of the global model but sometimes even cause the complete collapse of the convergence process. Therefore, how to simultaneously address the generalization challenges brought by non-independent and identically distributed data and the robustness risks caused by malicious attacks while protecting data privacy has become the core issue urgently needed to be solved in the field of federated learning.
[0004] To alleviate the insufficient adaptation of a single global model in heterogeneous environments, researchers have proposed the concept of personalized federated learning, allowing each client to locally retain certain differentiated models or features to achieve better performance in non-independent and identically distributed scenarios. Existing work has explored from various perspectives, mainly summarized into the following categories of ideas:
[0005] (1) Meta - learning or regularization - based methods: Representatives of such methods include Per - FedAvg, Ditto, FedProx, etc. They usually introduce a "two - layer" or "multiple" optimization mechanism during training. On the one hand, they maintain the iterative update of the global model to capture common features across clients; on the other hand, for each client, they introduce a local model or a local regularization term to make the local training fit the specific distribution of that client. Such meta - learning or regularization methods usually achieve good results in standard non - independent and identically - distributed environments and are relatively easy to implement. However, in extremely heterogeneous or high - risk scenarios, there may still be problems such as unstable training or weak defense against malicious interference.
[0006] (2) Distillation and contrast - learning - based methods: In a non - independent and identically - distributed federated learning environment, some studies choose to enhance the model's adaptability to diverse distributions through knowledge distillation or contrast learning. Typical methods such as MOON, FedKD, etc. generally introduce an alignment process between the distillation teacher model and the student model on the client or server side to inherit the general representation after global aggregation while maintaining local data features. These methods have shown good cross - domain adaptation ability and model compactness in theory and actual measurements. However, their defense ability against malicious attacks is usually poorly designed, and in extremely heterogeneous data distributions, the distillation effect is easily affected by label biases of different clients.
[0007] (3) Clustering or similarity - based methods: When the number of clients in the federated system is large and the distribution types are diverse, clients with similar data distributions can be grouped together for separate training, and then cooperation is carried out at the group or global level. Typical methods of this idea include CFL, etc. Such methods are relatively suitable for use in large - scale federated learning scenarios. However, during actual deployment, an additional clustering structure needs to be maintained, and both the communication and scheduling complexities are high. And when facing interference from malicious clients, it is difficult to accurately identify and defend against highly adversarial updates only relying on similarity metrics.
[0008] Under the above three common ideas, different algorithms have effectively alleviated some of the impacts brought by non - independent and identically - distributed data, making personalized federated learning achieve initial results in many scenarios. However, when the data distribution differences between clients further intensify, or there are various forms of malicious attacks, many methods still show insufficient defense ability or tend to ignore those clients with small data scales and particularly sparse distributions, thus causing problems of unfair overall performance. To further break through this bottleneck, it is necessary to take into account robust defense against potential attacks and performance balance among multiple clients in the personalized mechanism.
[0009] As federated learning is implemented in more open and heterogeneous environments, how to ensure robustness and fairness has become another key issue. Robustness research mainly focuses on whether the federated system can maintain the normal convergence and high accuracy of the model when it is interfered by malicious or abnormal clients. Most existing Byzantine defense algorithms start from the aggregation level, such as using geometric median or empirical loss screening methods to eliminate malicious gradients. However, these methods often rely too much on the statistical features during aggregation and lack further means of identifying and defending against abnormal samples or malicious distributions during the local training stage. If attackers tamper with labels or local features in a more concealed way, such aggregation strategies may be difficult to detect and eliminate in a timely manner. Corresponding to this is the fairness research of federated learning, that is, in a significantly heterogeneous client environment, how to avoid "averaging" or "marginalizing" a small number of disadvantaged nodes. Some methods try to improve the overall performance balance through multi-objective gradient descent or dynamic client screening, but if there is no supporting robust defense mechanism, they may still face the risk of model collapse in the face of high adversarial attacks. There are also studies that propose to increase the weights of small-scale clients or clients with special distributions during aggregation, but without a fine-grained attack detection or adaptive regulation strategy, it is very likely that malicious clients will gain more influence due to misjudgment, thus reversing the damage to the global model. Therefore, in an extremely heterogeneous and high-risk environment, making federated learning simultaneously possess robust robustness and good fairness remains a very difficult systematic problem.
[0010] There are many existing personalized federated learning solutions, which often address the issue of data distribution heterogeneity among clients at different levels. Typical methods include Per-FedAvg, MOON, and Ditto, etc. Per-FedAvg introduces the idea of meta-learning on the basis of FedAvg, combining global optimization with local fast adaptation, enabling clients to better fit local data through a few gradient updates after receiving the global model. MOON, on the other hand, uses contrastive learning loss to alleviate the gradient conflict among clients by comparing the differences between local models and the global model in the representation space. Ditto optimizes the global model and each local model in parallel in the loss function, thus taking into account both overall accuracy and personalized requirements. These methods have, to a certain extent, solved the instability brought about by non-independent and identically distributed data and can achieve better performance in conventional heterogeneous scenarios. However, in the face of more complex or extreme non-independent and identically distributed data and high-risk environments with various malicious interferences, relying solely on meta-learning or contrastive loss may not be sufficient to provide adequate robust defense means. In addition, for clients with particularly rare distributions or limited data scales, these methods are also prone to significant performance deviations, and overall fairness is difficult to guarantee. In recent years' research on personalized federated learning, FedPAC has also received attention. At the algorithm level, this method mainly attempts to enable clients to use the global classifier or general feature vector for auxiliary alignment while maintaining their own personalized updates by adding regularization terms and constraint terms in the local training stage. Specifically, FedPAC maintains a global feature centroid on the server to help local models reduce representation drift, and then combines the collaborative classifier among multiple clients to alleviate the ambiguity of the classification decision boundary. In addition, FedPAC penalizes the differences between local models and global features in the design of the loss function, thus strengthening the consistency in the federated process to a certain extent.
[0011] In addition to personalized federated learning algorithms, many research results have also emerged in the direction of robustness and fairness. From the perspective of robustness, common Byzantine defense aggregation methods such as geometric median, entropy-based filtering, or empirical loss reweighting, etc., can screen out suspicious updates on the server side, thus resisting malicious clients to a certain extent. However, these methods pay more attention to statistical filtering at the aggregation layer and have little to say about how to handle abnormal samples or attack interferences in local client training. Once the attack methods are diverse or concealed enough, simple aggregation screening is difficult to cope with. At the same time, the fairness research of federated learning often focuses on how to balance the final model accuracy of clients with different data volumes and different distribution types. There are already some multi-objective optimization, differential weighting, or gradient alignment strategies to reduce the performance variance of clients, but without a refined detection and defense mechanism against malicious clients, it is still difficult to continuously maintain the high quality and fair distribution of the model in complex environments.
[0012] Although federated learning has shown great potential in protecting data privacy, security compliance, etc. in recent years, the existing personalized federated learning and robust aggregation technologies still have the following main technical drawbacks:
[0013] (1) It is difficult to balance accuracy and efficient convergence in a highly heterogeneous environment: Most federated learning schemes based on traditional average aggregation or simple regularization tend to have problems such as a slow convergence rate or continuously high error during the training process when faced with extremely unbalanced data distributions and extremely skewed categories. On the one hand, severe gradient conflicts among clients make it difficult for the global model to find a consistent optimization direction during aggregation, resulting in high communication costs and an increase in the number of repeated iterations. On the other hand, a single global model or ordinary personalized methods are also prone to an increase in the misjudgment rate and model instability on some clients.
[0014] (2) Insufficient ability to identify and defend against malicious clients: Some existing methods use aggregation screening based on statistical features to combat Byzantine attacks. However, there are practical limitations in terms of computational complexity and detection accuracy: If the attack method is sufficiently concealed, simple outlier detection at the aggregation end is likely to fail; in a large-scale federated environment, maintaining a more complex filtering process will also significantly increase communication costs and computational overhead. At the same time, due to the lack of differential processing of suspicious samples locally on the client side, the model may overfit known abnormal patterns, resulting in insufficient generality in defending against unknown attacks and a sharp drop in the accuracy of the model when facing new attacks.
[0015] (3) Ignoring the fairness of small-scale or specially distributed clients: Many personalized federated learning methods only focus on improving the overall average accuracy or achieving good performance on mainstream clients, but often fail to take into account clients with scattered distributions, extremely small amounts of data, and limited contributions despite being online for a long time, resulting in a significant widening of the test performance gap between different nodes. From a technical perspective, if these disadvantaged clients are not differentiated or given the same weight without a strategy, it is extremely easy to lead to inefficient or error-prone model optimization for them; while blindly amplifying their influence may cause fluctuations in the overall convergence process. In addition, problems such as high consumption of computing power and bandwidth resources of edge devices themselves, or high execution costs due to excessive local repeated training rounds of the algorithm, have not been systematically solved. Summary of the Invention
[0016] In order to solve the problems existing in the prior art, namely, it is difficult to balance accuracy and efficient convergence in a highly heterogeneous environment, insufficient ability to identify and defend against malicious clients, and ignoring the fairness of small-scale or specially distributed clients, the present invention provides a personalized federated learning method based on margin constraint and trust degree aggregation.
[0017] From the dual perspectives of personalization and robustness, the present invention systematically solves the three major problems of generalization ability, attack defense, and performance fairness faced by federated learning in heterogeneous environments by using a joint mechanism of feature distillation alignment, margin-constrained cosine normalization, and dispersion regularization, and combining confidence index aggregation, providing new ideas and a practical framework for truly implemented federated learning solutions.
[0018] The objective of the present invention is to fundamentally balance the three major requirements of "generalization performance, robust defense, and fair distribution":
[0019] Objective 1: Strengthen accurate representation and stable convergence in heterogeneous environments;
[0020] By combining "feature distillation alignment" with "margin-constrained cosine normalization", clients can inherit global prior knowledge while retaining the differential expressions of their local data. Together with the "dispersion regularization" of suspicious samples, the risk of overfitting in small batches or single-class samples is reduced, and the high misjudgment of rare distributions is avoided, improving the overall convergence efficiency and accuracy.
[0021] Objective 2: Enhance the defense and adaptation capabilities against malicious clients;
[0022] At the local client side, through differential cosine constraints and dispersion management means, more flexible distance or margin regulation is imposed on potential malicious or abnormal samples to avoid the model overfitting to a small number of poisoning patterns. At the same time, "confidence index aggregation" is adopted on the server side, which can dynamically reduce the weight of suspicious clients on the global model, greatly improving the recognition accuracy of covert attacks and reducing the waste of resources caused by the sharp increase in computational volume.
[0023] Objective 3: Improve the performance balance among multiple clients;
[0024] Through an adaptive weighting strategy for high-quality clients and weak clients at the global aggregation level, clients with a small amount of data or extreme distributions can obtain reasonable support and protection, while the influence range of malicious updates is suppressed. This ensures the overall model accuracy while minimizing the performance gap among clients as much as possible, achieving a higher level of fairness and robustness.
[0025] In summary, the present invention aims to address the technical shortcomings of slow convergence, insufficient defense, and lack of fairness in the prior art in high-heterogeneity and high-risk federated environments, thereby achieving efficient cooperation and secure training for multi-scenario and multi-type clients in a systematic way, and providing a more perfect solution for the implementation of federated learning in practical applications.
[0026] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0027] A personalized federated learning method based on margin constraint and confidence aggregation provided by the present invention specifically includes the following steps:
[0028] S1: Feature distillation alignment;
[0029] Through knowledge distillation at the feature level, align the global model sent by the server with the client's local model, so that the client retains the learning of common features during the update process;
[0030] S2: Margin-constrained cosine normalization;
[0031] Set two different boundary values under the cosine metric, corresponding to the trusted distribution and the suspicious distribution respectively; when the sample is judged to be in the trusted distribution, the angle between its feature vector and the fixed direction vector must meet strict boundary requirements; when the sample is judged to be in the suspicious distribution, a larger angle space is reserved to reduce overfitting to unknown attacks;
[0032] S3: Dispersion regularization;
[0033] Identify suspicious samples in each mini-batch and apply a regularization term that widens the margin to their features;
[0034] S4: Confidence index aggregation;
[0035] By dynamically adjusting the weights of each client during aggregation, clients that contribute more to the global objective obtain higher weights.
[0036] Further, in step S1, assume that the feature extraction module of the teacher model is F T (·), and the feature extraction module of the student model is F S (·); for any input sample (x, y), the client uses the teacher model and the student model to extract the features of the middle layer or the penultimate layer respectively, and the teacher features and student features obtained are respectively expressed as:
[0037] f T (x) = F T (x),
[0038] f S (x) = F S (x),
[0039] To quantify the difference between the teacher features and the student features in the feature space, define the alignment loss function L Align :
[0040] L Align (x) = γ · ||F S (x) - f T (x)|| 2 ,
[0041] Among them, γ represents the weight of the alignment loss in the total loss. The larger this loss term is, the more significant the gap between the features extracted by the local model and the global prior features. Backpropagation will force the student features to approach the teacher features by increasing the gradient.
[0042] Furthermore, in step S2, let the fixed direction vector represent the weight vector of the credible direction, regarded as the convergence center of the target distribution; define two boundaries m0 and m1 (m0, m1 ∈ [-1, 1] & m0 > m1) corresponding to the angular restrictions of the credible class features and the suspicious class features respectively; the local student feature f S (x) and the server teacher feature f T (x) similarity is defined as:
[0043]
[0044] where f S,k (x) and f T,k (x) represent the outputs of the student and the teacher on the k-th dimensional feature respectively, d represents the dimension of the feature space; the larger the similarity score(x) is, the more consistent the features of the local model and the global model are on this sample, and this sample is more likely to be judged as the target distribution.
[0045] Even further, according to the definition of the similarity between the local student feature f S (x) and the server teacher feature f T (x), the margin-constrained cosine normalization loss function on a single sample x i is defined as:
[0046] L MCCS (x i ) = βlog(1 + e [σ(μ-score(x))] ),
[0047] where β is used to adjust the weight of this loss term in the total loss, σ is used to control the steepness of the gradient; μ represents an adjustable boundary parameter. If the sample x i is judged to be a credible distribution, μ = m0; if the sample x i is judged to be a suspicious distribution, μ = m1; for a mini-batch, average over all samples:
[0048]
[0049] where N represents the number of clients.
[0050] Furthermore, in step S3, assume that the selected set of suspicious samples is then the center vector of this set is expressed as:
[0051]
[0052] Among them, represents the feature vector of the m-th suspicious sample, and M represents the number of suspicious samples.
[0053] Furthermore, according to the representation of the central vector of the suspicious sample set, calculate the average dispersion of the suspicious samples in the feature space:
[0054]
[0055] When the average dispersion δ is too small, it means that these suspicious samples are very similar to each other, and the model may fit well to this feature set during training.
[0056] Furthermore, to maintain the dispersion of these suspicious samples during training, define the dispersion regularization as:
[0057]
[0058] where ∈ represents a very small constant to prevent division by zero; when the average dispersion δ → 0, L Spread → +∞, and the gradient penalty will force the suspicious class features to separate from each other; when the average dispersion δ is large enough, indicating that the suspicious samples are already sufficiently dispersed in the feature space, the contribution of this regularization term to the loss becomes smaller.
[0059] Furthermore, in step S4, to quantify the trust index of the client in this round of training, define the preliminary weight w' i as:
[0060]
[0061] where λ d is the balance coefficient, used to trade off between tightening the discrimination ability and the dispersion ability; L Spread represents the dispersion regularization loss; L MCCS represents the margin constraint cosine normalization loss; if the client performs well in distinguishing between trusted and suspicious samples, its loss value is low, making the preliminary weight w' i amplified; otherwise, it decays exponentially.
[0062] Furthermore, according to the defined preliminary weight w' i perform normalization to obtain the aggregated weight α of the client in the r-th round of federated training i :
[0063]
[0064] where, represents the set of clients participating in the training in the r-th round, Represents the accumulation of the initial weights of all clients.
[0065] Furthermore, the server aggregates all client updates with weight α i into a new global model Θ (r+1 ):
[0066]
[0067] where, represents the model parameters updated by client i after the completion of the r-th round of local training, and Θ (r+1) represents the global model to be distributed in the (r + 1)-th round.
[0068] The beneficial effects of the present invention are:
[0069] Based on the absorption of previous research on Byzantine defense, personalized strategies, and multi-client performance balance, the present invention endeavors to start from two key aspects of "feature learning" and "aggregation scheduling" to balance the defense against known / unknown attacks and the model effects of each node, thereby achieving a more comprehensive federated learning goal. Therefore, while targeting client personalization, the present invention further increases the management of the dispersion of suspicious samples and the differential cosine constraints for trusted / suspicious classes, attempting to maintain robust convergence and overall balance in the face of high risks and large-scale heterogeneity.
[0070] Compared with the prior art, the present invention has the following advantages:
[0071] 1. Achieve high accuracy and stable convergence in non-i.i.d. scenarios;
[0072] On the client side, the present invention enhances the effective inheritance of global prior knowledge by core modules such as feature distillation alignment and margin-constrained cosine normalization. Compared with traditional single average aggregation or simple meta-learning methods, the present invention can not only maintain a stable convergence process in highly heterogeneous data scenarios but also perform better in terms of final accuracy. For clients with extreme class imbalance or a large disparity in data volume, it can also quickly adapt and achieve a high accuracy rate in global testing.
[0073] 2. Significantly improve the robustness and defense ability against malicious attacks;
[0074] By adopting a dual strategy of margin-constrained cosine normalization and dispersion regularization in the local training phase, the model can effectively distinguish between trustworthy samples and suspicious samples, and retain additional dispersion in the feature space of suspicious samples. This measure avoids the over-memorization of known attack patterns by clients and reserves sufficient adaptation space for potential new attacks. On the server side, "confidence index aggregation" further dynamically amplifies high-quality client updates and suppresses malicious or poorly performing updates, thus maintaining a high performance retention rate under various attack scenarios.
[0075] 3. Consider the performance fairness of multiple clients;
[0076] In traditional federated learning, clients with small amounts of data or extremely skewed distributions are often "swamped" by mainstream clients, resulting in a high degree of dispersion in test results and insufficient fairness. The present invention enables weak clients to inherit certain global knowledge through feature distillation alignment; and reasonably allocates weights at the global level with confidence index aggregation to maintain appropriate support for weak clients. Such a mechanism reduces the attack risk while ensuring that client models can achieve stable performance growth in most or even small data scenarios, significantly reducing the performance variance between clients.
[0077] 4. Meet the implementation requirements of multiple industries and have broad application prospects;
[0078] In view of its consideration of high heterogeneity, strong security defense, and fairness, the present invention is particularly suitable for fields such as healthcare, the Internet of Things, and financial risk control, which are highly sensitive to privacy and security and have complex and diverse data distributions. By flexibly adding and adjusting four key functional modules in the "local learning + server aggregation" architecture, a federated learning solution that is feasible for deployment, flexible for expansion, and stable in performance in real-world scenarios can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 It is a flowchart of the personalized federated learning method based on margin constraint and confidence aggregation provided by the present invention.
[0080] Figure 2 It is a flowchart of local training for clients.
[0081] Figure 3 It is the theoretical decision boundary of margin-constrained cosine normalization.
[0082] Figure 4 It is the effect of dispersion regularization. DETAILED DESCRIPTION OF THE INVENTION
[0083] The following further elaborates on the present invention in conjunction with the accompanying drawings.
[0084] A personalized federated learning method based on margin constraint and trust degree aggregation is provided by the present invention, which is used to improve the generalization ability, robustness and fairness of the model simultaneously in the scenario of non-independent and identically distributed data with potential malicious clients. The technical solution is mainly divided into two major stages: local training on the client side and aggregation on the server side, and the following four key functional modules cooperate closely:
[0085] (1) Feature distillation alignment (implemented on the client side);
[0086] (2) Margin-constrained cosine normalization (implemented on the client side);
[0087] (3) Dispersion regularization (implemented on the client side);
[0088] (4) Trust degree index aggregation (implemented on the server side).
[0089] In each round of federated training, the server first sends the latest global model to each client; subsequently, each client completes the update process involved in the above multiple key functional modules locally; finally, each client uploads the updated model or gradient information back to the server, and the server then performs trust degree index aggregation to obtain a new round of global model. After multiple rounds of cycling, when the global model converges or reaches the set threshold, the final model can be output.
[0090] The design concept of the present invention lies in effectively distinguishing between trusted and suspicious samples through margin-constrained cosine normalization and dispersion regularization on the client side, and then supplemented by trust degree index aggregation on the server side to form an overall solution for defense and scheduling from the dual levels of "local + aggregation".
[0091] A personalized federated learning method based on margin constraint and trust degree aggregation provided by the present invention, its key points and innovations are as follows:
[0092] 1. Personalized optimization strategy based on feature distillation;
[0093] In the heterogeneous scenario of federated learning, clients often struggle to fully benefit from the global model due to significant differences in data scale or class distribution. To address this issue, the present invention introduces a feature distillation alignment module in the local training of clients to more specifically compensate for the deficiencies of personalized models. Specifically, after receiving the global model, the client treats it as a teacher model and aligns it with the local model through intermediate feature layers to form a feature distillation path. Compared with previous solutions that only perform distillation at the output layer, the present invention starts from deep feature representations, can better retain global prior knowledge, and overcome the local overfitting phenomenon caused by insufficient training data or unbalanced distributions. Through this feature distillation mechanism, the generalization ability of the client model is significantly improved, and the learning of the general representation is continuously consolidated in multiple iterations, thus providing a more solid feature foundation for subsequent modules.
[0094] 2. Margin-constrained cosine normalization;
[0095] Based on traditional softmax normalization, the present invention proposes a differential cosine metric strategy specifically for trusted and suspicious samples. By imposing a stricter margin on trusted samples, their features are highly concentrated around the target class center, improving the accuracy of classification decisions. For suspicious samples, a looser cosine space is reserved to prevent the model from overfitting to local abnormal distributions. This innovation can effectively mitigate the contamination of known attack patterns on the discriminative boundary and leave enough adaptation margin for unknown or rare attacks. Margin-constrained cosine normalization significantly enhances the robustness of the present invention against malicious behaviors in non-i.i.d. environments and has a more flexible adjustable parameter design compared to existing methods, enabling refined discrimination and constraint of trusted / suspicious samples according to different scenarios.
[0096] 3. Dispersion regularization;
[0097] In the processing of potential suspicious samples, the present invention not only relies on simple masking or elimination strategies but also additionally designs dispersion regularization to handle more complex attack patterns. The basic principle is that when suspicious samples are too close in the feature space, the model is prone to highly fit them, resulting in a decline in the defense ability against other types of attacks. Therefore, in each mini-batch training, the present invention dynamically calculates the average dispersion between suspicious samples and adds a regularization penalty when the distance is too close, forcing them to spread out in the feature space. This idea breaks through the binary processing mode of traditional anomaly detection, which can only be either present or absent, and can retain the information of suspicious samples while reducing their rigid dependence on the model in multiple iterations, thereby balancing the adaptation ability to diverse malicious distributions and the generalization performance of the global model.
[0098] 4. Trust index aggregation;
[0099] On the server side, the trust index aggregation proposed by the present invention breaks the limitations of traditional average aggregation or static weighting, and measures the performance of the client at the end of local training through a dynamic exponential function. If the loss metric of a certain client is relatively low, it indicates that the quality of its model update is better, and the server assigns a higher weight to it; if both the margin constraint cosine normalization loss and the dispersion regularization loss of the client are at high values, it means that it is potentially malicious or performs poorly, and the server weakens its contribution to the global model through exponential decay. In contrast, existing Byzantine defenses or entropy-based reweighting schemes often require additional complex detection processes or global auxiliary verification data, and cannot be flexibly scheduled in large-scale heterogeneous scenarios. The trust index aggregation of the present invention greatly improves the global aggregation efficiency in high-risk environments because of its simple implementation while taking into account robustness and fairness.
[0100] After organically integrating the above four core invention points (feature distillation alignment, margin constraint cosine normalization, dispersion regularization, trust index aggregation), the present invention forms a dual defense and adaptation mechanism that echoes each other between the client and the server. The multiple loss terms on the client side ensure the reinforcement learning of trusted samples and the decentralized management of suspicious samples, while the aggregation strategy on the server side adaptively amplifies high-quality contributions and suppresses potential attacks. The two cycle and iterate, and finally enable the global model to still achieve excellent generalization performance, robust adversarial ability and performance balance among multiple clients in high-heterogeneity and potentially malicious scenarios. This systematic technical idea provides important practical value for actual Internet of Things, healthcare, financial risk control and more application scenarios where data is highly dispersed and privacy-sensitive, and also makes up for the gap in the existing technology in simultaneously considering the three demands.
[0101] A personalized federated learning method based on margin constraint and trust aggregation provided by the present invention has a specific implementation process as Figure 1 shown as follows:
[0102] I. Client local training stage;
[0103] On the client side, the present invention collaboratively improves the expression ability in heterogeneous environments and the defense ability against malicious clients through three key functional modules. As Figure 2 shown, its specific implementation process is as follows:
[0104] 1. Feature distillation alignment;
[0105] In a federated learning environment with strong heterogeneity and potential malicious attacks, the data distributions of different clients vary significantly. If only relying on limited local data for training, it is often difficult for the model to fully inherit the global prior features. Therefore, feature distillation alignment aligns the global model sent by the server with the client's local model through knowledge distillation at the feature level, enabling the client to retain the learning of common features during the update process.
[0106] In traditional knowledge distillation, the distillation loss is usually imposed by comparing the prediction probabilities or logits of the teacher model and the student model for the same input sample. However, the distribution heterogeneity in federated learning may lead to distribution drift or class bias when directly comparing the prediction distributions. If the client data classes do not exactly match those during the teacher model training, the distillation effect based on the output layer will be limited. In the present invention, the distillation object is changed to the feature representation layer. By comparing the intermediate layer features of the teacher model and the student model on the same input sample (x, y), the advantages of the global model in general feature abstraction can be better captured while retaining the differentiated features of the local data. Assume that the feature extraction module of the teacher model is F T (·), and the feature extraction module of the student model is F S (·), and the local dataset is D i ={(x n , y n )}. For any input sample (x, y), the teacher feature and the student feature can be respectively expressed as:
[0107] f T (x) = F T (x),
[0108] f S (x) = F S (x),
[0109] Subsequently, to quantify the difference between the two in the feature space, an alignment loss function L Align is defined as follows:
[0110] L Align (x) = γ · ||f S (x) - f T (x)|| 2 ,
[0111] where γ represents the weight of the alignment loss in the total loss. The larger this loss term is, the more significant the gap between the features extracted by the local model and the global prior features. Backpropagation will force the student features to approach the teacher features by increasing the gradient.
[0112] The pseudocode of feature distillation alignment in this process is shown in Algorithm 1.
[0113]
[0114]
[0115] In the specific implementation process, for each mini-batch, the client first extracts the features of the middle layer or the penultimate layer using the teacher model and the student model respectively, then calculates the difference of their two-norms as the distillation loss, and performs a weighted sum of this loss and other subsequent losses (such as margin-constrained cosine normalization), and jointly backpropagates to update the parameters of the student model. Through this process, the client can inherit the global prior features and improve the generalization ability under small-scale or biased data.
[0116] The present invention establishes "feature layer" distillation between the local model and the server-side global model, and transfers the representation knowledge of the teacher model (global) at the general level to the student model (local), so as to offset the model instability caused by limited local data scale or distribution deviation. In each mini-batch, the client simultaneously extracts the middle layer features of the same input sample using the global model and the local model, calculates the feature difference between the two and uses it as part of the distillation loss for backpropagation. Different from the traditional distillation only at the "output layer", the feature layer alignment in the present invention can better adapt to non-independent and identically distributed data, improve the generalization performance, and is more robust for training when the local label distribution is severely unbalanced.
[0117] 2. Margin-constrained cosine normalization;
[0118] In the traditional softmax normalization or softmax normalization with additive margin methods, only a relatively unified margin is imposed in the cosine metric space during classification. However, this method cannot effectively distinguish between reliable class samples and potentially malicious or abnormal samples. The present invention proposes a scheme for distinguishing different constraint strengths between reliable classes and suspicious classes in the cosine space, that is, margin-constrained cosine normalization. Through the differential design of the present invention in the cosine metric and angular margin ideas, the performance and stability of federated learning in the face of malicious interference and extremely unbalanced distributions can be effectively improved without destroying the basic advantages of the cosine metric.
[0119] The present invention sets different margins for reliable samples and suspicious samples through the margin-constrained cosine normalization method, making the feature vectors of reliable classes more compact and the feature space of suspicious classes relatively loose, thereby enhancing the defense ability of the model against new attacks or abnormal distributions.
[0120] Margin-constrained cosine normalization sets two different boundary values under the cosine metric, corresponding to the reliable distribution and the suspicious distribution respectively. When a sample is judged to be a reliable distribution, its feature vector and the fixed direction vector The angle between them must meet more stringent boundary requirements. When a sample is judged to have a suspicious distribution, a larger angular space is reserved to reduce overfitting to unknown attacks. Specifically, let the fixed direction vector represent the weight vector of the credible direction, which can be regarded as the convergence center of the target distribution. At the same time, the present invention defines two boundaries m0 and m1 (m0, m1 ∈ [-1, 1] & m0 > m1), corresponding to the angular limits of the credible class features and the suspicious class features respectively. The similarity between the local student feature f S (x) and the server teacher feature f T (x) is defined by the following formula:
[0121]
[0122] where f S,k (x) and f T,k (x) respectively represent the outputs of the student and the teacher on the k-th dimensional feature, and d represents the dimension of the feature space. When the similarity score(x) is larger, it indicates that the features of the local model and the global model on this sample are more consistent, which also means that this sample is more likely to be judged as the target distribution. Based on the above settings, the margin-constrained cosine normalization loss function on a single sample x i can be defined by the following formula:
[0123] L MCCS (x i ) = βlog(1 + e [σ(μ-score(x))] ),
[0124] where β is used to adjust the weight of this loss term in the total loss, and σ is used to control the steepness of the gradient; μ represents the adjustable boundary parameter. If the sample x i is judged to have a credible distribution, μ = m0; if the sample x i is judged to have a suspicious distribution, μ = m1. For a mini-batch, the above formula can be averaged over all samples:
[0125]
[0126] where N represents the number of clients.
[0127] From the perspective of the theoretical decision boundary, L = log(1 + e γ) It shows exponential growth when γ > 0, and gradually converges to 0 starting from ln2 when γ ≤ 0; when the similarity between the feature and the direction vector is not sufficient to meet the boundary requirements, the loss will increase significantly to enhance the discrimination ability. In the present invention, if the sample is correctly assigned to the credible class (γ ≤ 0), the loss will naturally tend to be small, thus encouraging a more robust and accurate classification boundary after multiple rounds of iteration. For margin-constrained cosine normalization, if the similarity score(x) is higher than the boundary parameter μ, then σ(μ - score(x)) is negative, making log(1 + e [σ(μ-score(x))] ) approach a smaller constant, indicating that the model has met or exceeded the boundary requirements for this sample; conversely, when the similarity score(x) is not sufficient to reach the boundary parameter μ, the loss exponential term σ(μ - score(x)) is activated, significantly enlarging the gradient, urging the student feature and the teacher feature to be further tightened on the credible class, or reserving a more flexible distribution space for the suspicious class.
[0128] Margin-constrained cosine normalization not only adjusts the angular margin but also allows for a more flexible strategy to be applied to suspicious samples, making it more suitable for the requirements of federated learning in high-heterogeneous and high-risk environments. Finally, the theoretical decision boundary of margin-constrained cosine normalization is plotted, as Figure 3 shown, and it can be intuitively seen that the latter has flexibility in the regulation strength between the credible distribution and the suspicious distribution. Margin-constrained cosine normalization can differentially treat credible samples and suspicious samples in the cosine metric space while taking into account the need for local teacher-student feature alignment, thus showing better robustness and discrimination ability in the non-i.i.d. scenario of federated learning.
[0129] The pseudocode of margin-constrained cosine normalization in this process is shown in Algorithm 2.
[0130]
[0131]
[0132] In the specific implementation process, the client can, after feature distillation alignment, perform a credible / suspicious classification judgment on the samples in the current mini-batch (such as using the similarity between the local student feature and the server teacher feature, historical loss, etc.), and then calculate the cosine losses with different boundaries and merge them into the total loss for backpropagation. This process significantly enhances the robustness of the model to adverse distributions such as possible poisoning samples and label noise.
[0133] The present invention adopts differentiated margins for the characterization distribution of credible samples and suspicious samples in cosine space, so that the credible class feature vector is more compact with the target class center, while the suspicious class features are kept at a larger interval, so as to reduce overfitting of known anomalies and retain the defense margin against unknown attacks. On the basis of soft maximum normalization, two different margin parameters are set and applied to credible samples and suspicious samples respectively to form differentiated classification boundaries; during training, samples are judged as credible / suspicious according to indicators such as global feature alignment or historical loss. The discriminability and robustness of the model in highly heterogeneous environments are improved to avoid recognition errors or over-tightening of special distributions or suspicious samples when only a single margin is used.
[0134] 3. Dispersion regularization;
[0135] In a highly heterogeneous federated learning environment where suspicious samples may exist, simply imposing stronger aggregation constraints on trusted samples is not enough to deal with unknown attacks or rare distributions. To this end, the present invention additionally introduces dispersion regularization, which deliberately retains a higher dispersion for suspicious samples from a feature perspective to avoid overfitting a single attack mode. In actual training, some samples may have obviously abnormal loss values, gradient features, or prediction deviations, which often indicate potential malicious attacks. If the same compactness constraint as the target class is imposed on these samples, the model is prone to over-learning known anomalies, making it difficult to effectively respond when unknown attacks occur. Based on this, dispersion regularization will identify suspicious samples in each mini-batch and apply regularization terms to their features to widen the intervals. Assume that the selected set of suspicious samples is Then the center vector of the set is expressed as follows:
[0136]
[0137] in, represents the feature vector of the mth suspicious sample, and M represents the number of suspicious samples. Next, the average dispersion, i.e., the average dispersion of this batch of suspicious samples in the feature space, is calculated according to the following formula:
[0138]
[0139] When the average dispersion δ is too small, it means that these suspicious samples are very similar to each other, and the model may focus on this feature during training, reducing the versatility for new or different attack modes.
[0140] In order to maintain the dispersion of these suspicious samples during the training process, the present invention defines dispersion regularization as shown in the following formula:
[0141]
[0142] Where ∈ represents a very small constant to prevent division by zero. When the average dispersion δ→0, L Spread →+∞, the gradient penalty will force the suspicious class features to separate from each other; when the average dispersion δ is large enough, it means that the suspicious samples have been fully dispersed in the feature space, and the contribution of the regularization term to the loss becomes smaller.
[0143] In margin-constrained cosine normalization, the trusted samples and suspicious samples use different cosine boundaries, so that the trusted samples converge toward the target direction, while retaining greater flexibility for suspicious samples. Dispersion regularization further constrains the dispersion of suspicious samples identified in local training, reducing the phenomenon of abnormal features being too close to each other from a feature perspective, such as Figure 4 The two complement each other, allowing the model to maintain the ability to adapt to unknown distributions or new attacks while taking into account the accuracy of the main task.
[0144] The pseudo code of dispersion regularization in this process is shown in Algorithm 3.
[0145]
[0146]
[0147] In the specific implementation process, dispersion regularization is calculated for suspicious samples in each small batch. Only when the features of these samples are overly concentrated with each other, the regularization term is activated and produces a significant gradient, which encourages the model to actively spread their positions in the feature space and enhance its adaptability when facing unknown abnormal distributions.
[0148] If suspicious samples are over-clustered in the feature space, the model will easily "remember" known attack patterns and lose its ability to defend against other potential attack methods. The present invention uses additional regularization terms to ensure that suspicious samples maintain sufficient dispersion and reduce overfitting of abnormal patterns. First, the center vector and the average distance between samples of the suspicious sample set are calculated. If the dispersion is too low, a penalty term is added to the total loss to force these suspicious class features to move away from each other. Unlike simply shielding suspicious samples or eliminating them outright, the present invention allows the model to have a wider adaptability to unknown attacks while retaining suspicious sample information; significantly enhancing the general robustness of the model.
[0149] In summary, the present invention comprehensively utilizes key mechanisms such as feature distillation alignment, margin-constrained cosine normalization, and dispersion regularization, starting from the overall process from client-side local training to server-side aggregation, to effectively cope with the federated learning environment with strong heterogeneity and potential malicious attacks. By applying differentiated cosine metric constraints to trusted and suspicious samples in the local training phase and inheriting global prior knowledge at the feature level, the present invention can significantly reduce the risk of overfitting to known attacks or noise patterns, while reserving sufficient defense and adaptation space for unknown distributions.
[0150] II. Server-side Aggregation Phase;
[0151] In the present invention, the client first conducts local training through three key functional modules: feature distillation alignment, margin-constrained cosine normalization, and dispersion regularization, aiming to obtain a more robust and well-discriminative personalized model in a strongly heterogeneous and potentially malicious attack environment. However, if the server still adopts the traditional uniform aggregation strategy, high-quality updates and malicious updates will be regarded as equally important, weakening the positive guiding role of high-quality clients in the convergence of the global model. Therefore, the present invention proposes a trust index aggregation strategy, which dynamically adjusts the weights of each client during aggregation, enabling clients that contribute more to the global objective to obtain higher weights, thus taking into account both robustness and fairness.
[0152] To quantify the trust index of the client in this round of training, the preliminary weight w' is defined i as follows:
[0153]
[0154] where λ d is the balance coefficient, used to trade off between tightening the discriminative ability and the dispersion ability. If the client performs well in distinguishing between trustworthy and suspicious samples, its loss value is low, causing the preliminary weight w' i to be amplified; otherwise, it decays exponentially. Subsequently, normalization is performed to obtain the aggregation weight α i of the client in the r-th round of federated training, as follows:
[0155]
[0156] where represents the set of clients participating in training in the r-th round, represents the sum of the preliminary weights of all clients.
[0157] Finally, the server aggregates all client updates according to the weight α i into a new global model Θ (r+1) as follows:
[0158]
[0159] where represents the updated model parameters of client i after the local training in the r-th round, and Θ (r+1) represents the global model to be distributed in the (r + 1)-th round.
[0160] Through trust index aggregation, high-quality clients will obtain higher weights on the server side and have a greater influence on the update of the global model. Malicious clients, on the other hand, will have their weight indices decay due to higher losses, thus effectively suppressing their destructive impact on the global model.
[0161] The pseudocode of trust index aggregation in this process is shown in Algorithm 4.
[0162]
[0163] In an actual system, the server side can perform statistics based on the losses or other performance metrics uploaded by the clients and adjust the weights through a simple exponential function. This mechanism not only retains an implementation interface similar to classical federated optimization methods but also can adaptively amplify the contributions of high-quality clients and suppress suspicious updates in extremely non-i.i.d. or multi-attack type scenarios, taking into account both global robustness and performance fairness.
[0164] The present invention performs exponential amplification / decay on the updates uploaded by each client on the server side, so that high-quality or highly credible clients account for a greater weight in the global model update, and the influence of potentially malicious or poorly performing clients is relatively weakened. The loss metric of each client at the end of local training is incorporated into the exponential weight function, and then normalized to obtain the aggregated weights. Finally, the client parameters are weighted and summed to update the global model. Compared with traditional average aggregation or static weighting, the present invention can dynamically identify and suppress malicious updates, and can also adaptively amplify small-scale but excellent clients, taking into account overall robustness and fairness among multiple clients.
[0165] A personalized federated learning method based on margin constraint and trust aggregation provided by the present invention, and its complete pseudocode implementation can be seen in Algorithm 5 to more detailedly present the interaction and execution process of all the above key functional modules.
[0166]
[0167] In summary, the present invention comprehensively utilizes key mechanisms such as feature distillation alignment, margin constraint cosine normalization, and dispersion regularization, starting from the overall process from client local training to server-side aggregation, and effectively dealing with the federated learning environment with strong heterogeneity and potential malicious attacks. On this basis, combined with the trust index aggregation strategy on the server side, it can not only greatly strengthen the positive contribution of high-quality clients to the global model but also, to a certain extent, suppress the interference of malicious or poorly performing clients, thus taking into account the overall robustness of the system and the fairness among multiple clients. After multiple rounds of iteration, the overall technical solution proposed by the present invention can converge to a global model with excellent generalization ability, strong adversarial performance, and good distribution adaptability in various scenarios.
[0168] The above four major innovative points of the present invention (feature distillation alignment, margin-constrained cosine normalization, dispersion regularization, and confidence index aggregation) run through the entire process of the present invention from local client training to server-side aggregation. Through joint optimization at the feature learning and aggregation scheduling levels, it effectively solves the problems such as difficult model convergence, insufficient discriminative power, and poor fairness among multiple clients under non-independent and identically distributed and potential malicious interference.
[0169] In addition, to achieve the core goal of the present invention, "simultaneously improving the generalization, robustness, and fairness of federated learning in high-heterogeneous and potentially malicious scenarios", in addition to the foregoing technical solutions, there are also the following alternative or expandable implementation manners. These alternative solutions usually differ in specific implementation details or module combinations, but they can still meet the purpose of the present invention in terms of the overall framework and functional requirements:
[0170] (1) Adjustment of the distillation layer: In the feature distillation alignment part, the present invention adopts the feature alignment method of the feature layer (the last layer) to balance the global prior and local differences under non-independent and identically distributed conditions. It is also possible to adjust the distillation target to more complex strategies such as multi-layer features to further strengthen the client's absorption of global knowledge. Even so, its core principle is still "retaining the client's personalized features while inheriting the general representation", so it can achieve the same technical effects.
[0171] (2) Multiple client training-server aggregation interaction methods: The currently described process is mainly based on the synchronous communication mode of fixed rounds. After the clients complete local training, they uniformly submit the updates to the server. If a semi-asynchronous or fully asynchronous mode is adopted, as long as the local update mechanisms for "feature distillation alignment", "margin-constrained cosine normalization", and "dispersion regularization" are retained, and the corresponding "confidence regulation" principle is adopted on the server side, the same robust and fair effects as those of the present invention can be obtained in an asynchronous environment. Although the communication timings are different, the overall mechanism still meets the purpose of the invention.
[0172] The above are only the preferred implementation manners of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A personalized federated learning method based on margin constraints and trust aggregation, characterized by: The following steps are involved: S1: Feature distillation alignment; Through feature-level knowledge distillation, the global model sent by the server is aligned with the client's local model, allowing the client to retain the learning of common features during the update process; S2: margin constrained cosine normalization; Two different boundary values are set under the cosine metric, corresponding to the credible distribution and the suspicious distribution respectively; when the sample is judged to be a credible distribution, The angle between its feature vector and the fixed direction vector must meet strict boundary requirements; when a sample is judged to be a suspicious distribution, a larger angle space is reserved to reduce overfitting of unknown attacks; S3: dispersion regularization; In each mini-batch, suspicious samples are identified and regularization terms that widen the interval are applied to their features; S4: Trust index aggregation; By dynamically adjusting the weight of each client during aggregation, clients that contribute more to the global goal will receive higher weights.
2. The personalized federated learning method based on margin constraint and trust aggregation according to claim 1 is characterized in that: In step S1, it is assumed that the feature extraction module of the teacher model is F T (·), the feature extraction module of the student model is F S (·); For any input sample x, y), the client uses the teacher model and the student model to extract the features of the middle layer or the second-to-last layer respectively. The obtained teacher features and student features are expressed as: f T (x=F T x), f S x=F S x, In order to quantify the difference between teacher features and student features in the feature space, the alignment loss function L is defined Align : Among them, γ represents the weight of the alignment loss in the total loss. The larger the loss, the more significant the gap between the features extracted by the local model and the global prior features. Back propagation will force the student features to move closer to the teacher features by increasing the gradient.
3. The personalized federated learning method based on margin constraint and trust aggregation according to claim 1 is characterized in that: In step S2, let the fixed direction vector The weight vector representing the credible direction is regarded as the convergence center of the target distribution; two boundaries m0 and m1 (m0, m1∈[-1,1]&m0>m1) are defined to correspond to the angle limits of the credible class features and the suspicious class features respectively; the local student feature f S (x) and server teacher feature f T (The similarity of x is defined as: Among them, f S,k (x) and f T,k x) represents the output of the student and teacher on the k-th dimension feature, d represents the dimension of the feature space; the larger the similarity scorex), the more consistent the local model and the global model are on the features of this sample, and the sample is more likely to be judged as the target distribution.
4. The personalized federated learning method based on margin constraints and trust aggregation according to claim 3 is characterized in that: According to the characteristics of local students S (x) and server teacher feature f T The similarity definition of (x) is to normalize the margin constraint cosine on a single sample x i The loss function on is defined as: THE MCCS x i )=βlog1+e [σ(μ-score(x))] , Among them, β is used to adjust the weight of the loss term in the total loss, σ is used to adjust the steepness of the gradient; μ represents an adjustable boundary parameter. If the sample x i It is judged to be a credible distribution, μ = m0; if the sample x i It is judged to be a suspicious distribution, μ = m1; for a small batch, average all samples: Where N represents the number of clients.
5. The personalized federated learning method based on margin constraints and trust aggregation according to claim 1 is characterized in that: In step S3, it is assumed that the selected suspicious sample set is Then the center vector of the set is expressed as: in, represents the feature vector of the mth suspicious sample, and M represents the number of suspicious samples.
6. The personalized federated learning method based on margin constraints and trust aggregation according to claim 5 is characterized in that: According to the central vector representation of the suspicious sample set, the average dispersion of the suspicious samples in the feature space is calculated: When the average dispersion δ is too small, it means that these suspicious samples are very similar to each other, and the model may fit the feature set during training.
7. The personalized federated learning method based on margin constraints and trust aggregation according to claim 6 is characterized in that: In order to maintain the dispersion of these suspicious samples during training, the dispersion regularization is defined as: Where, ∈ represents a very small constant to prevent division by zero; when the average dispersion δ→0, L Spread →+∞, the gradient penalty will force the suspicious class features to separate from each other; when the average dispersion δ is large enough, it means that the suspicious samples have been fully dispersed in the feature space, and the contribution of the regularization term to the loss becomes smaller.
8. The personalized federated learning method based on margin constraint and trust aggregation according to claim 1, characterized in that: In step S4, in order to quantify the trust index of the client in this round of training, a preliminary weight w' is defined i for: Among them, λ d is the balance coefficient, which is used to balance the ability to tighten the discrimination and the ability to disperse. Spread represents the dispersion regularization loss; L MCCS represents the margin constrained cosine normalized loss; if the client performs well in distinguishing between credible and suspicious samples, its loss value is low, making the initial weight w' i is amplified; otherwise it decays exponentially.
9. The personalized federated learning method based on margin constraints and trust aggregation according to claim 8, characterized in that: According to the initial weight w' defined i Normalize to obtain the aggregate weight α of the client in the rth round of federated training i : in, represents the set of clients participating in the training in round r, Represents the sum of the initial weights of all clients.
10. The personalized federated learning method based on margin constraints and trust aggregation according to claim 9, characterized in that: The server updates all clients according to the weight α i Aggregate into a new global model Θ (r+1) : in, represents the model parameters updated by client i after the rth round of local training, Θ (r+1) Indicates the global model to be issued in the r+1th round.
Citation Information
Cited By
Federal learning optimization method based on industrial small sample defect detection
CN120706512A