Prototype-guided Federated Consistency Representation Learning System and Method

By introducing prototype-guided in-source characterization calibration and cross-origin consistent characterization learning modules in federated learning, the feature space consistency problem of client models without sharing private data is solved, and high accuracy and universal image classification effect is achieved.

CN119229223BActive Publication Date: 2025-05-30SHANDONG RES INST OF IND TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411764345.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-30
Estimated Expiration
2044-12-04

Smart Images

  • Figure CN119229223B_ABST
    Figure CN119229223B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of federated learning, and specifically relates to a federated consistent representation learning system and method based on prototype guidance. The system includes two main modules: an in-source representation calibration module and a cross-source consistent representation learning module. It first uses the in-source representation correction module to improve local training and correct the feature distribution on unbalanced data. This helps to alleviate the significant differences in the feature space caused by dataset bias. At the same time, it provides prototype information to the server, including cluster prototypes, cluster variances, and attention scores. Then, the cross-source consistent representation learning module uses the prototype information obtained from all clients to learn generalized projections and classifiers. The algorithm first generates extended features using statistical knowledge to refine the feature space and improve diversity. Subsequently, features from different sources are mapped to a unified space for comparison and classification, and the interference of outliers is eliminated according to the attention scores.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and particularly relates to a federated consistency representation learning system and method based on prototype guidance. Background Art

[0002] Federated learning supports collaborative modeling using data from different sources. It shares model parameters rather than raw data between data sources to ensure privacy and security. This significantly improves the effective utilization of isolated data, enabling them to contribute to collaborative decision-making and learn a general model. However, existing research emphasizes that the heterogeneity of data distributions among clients may lead to a reduction in the effectiveness of collaborative modeling. This is mainly because it becomes challenging to learn a consistent feature space when dealing with imbalanced data within clients and inconsistent distributions between clients, making it very difficult to integrate multiple learners with inconsistent objectives into a significant model.

[0003] The inventors found that the prior art has the following technical defects: existing federated learning methods cannot guide client models to learn a consistent feature space using private data without sharing user privacy data, and they ignore the imbalance of training data, which may hinder effective representation learning and calibration for categories with a limited sample size, resulting in low image classification accuracy of the aggregated model and a lack of universality. Summary of the Invention

[0004] In view of the above problems existing in the prior art, the present invention provides a federated consistency representation learning system and method based on prototype guidance, which significantly improves the image classification effect in the federated learning scenario.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows: In a first aspect, the present invention provides a federated consistency representation learning system based on prototype guidance, including: an intra-source representation calibration module, which is used to model the representation distribution of clients, improve local training, and correct the feature distribution on imbalanced data; at the same time, it provides prototype information to the server, including cluster prototypes, cluster variances, and attention scores;

[0006] Constraining the representation learning of the local model using the class-aware text representation output by the pre-trained CLIP model. The client stores a private data set and a private model. After training, the trained private model is used to extract features from the private data set, and the extracted features are clustered. The mean value of the features within the cluster is regarded as a prototype, and the variance of the features within the cluster in each dimension and the importance score of the cluster prototype are calculated;

[0007] Cross - source consistent representation learning module for learning generalization projection and classifier; the client sends the private model and prototype information to the server; the server receives the private model and prototype information uploaded by the client, generates enhanced features using statistical knowledge, refines the feature space, and improves diversity; subsequently, the cross - source feature alignment module based on relationship consistency maps features from different sources to a unified space for comparison and classification, sums the parameters of the private models, and then takes the average to obtain the global model. The server calibrates the global projection head and global classifier in the global model using the prototypes of all parties, sends the calibrated global model to the client, and eliminates the interference of outliers according to the attention scores.

[0008] In a typical implementation, the in - source representation calibration module includes two processes: knowledge - guided representation calibration and clustering - driven class pattern modeling.

[0009] Furthermore, the knowledge - guided representation calibration is specifically as follows: using the fixed category - aware text features learned from the pre - trained CLIP as the optimization target for local feature learning to standardize the client's feature learning; it performs supervised prototype contrast learning to maximize the consistency between the image features and text prototypes in the latent space. The objective loss is defined as:

[0010] ,

[0011] where, represents the image features of class in the client , , is the text encoder, represents the number of classes, represents that when the model's prediction of the image is , it takes 1, otherwise it takes 0, is the number of training data in the client , represents the temperature parameter; meanwhile, the empirical loss is used to further ensure the classification ability of the model, that is,

[0012] ,

[0013] where, represents the -th element in the model output vector. The dimension of the model output vector is consistent with the number of classes. The -th element represents the prediction probability of the -th class, which is a numerical value; similarly, represents the elements is the label of the image

[0014] In a typical implementation, the cluster-driven class pattern modeling is specifically as follows: using the fixed model learned from the intra-source representation calibration module extract features from all training data and use the k-means clustering method to mine different patterns in the latent space, that is

[0015] ,

[0016] where represents the data of category in the client , is a hyperparameter that represents the number of clusters represents the features of the cluster of category .

[0017] In addition, to obtain a more accurate distribution in the feature space, this module calculates the mean and variance of the features within the cluster , that is

[0018] ,

[0019] ,

[0020] where represents the size of the cluster

[0021] In addition, considering that the imbalance of data distribution leads to limited capabilities, in order to learn the discriminative representation of the minority sample classes, the intra-source representation calibration module further evaluates the importance of all clusters to reduce the interference of abnormal features on model calibration. It includes three factors, including the cluster size , the cluster compactness and the minimum distance to the cluster centers of other classes . For the cluster , , , , where is the cluster center of the cluster different from the cluster , represents the data features belonging to the cluster . Essentially, the larger and more compact a cluster is, and the farther it is from the centers of other clusters, the more important it is. Therefore, the importance score of the cluster can be expressed as . The client will use the triple and the local model Uploaded to the server, where is the number of clusters in the client . It should be noted that the cluster center is also known as the local prototype.

[0022] In a typical implementation, the cross - source consistent representation learning module obtains all the local models and the set of local prototypes uploaded from the client, and aligns the prototype features from heterogeneous spaces, mainly including two processes: class - aware region calibration and cross - source feature alignment.

[0023] Furthermore, the class - aware region calibration is specifically as follows: Using knowledge transfer technology to calibrate the class - aware region, it uses a Gaussian model to generate extended features based on variance, and fuses the clustering variance with the highest score for the corresponding client and the variances except to transfer important knowledge to other class features in the client , that is,

[0024] ,

[0025] where is the local prototype, represents the enhanced feature, represents the th enhanced feature, is the number of enhanced features , represents the fused variance, represents the clustering variance with the highest score for the corresponding client, represents generating enhanced features that satisfy the Gaussian distribution with a mean of 0 and a variance of .

[0026] It should be noted that compared with the point - to - point method in the existing method, the generated set of enhanced features forms a region, which helps the cross - source feature alignment module to achieve region - to - region alignment.

[0027] Furthermore, the cross - source feature alignment is specifically as follows: For the original global model , the global feature extractor does not need to be retrained, while the projection head and the classifier need to be calibrated, that is . Map the locally learned prototypes and the enhanced features from the class - aware region calibration module to a new space designed for cross - source collaborative classification. At the same time, it uses a two - level regularization method to refine the representation learning, including local consistency matching and complementary consistency matching, which can more effectively emphasize the commonality of intra - class features and the difference of inter - class features, and eliminate client - specific information.

[0028] For the local consistency matching level, it promotes the learning process by imposing constraints on the consistency of the relationships between local representations, thereby guiding the model to acquire features that remain invariant across different clients. Taking these 3 classes as an example, the local consistency matching is expressed in the following way:

[0029] ,

[0030] where, is the calibrated projection head mapped features. If is the local prototype of class in client , then ; if is the enhanced feature, then , , represents the dot product; represents the global prototype of class , which is the average of the local class prototypes of all clients.

[0031] ,

[0032] where, represents the distance-based consistency matching loss, is the Euclidean distance.

[0033] Overall, the local matching loss is defined as:

[0034] .

[0035] For the complementary consistency matching level, it utilizes the complementarity of features from different sources to promote the model to learn consistent features across clients, enabling the model to transcend the limitations of a single perspective and achieve a more comprehensive learning level. This can be defined as:

[0036] .

[0037] In addition, to enhance the robustness of model calibration and maintain clear decision boundaries, this module utilizes the importance scores of all features in a weighted manner

[0038] to shift the focus of the model away from features of lower quality, thereby designing a weighted supervised classification loss, defined as follows:

[0039] where, is the attention score learned from the source internal representation calibration module, is passed through the classifier to learn the prediction.

[0040] Furthermore, the cross-source consistent representation learning module sends the calibrated global model to all clients.

[0041] In a second aspect, the present invention provides a prototype-guided federated consistent representation learning method, including: training a private model using private data in a client, and using the class-aware text representation output by a pre-trained CLIP model to constrain the representation learning of the local model. The private dataset and the private model are stored in the client. After training, the trained private model is used to extract features from the private dataset, the extracted features are clustered, the mean of the features within the cluster is regarded as the prototype, and the variance of the features within the cluster in each dimension and the importance score of the cluster prototype are calculated; the client sends the private model and the prototype information to the server; the server receives the private model and the prototype information uploaded by the client, sums the parameters of the private model, and then takes the average to obtain the global model; the server uses the prototypes of all parties to calibrate the global projection head and the global classifier in the global model, and sends the calibrated global model to the client.

[0042] In a typical embodiment, there are at least two clients and one server.

[0043] In a typical embodiment, the private models in all clients have the same structure.

[0044] In a typical embodiment, the training objective of the client is to calibrate the local distribution of the client to alleviate the significant differences in the feature space between clients caused by data distribution imbalance. The overall optimization objective of the client training strategy is:

[0045] ,

[0046] where is the weighting parameter;

[0047] The server further reduces the differences in feature distributions in different spaces, and the server optimizes the following objective function:

[0048] ,

[0049] where is the weight parameter.

[0050] In a typical embodiment, the relevant information is the cluster prototype, the within-cluster variance, and the importance score of the cluster prototype.

[0051] In a typical embodiment, in the in-source representation calibration module, the weighting parameter is 0.1 to 5, the temperature parameter is 0.5, and the number of clusters is 1 to 3; in the cross-source consistent representation learning module, the weight parameter is 0.1 to 5, and the number of extended features is 1 to 8.

[0052] Furthermore, in the in-source representation calibration module, the weighting parameter , the temperature parameter , and the number of clusters .

[0053] In the cross-source consistent representation learning module, the weight parameter , and the number of extended features .

[0054] The present invention has achieved the following technical effects:

[0055] The present invention proposes a new prototype-guided federated consistency representation learning method called FedCRL, which includes two main modules: an in-source representation calibration module and a cross-source consistent representation learning module. To promote the alignment of features across different spaces, the in-source representation calibration module uses pre-trained text representations to guide the local learning of each client to reduce the large differences in different source feature spaces; subsequently, clustering is used to model the class patterns and provide prototype information to the server to assist in model calibration. Subsequently, to enhance the robustness of calibration, the present invention develops a class-aware region calibration method in the cross-source consistent representation learning module, which reconstructs the feature space of each source by using knowledge transfer technology, alleviates the adverse effects brought by data distribution imbalance, and increases the sample diversity of each category. At the same time, the present invention uses two-layer constraints to achieve effective alignment of cross-source features.

[0056] It should be noted that the cross-source consistent representation learning module is a highly adaptable tool that can be easily integrated into various algorithms. Experiments were conducted on four datasets, including performance comparison, ablation study of key components, in-depth analysis of the effectiveness of cross-source consistent representation learning, and case studies. The experimental results show that the present invention can calibrate heterogeneous representations from different data sources into a unified space, and the performance is better than existing methods. The present invention calibrates heterogeneous features between clients into a unified space without sharing user privacy data, and the calibrated model has high image classification accuracy and universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] FIG. 1 is a schematic diagram of a prototype-guided federated consistency representation learning system according to Embodiment 1 of the present invention;

[0058] FIG. 2 is a schematic diagram showing poor client representation learning in federated learning leading to poor performance of the federated learning model according to Embodiment 1 of the present invention;

[0059] Figure 3 is an illustration of the framework of a prototype-guided federated consistency representation learning method in Embodiment 1 of the present invention;

[0060] Figure 4 is a schematic diagram of the robustness analysis of FedCRL with different hyperparameters in Embodiment 1 of the present invention;

[0061] Figure 5 is a schematic diagram of the relationship between the number of generated enhanced features and the classification results in Embodiment 1 of the present invention;

[0062] Figure 6 is a schematic diagram of the effectiveness of the in-source representation calibration module in Embodiment 1 of the present invention;

[0063] Figure 7 is an error analysis diagram of the cross-source consistent representation learning method proposed in Embodiment 1 of the present invention;

[0064] Figure 8 is a schematic diagram of the feature distribution of FedCSPC in Embodiment 1 of the present invention;

[0065] Figure 9 is a schematic diagram of the feature distribution after cross-source feature alignment in FedCRL in Embodiment 1 of the present invention. Detailed implementation manners

[0066] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0067] It should be noted that the terms used herein are only for describing the specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0068] The present invention will be further described below in conjunction with the embodiments.

[0069] Embodiment 1

[0070] As shown in Figure 1, the present invention provides a prototype-guided federated consistent representation learning system to enhance the generalization ability of the global model, including: an intra-source representation calibration module for modeling the representation distribution of the client and providing prototype information to the server; a cross-source consistent representation learning module for learning a generalization projection and a classifier, which first generates enhanced features using statistical knowledge, refines the feature space, and improves diversity; subsequently, maps features from different sources to a unified space for comparison and classification, and eliminates the interference of outliers according to the attention score, enabling the global model to generalize to all clients.

[0071] 1. Intra-source representation calibration module

[0072] The intra-source representation calibration module aims to use pre-trained knowledge to help reduce the difference in the feature space between clients and provide prototype information about the representation to the server, which helps model calibration. It has two main processes: knowledge-guided representation calibration and modeling prototype representation for data through clustering.

[0073] (1) Knowledge-guided representation calibration

[0074] Intuitively, the change of the target is not conducive to optimizing the representation of categories with fewer samples. Therefore, the Knowledge-guided Representation Calibration (KGRA) module uses a set of fixed prototypes to standardize the feature learning of the client. Inspired by the pre-trained vision-language model, the KGRA module uses the fixed category-aware text features learned from the pre-trained CLIP model as the optimization target for local feature learning, ensuring the consistency of all training rounds, promoting cooperation to make up for its own deficiencies, and avoiding error accumulation. Specifically, it performs supervised prototype contrast learning to maximize the consistency between the image features and the text prototype in the latent space, and the objective loss is defined as:

[0075] ,

[0076] where, represents the image feature of class in client , , is the text encoder, represents the number of classes, represents that when the model's prediction of the image is , it takes 1, otherwise it takes 0, is the number of training data in client , represents the temperature parameter; at the same time, the empirical loss is used to further ensure the classification ability of the model, that is,

[0077] ,

[0078] wherein, represents the -th element in the model output vector. The dimension of the model output vector is consistent with the number of categories. The -th element represents the prediction probability for the -th category, which is a numerical value. Similarly, represents the -th element in the model output vector, which is the label of the image.

[0079] (2) Cluster-driven Class Pattern Modeling

[0080] It can be understood that despite efforts to enhance feature learning, intra-class differences and inter-class overlaps still occur in the feature space, which generate noise information that interferes with prototype modeling. Therefore, the Cluster-driven Class Pattern Modeling (CPM) module aims to explore the consistency and diversity represented in each class and evaluate the importance of each prototype. Specifically, the CPM module uses the fixed model learned from the KGRA module to extract features from all training data and uses the k-means clustering method to mine different patterns in the latent space, i.e.,

[0081] ,

[0082] wherein, represents the data of category in client , is a hyperparameter that represents the number of clusters, represents the features of the cluster of category . In addition, to obtain a more accurate distribution in the feature space, this module calculates the mean and variance of the features within the cluster , i.e.,

[0083] ,

[0084] ,

[0085] where represents the size of the cluster.

[0086] In addition, considering that the imbalance of data distribution leads to limited capabilities, in order to learn discriminative representations of minority sample classes, the Intra-source Representation Calibration module further evaluates the importance of all clusters to reduce the interference of abnormal features on model calibration. It includes three factors, including cluster size , clustering compactness and the minimum distance to the cluster centers of other categories . For a cluster , , , , where is the cluster center of a different category from cluster , represents the data characteristics belonging to cluster . Essentially, the larger and more compact a cluster is, and the farther it is from the centers of other clusters, the more important it is. Therefore, the importance score of cluster can be expressed as . The client uploads the triple and the local model to the server, where is the number of clusters in the client . It should be noted that the cluster center is also called the local prototype.

[0087] 2. Cross - source consistent representation learning module

[0088] Generally, simply using a local prototype set containing only a few samples and some noisy samples to train a model may reduce its generalization ability. Therefore, the design of the cross - source consistent representation learning module is to improve the learning process, focusing on enriching the latent space and reducing the attention to noise. Specifically, it has two main processes: class - aware region calibration and cross - source feature alignment based on relational consistency.

[0089] (1) Class - aware region calibration

[0090] To increase the diversity of the latent space, the class - aware region calibration (CRC) module uses knowledge transfer techniques to calibrate the class - aware region, thereby further alleviating the adverse effects brought by data distribution imbalance. Specifically, the CRC module uses a Gaussian model to generate extended features based on variance, fusing the clustering variance with the highest score for the corresponding client and the variances except to transfer important knowledge to other category features in the client , that is,

[0091] ,

[0092] where is the local prototype, represents the enhanced feature, represents the th enhanced feature, is the enhanced feature The quantity of represents the fusion variance represents the clustering variance with the highest score for the corresponding client represents generating enhanced features that satisfy the Gaussian distribution with a mean of 0 and a variance of ,

[0093] It should be noted that, compared with the point-to-point method in the existing method, a set of generated enhanced features forms a region, which helps the cross-source feature alignment module achieve region-to-region alignment.

[0094] (2) Cross-source feature alignment based on relational consistency

[0095] Obviously, as the data distribution imbalance intensifies, it becomes challenging to completely eliminate functional heterogeneity in the clients. Therefore, the cross-source feature alignment (CSFA) module based on relational consistency is used to retrain the generalized projection and the classifier so that the global model can achieve unified feature learning for samples with the same label across data sources, that is, for the original global model , the global feature extractor does not need to be retrained, while the projection head and the classifier need to be calibrated, that is . Specifically, the CSFA module maps the locally learned prototypes and the enhanced features from the CRC module to a new space designed for cross-source collaborative classification. At the same time, it adopts a two-level regularization method to refine the representation learning, including local consistency matching and complementary consistency matching, which can more effectively emphasize the commonality of intra-class features and the difference of inter-class features, and eliminate client-specific information.

[0096] For the local consistency matching level, it promotes the learning process by imposing constraints on the consistency of the mutual relationships between local representations, thereby guiding the model to obtain features that remain unchanged across different clients. Taking these 3 classes as an example, the local consistency matching is expressed in the following way:

[0097] ,

[0098] where is the feature after being mapped by the calibrated projection head . If is the local prototype of class in client , then ; if is the enhanced feature, then , , represents the dot product; represents the global prototype of the class , which is the average of all client local class prototypes.

[0099] ,

[0100] Among them, represents the distance-based consistency matching loss, which is the Euclidean distance.

[0101] Generally speaking, the local matching loss is defined as:

[0102] .

[0103] For the complementary consistency matching level, it utilizes the complementarity of features from different sources to promote the model to learn consistent features across clients, enabling the model to transcend the limitations of a single perspective and achieve a more comprehensive learning level. This can be defined as:

[0104] .

[0105] In addition, to enhance the robustness of model calibration and maintain a clear decision boundary, this module utilizes the importance scores of all features in a weighted manner to shift the focus of the model away from features with lower quality, thus designing a weighted supervised classification loss, defined as follows:

[0106] ,

[0107] Among them, is the attention score learned from the in-source representation calibration module, is the prediction learned through the classifier .

[0108] The cross-source consistent representation learning module sends the calibrated global model to all clients.

[0109] Example 2

[0110] As shown in Figure 3, this embodiment provides a prototype-guided federated consistent representation learning method, including: training a private model using private data in two clients, and using the class-aware text representation output by a pre-trained CLIP model to constrain the representation learning of the local model. The private datasets and private models are stored in the clients, and the private models in all clients have the same structure. After training, use the trained private model to extract features from the private dataset, cluster the extracted features, and the mean value of the features within the cluster is regarded as the prototype, and calculate the variance of the features within the cluster in each dimension and the importance score of the cluster prototype. The client sends the private model and prototype information (cluster prototype, within-cluster variance, and importance score of the cluster prototype) to a server; the server receives the private model and prototype information uploaded by the client, sums the parameters of the private model, and then takes the average to obtain the global model; the server calibrates the global projection head and global classifier in the global model using the prototypes of all parties, and sends the calibrated global model to the client.

[0111] This method works collaboratively at the client and server levels to alleviate the challenges brought by cross-client class imbalance. It has the following training strategies.

[0112] In the client, the goal is to calibrate the local distribution of the client to alleviate the significant differences in the feature space between clients caused by the unbalanced data distribution. The optimization objective loss is defined as:

[0113] ,

[0114] where is the weighting parameter.

[0115] In the server, FedCRL aims to further reduce the distribution differences in the feature space among different clients and optimize the following objective function:

[0116] ,

[0117] where is the weight parameter.

[0118] In summary, the present invention includes three main contributions: a new cross-source consistency representation learning framework FedCRL in federated learning is proposed, which can alleviate the negative impact brought by data imbalance and improve the feature alignment between data sources. The proposed cross-source consistency representation learning module represents an orthogonal enhancement to client-based methods. Its plug-and-play feature enables it to effectively combine multiple strategies, enhancing the generalization performance of the corresponding model. Experimental results show that unbalanced training data will weaken the benefits brought by cross-silo feature alignment to federated learning. This finding ensures the effectiveness of the proposed in-silo representation calibration and cross-source consistency representation learning modules, enhancing the representation learning of clients and cross-source feature alignment on the server respectively.

[0119] To alleviate the negative impact of the differences between different feature spaces on model aggregation in federated learning, existing research can be roughly divided into two categories, including knowledge distillation-based methods and model calibration-based methods.

[0120] It should be understood that the knowledge distillation-based method is a promising strategy for alleviating data heterogeneity in federated learning, aiming to guide clients to learn consistent knowledge and construct similar feature spaces. Generally, in this research field, traditional methods need to use additional information as a regularization norm for local updates. In this context, regularization techniques play an important role. For example, Moon uses contrastive regularization to penalize the inconsistency between local and global feature spaces. FedProc, FedProto, and FPL construct prototypes for each class according to the sample representation to represent the center of the intra-class representation. Then, it guides the local training process by constraining the representations of all clients to converge to these prototypes. In addition, using a classifier to guide the calibration of the feature space is also an effective strategy. For example, FedETF uses a fixed simplex equiangular tight frame classifier to encourage all clients to learn unified and optimal feature representations. FedFA uses feature anchors to simultaneously optimize the feature space and calibrate the classifier, promoting a virtuous cycle between feature space and classifier updates. Although the results of these methods are positive, further exploration of unbalanced data is still needed because imbalance usually accumulates errors in training iterations.

[0121] It should be understood that in order to improve the performance of model aggregation in federated learning, methods based on model calibration have become the focus of numerous studies. Different from knowledge distillation, they focus on improvements on the server side, aiming to alleviate the bias problem caused by the weighted average of model parameters. Traditional methods along this line include global classifier calibration, projection head retraining, and global model fine-tuning. They all hope to obtain a general model to adapt to all data from different sources. For example, CCVR fuses the mean and variance of sample features obtained from clients and uses a Gaussian mixture model to generate virtual features for retraining the global classifier. Creff and CLIP2FL generate a series of features whose gradients are consistent with the actual data to fine-tune the classifier. FedFTG fine-tunes the entire global model by using a generator to explore the input space. In addition, the goal of FedCSPC is to map different features to a common space for alignment and classification. From the previous analysis, their performance is closely related to the quality of local feature information, as shown in Figure 2. Therefore, alleviating unbalanced factors to enhance representation learning is a promising method.

[0122] 1. Performance Comparison

[0123] To verify the effectiveness of the algorithm

[0124] In this invention, three datasets, including the CIFAR10, CIFAR100, and TinyImageNet datasets, are used in the experiment. Table 1 below gives the statistical details of the datasets and uses the Dirichlet distribution at that time to partition the datasets. when partitioning the datasets.

[0125] Table 1 Statistical Information of the Datasets Used

[0126]

[0127] This invention uses the Top-1 accuracy to evaluate the performance of the method, and the formula is as follows:

[0128] ,

[0129] where , are the number of correct predictions and the total number of samples, respectively.

[0130] For all methods, set the number of clients and , the sampling ratio , the number of local training epochs , the batch size , the communication rounds on the CIFAR10 and CIFAR100 datasets , communication rounds on the TinyImageNet dataset , and the learning rate of the SGD optimizer , the weight decay is set to .

[0131] In the in-source representation calibration module, the weighting parameter , the temperature parameter , the number of clusters ; in the cross-source consistent representation learning module, the weight parameter , the number of extended features .

[0132] The present invention is compared with nine state-of-the-art methods, including FedAvg, MOON, CCVR, FedNTD, FedProc, FedDecorr, FedETF, FedCSPC, and CLIP2FL. The network architectures used for all methods include an image encoder, a projection head, and a classifier. For all datasets, the present invention uses a 2-layer multi-layer perceptron (MLP) as the projection head, and the classifier is a 1-layer fully connected layer. For the CIFAR10 dataset, the present invention uses a convolutional neural network including two 5x5 convolutional layers followed by 2x2 max pooling, and two fully connected layers with ReLU functions as the image encoder. For other datasets, the present invention uses ResNet18 (excluding its last fully connected layer) as the image encoder. The following can be observed from Table 2 below:

[0133] FedCRL is a general framework that can incorporate various knowledge extraction-based methods, such as FedAvg and FedETF, and bring performance improvements to them, demonstrating its model-agnostic ability; model calibration-based methods are generally superior to knowledge distillation-based methods, which the present invention proves, because it endeavors to utilize information from multiple sources to train a general model; FedETF uses a unified simplex equiangular tight frame classifier and tends to produce better results than data-driven knowledge-based methods (FedProc, FedNTD, MOON). This may be because they avoid the problem of poor knowledge quality caused by data differences and the inherent limitations of the model itself. As the number of data sources increases, the performance tends to decline. This is due to the widening gap between data distributions. FedCRL maintains its performance advantage, fully demonstrating the effectiveness of its calibration mechanism.

[0134] Table 2 Performance comparison of FedCRL and existing methods on CIFAR10, CIFAR100, and TinyImageNet datasets

[0135] 。

[0136] 2. Ablation Experiments

[0137] Table 3 Study on the ablation effect of different components of FedCRL on CIFAR10 and CIFAR100 datasets

[0138] 。

[0139] As can be seen from Table 3 above, the Intra-Source Representation Calibration (ISRC) module plays a key role, contributing an average performance improvement of 1.2% to the baseline method, which verifies that providing unified guidance to different clients helps improve their collaborative results; the cooperation between the Class-Aware Region Calibration (CRC) module and the ISRC module improves the accuracy of the baseline method by approximately 3% in all cases; generally speaking, the Cross-Source Feature Alignment (CSFA) module can also produce good results without the support of CRC because it reduces outlier interference in a weighted manner compared with existing methods; combining the ISRC, CRC, and CSFA modules can produce the best performance. This is reasonable because the ISRC module provides reliable functional information, the CSFA module optimizes the functional space, and uses this information to reduce the distribution differences across different spaces.

[0140] 3. Robustness of FedCRL to Hyperparameters

[0141] As shown in Figure 4, FedCRL shows highly consistent results in all cases. This indicates that FedCRL is relatively robust and shows little sensitivity to the selection of hyperparameters within a wide range. It is worth noting that the worst result shown in Figure 4 is still better than the baseline. In addition, it is also found that when the number of clusters is set to 2 when it is small, the result is the best. This is because a single cluster cannot model the variability within a class and thus ignores the interference of outliers. As the number of clusters increases, the significance of important features may decrease, which may shift the attention to noise features.

[0142] 4. Relationship between Generating Different Numbers of Augmented Samples and Classification Results

[0143] As shown in Figure 5, in most cases, the more augmented features are generated, the greater the performance gain. This is because enriching the feature space through knowledge transfer can effectively simulate the real feature distribution. This can increase the information required to train a generalized model to alleviate the overfitting problem caused by insufficient feature numbers. It is worth noting that using fewer augmented samples can still improve the performance by approximately 3%. In addition, the inventors observed that when When this happens, the performance of CIFAR100 degrades. Part of this degradation is due to the fact that CIFAR100 has far fewer samples per class than CIFAR10, making feature learning difficult. Additionally, the greater number of classes in CIFAR100 increases the chance of overlapping class distributions, resulting in misleading information. Therefore, enhancing feature learning is of considerable significance in complex tasks.

[0144] 5. Intra-source Representation Calibration

[0145] This section aims to evaluate the impact of spatial distribution calibration on representation learning, prototype modeling, and model performance. The inventors selected two clients with different data distributions and class omissions and used t-SNE to visualize the feature distributions of two classes in the training test set. As shown in Figure 6, compared with FedAvg, FedProc and FedCRL learn more discriminative representations under the guidance of knowledge, which is particularly evident for classes with the majority of samples (e.g., the sand-brown class). However, FedProc faces challenges for classes characterized by a limited number of samples, which is due to the error accumulation caused by the following reasons: the differences in the target space between different training rounds. Notably, FedCRL uses consistent functions to guide local training, ensuring similar representations for shared classes even in the presence of missing classes. These factors contribute to the superior performance of FedCRL over other methods. Additionally, FedCRL also evaluates the importance of the generated prototypes. The results show that FedCRL assigns lower weights to prototypes in the overlapping regions of different classes. This helps the CSCRL module reduce the interference of outliers during the model calibration process.

[0146] 6. Error Analysis of FedCRL

[0147] This section presents a case study based on GradCAM visualization to delve into the working mechanism of FedCRL. As shown in Figure 7(a), incorporating the ISRC and CSCRL modules further enhances the discriminative ability of the basic method. This is reasonable because they can utilize a large number of samples to improve representation learning and cross-source feature alignment. Figure 7(b) shows that a limited sample size may lead to local model failure, thereby undermining the cooperation effectiveness. The ISRC module enhances representation learning through distribution calibration, helping to align heterogeneous features within the CSCRL module to correct prediction errors. Additionally, it is found that model calibration may fail. Although accurate predictions are made before calibration, their reliability is questionable. It cannot accurately focus on the target, and the low-quality representations it learns hinder the improvement of subsequent calibration performance, as shown in Figure 7(c). Finally, Figure 7Figure (d) illustrates the situation where these methods always mispredict classes using a small number of samples. It is worth noting that the CSCRL module promotes the model's attention to the target region and reduces the prediction difference between the ground truth and TOP-1. These findings verify the negative impact of data imbalance on joint learning and also confirm the effectiveness of the proposed framework.

[0148] 7. Cross-Source Consistent Representation Learning

[0149] Two clients are randomly selected, and two shared classes (birds with fewer samples and airplanes with more samples) are selected from them. Then, 200 samples are drawn from the test set of each class. Figures 8 and 9 illustrate the representation distributions of FedCSPC and FedCRL for the given samples, as well as their CKA similarities and model performances. The results show that, compared with FedCSPC, the representation distribution learned by FedCRL is more compact within classes and more discriminative between classes. At the same time, FedCRL can reduce the heterogeneity of the cross-client feature space before calibration, laying a solid foundation for cross-source feature comparison. This is also reflected in the CKA similarity. In contrast, compared with birds, FedCSPC achieves better alignment of the airplane class between the two clients. This is because the significant differences caused by the limited representation quality learned from a small number of samples hinder feature alignment. In addition, the inventors note that distribution differences may lead to a decline in collaborative performance (see Figure 8). Model calibration usually enhances the personalization ability of the local model by leveraging the knowledge of other clients to make up for its own deficiencies.

[0150] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A prototype-guided federated consistent representation learning system, characterized in that: include: In-source representation calibration module, used to model the representation distribution of the client, improve local training, and correct feature distribution on imbalanced data; At the same time, it provides prototype information to the server, including cluster prototype, cluster variance, and attention score; The class-aware text representation output by the pre-trained CLIP model is used to constrain the representation learning of the local model. The client stores a private data set and a private model. After the training is completed, the trained private model is used to extract features from the private data set. The extracted features are clustered, and the feature means within the cluster are regarded as prototypes. The variance of the features within the cluster in each dimension and the importance score of the cluster prototype are calculated. A cross-source consistent representation learning module is used to learn generalized projections and classifiers; the client sends private models and prototype information to the server; The server receives the private model and prototype information uploaded by the client, and uses statistical knowledge to generate enhanced features, refine the feature space, and improve diversity; Subsequently, the cross-source feature alignment module based on relational consistency maps the features from different sources to a unified space for comparison and classification, sums the parameters of the private models, and then takes the average to obtain the global model. The server uses the prototypes of all parties to calibrate the global projection head and global classifier in the global model, sends the calibrated global model to the client, and eliminates the interference of outliers based on the attention scores.

2. A prototype-guided federated consistent representation learning system according to claim 1, characterized in that: The in-source representation calibration module includes two processes: knowledge-guided representation calibration and clustering-driven class pattern modeling.

3. A prototype-guided federated consistent representation learning system according to claim 2, characterized in that: The knowledge-guided representation calibration specifically uses fixed category-aware text features learned from pre-trained CLIP as the optimization target for local feature learning to normalize the client's feature learning; it performs supervised prototype contrastive learning to maximize the image features in the latent space. and text prototype The consistency between them, the target loss is defined as: , in, Represents the client Medium The image features, , is a text encoder, represents the number of classes, The model's prediction for the image is When , it takes 1, otherwise it takes 0. For Clients The number of training data in represents the temperature parameter; at the same time, the empirical loss is used to further ensure the classification ability of the model, that is, , in, Represents the first elements, the dimension of the model's output vector is consistent with the number of categories, The element represents the The predicted probability of the class is a numerical value; similarly, Represents the first elements, is the label of the image.

4. A prototype-guided federated consistent representation learning system according to claim 2, characterized in that: The cluster-driven class pattern modeling is specifically: using a fixed model learned from the source representation calibration module Features are extracted from all training data and k-means clustering method is used to mine different patterns in the latent space, i.e., , in, Represents the client Medium Category data, is a hyperparameter that represents the number of clusters. Indicates category The characteristics of the cluster.

5. The prototype-guided federated consistent representation learning system according to claim 1, characterized in that: The cross-source consistent representation learning module obtains all local models and local prototype sets uploaded from the client, and aligns the prototype features from heterogeneous spaces, which mainly includes two processes: class-aware region calibration and cross-source feature alignment.

6. A prototype-guided federated consistent representation learning system according to claim 5, characterized in that: The class perception region calibration is specifically: using knowledge transfer technology to calibrate the class perception region, which uses a Gaussian model Generate extended features based on variance, and set the cluster variance corresponding to the client with the highest score and Divide Variance fusion outside the client to transfer important knowledge to other class features in the client ,Right now, , in, is a local prototype, Represents enhanced features, Indicates Enhanced features, It is an enhanced feature The number of represents the fusion variance, represents the cluster variance corresponding to the client with the highest score, It means that the mean is 0 and the variance is , generating enhanced features that satisfy Gaussian distribution.

7. The prototype-guided federated consistent representation learning system according to claim 5, characterized in that: The cross-source feature alignment is specifically as follows: for the original global model , global feature extractor No retraining is required, and the projection head and classifier Need to calibrate, that is .

8. The prototype-guided federated consistent representation learning system according to claim 7, characterized in that: The cross-source feature alignment adopts a two-level regularization method to refine representation learning, including local consistency matching and complementary consistency matching.

9. The prototype-guided federated consistent representation learning system according to claim 8, characterized in that: by Taking these three classes as an example, the local consistency matching is expressed in the following way: , in, Is a calibrated projection head After mapping, if Is the client Medium The local prototype of ;if is an enhanced feature, then , , represents the dot product; Indicates category The global prototype of all client local classes The average value of the prototype; , in, represents the distance-based consistency matching loss, is the Euclidean distance; The local matching loss is defined as: ; The complementary consistency match is defined as: 。 10. A prototype-guided federated consistency representation learning method, characterized in that: The system according to any one of claims 1 to 9 comprises: training a private model using private data in a client, and using a class-aware text representation output by a pre-trained CLIP model to constrain the representation learning of a local model, wherein a private data set and a private model are stored in the client, and after the training, the trained private model is used to extract features from the private data set, the extracted features are clustered, the feature means within the cluster are regarded as prototypes, and the variance of the features within the cluster in each dimension and the importance score of the cluster prototype are calculated; the client sends the private model and prototype information to the server; the server receives the private model and prototype information uploaded by the client, sums the parameters of the private model, and then takes the average value to obtain a global model; the server uses the prototypes of all parties to calibrate the global projection head and the global classifier in the global model, and sends the calibrated global model to the client.

Citation Information

Patent Citations

  • Federal learning system and method based on prototype guide cross training mechanism

    CN116452955A

  • Generator-based prototype confrontation method for heterogeneous federated learning

    CN117787385A