Federated learning method based on attention-guided feature distillation and prototype contrastive alignment
By adopting attention-guided feature distillation and prototype contrast alignment methods in federated learning, the conflict between data heterogeneity and personalization and generalization is solved, and the generalization ability and personalization accuracy of the model are improved.
Patent Information
- Application Number
- CN202510004062.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-02
AI Technical Summary
In federated learning, the conflict between data heterogeneity and personalization and generalization leads to poor model performance, especially when the client has extreme label distribution bias.
The federated learning method based on attention-guided feature distillation and prototype contrast alignment is adopted to guide students' model learning through attention mechanisms, retain historical local knowledge, and optimize the generalization ability of the model through prototype comparison and alignment.
It effectively resolves the conflict between data heterogeneity and personalization and generalization, improves the generalization ability of the model, and maintains personalized accuracy, reducing the loss of personalized knowledge caused by model updates.
Smart Images

Figure CN119416915B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and in particular to a federated learning method based on attention-guided feature distillation and prototype contrast alignment. Background Art
[0002] Federated learning (FL) is a distributed machine learning paradigm that aims to address issues such as data privacy and communication limitations by collaboratively learning a global model. In a federated learning environment, multiple edge devices train models on local data and then send model updates to a central server for aggregation, avoiding the need to transmit local data to the cloud. As an efficient communication and privacy-preserving learning solution, federated learning has shown great potential in real-world edge computing applications.
[0003] However, learning a single global model may fail when data is not independent and identically distributed across devices, or when each edge device performs different tasks. In these heterogeneous data scenarios, the performance of the global model on a specific client may be poor, and may even be lower than that of a model trained only locally. In addition, for multi-task applications, each edge device may have a different optimal model.
[0004] Furthermore, personalized federated learning (PFL) has attracted widespread attention to address the challenges in federated learning from a client-specific perspective. In PFL, each client can train a personalized model that performs better on its local task, benefiting from collaborative learning compared to non-collaborative local training. However, a key challenge of PFL is to strike a balance between sharing global knowledge and retaining local task-specific knowledge.
[0005] At present, in order to address this challenge, some relevant research results have emerged: methods based on parameter decoupling, such as FedRep and FedPer, divide model parameters into personalized and shared components. In these methods, only shared components are aggregated and transmitted to the server for communication between clients. Research on parameter decoupling has shown that dividing the training process into representation learning and prediction stages brings many benefits. For example, Mendieta et al. and Yu et al. have noticed that the prediction stage is closely related to the task type. These insights have promoted the trend of dividing model parameters into feature extractors (backbones) and classifiers (heads). Under this strategy, the backbone of the model is designated as a shared component to synchronize feature extraction capabilities between clients, while the classifier is customized as a personalized component, focusing on client-specific, task-related classification. This decoupling method retains historical local knowledge to a certain extent, especially FedRep, where only updating the head parameters can alleviate the problem of forgetting historical information to a certain extent. On the other hand, in the representation learning method, FedProto introduces the concept of using representations as sharable information, which not only respects privacy, but also enhances personalization by providing additional knowledge. Based on this concept, recent studies have shown that FedGH and FedPAC have achieved significant improvements by combining shared representations with global models to effectively extract insights from both models and representations. These methods are characterized by leveraging knowledge between representations with the same label, such as FedProto and FedPAC, which achieve regularization by calculating the distance between representations of the same label, while FedGH uses local representations to supervise the training of the global projection layer and also focuses on the alignment of representations of the same label.
[0006] However, existing methods are limited by the lack of cross-label collaboration, which may limit the generalization ability of FL when the client has extreme label distribution bias, resulting in a conflict between personalization and generalization. Therefore, designing a method that can solve the problem of forgetting historical information while resolving the conflict between personalization and generalization has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0007] The purpose of the present invention is to provide a federated learning method based on attention-guided feature distillation and prototype contrast alignment to solve the data heterogeneity, the conflict between personalization and generalization, and the problem of forgetting historical information in federated learning.
[0008] The technical solution adopted by the present invention is:
[0009] The federated learning method based on attention-guided feature distillation and prototype contrast alignment includes the following steps:
[0010] S1. Initialize global feature extractor: The server first initializes a global feature extractor, which is used to receive local model updates from each client and perform model aggregation;
[0011] S2, client local training: In each communication round, the server sends the global feature extractor to the selected client. Each client initializes the local feature extractor using the global feature extractor sent by the server and trains it using local private data;
[0012] S3. During the local model training process, each client adopts an attention-guided feature distillation method to retain and utilize historical local knowledge; specifically, the client uses the local feature extractor of the previous round as the teacher model and the global feature extractor of the current round as the student model, and guides the learning of the student model through the attention mechanism, thereby effectively retaining personalized historical knowledge.
[0013] S4. After feature distillation, the client performs prototype comparison and alignment operations to obtain updated local model parameters; that is, the local prototype is adjusted according to the global prototype received from the server, and the generalization ability of the model is optimized through prototype comparison learning loss and prototype alignment loss; this step aims to enhance the model's ability to distinguish data from different clients by dynamically adjusting the relationship between the global prototype and the local prototype.
[0014] S5. The client uploads the updated local model parameters and the calculated local prototype to the server. The server aggregates the model parameters and prototypes uploaded by all clients to update the global model parameters and global prototype to prepare for the next round of federated learning process.
[0015] S6. The server sends the updated global model parameters and global prototype back to the client, and the client uses the updated global parameters to initialize a new round of local model training; the iteration is performed until the model performance reaches a satisfactory level or the predetermined communication round is reached.
[0016] Furthermore, the method is based on two metrics: attention-guided feature distillation and prototype contrast alignment mechanism.
[0017] Furthermore, in step S3, the feature distillation based on attention guidance is used to achieve higher performance and stronger generalization ability, thereby reducing the loss of personality knowledge caused by model updating, which specifically includes the following steps:
[0018] S3-1. Use the local model from the previous round As the teacher model, the global model of the current round as a student model;
[0019] S3-2. Use the convolutional block attention module to generate channel and spatial attention maps to accurately guide features and improve the effectiveness of the distillation process. The channel and spatial attention modules are expressed as:
[0020]
[0021]
[0022] in, represents the feature map extracted in the block, Presentation layer, It is the feature map after channel attention modulation, and Represent global average pooling and global maximum pooling respectively, and Respectively represent the global average pooling and maximum pooling along the channel direction, represents a multilayer perceptron, represents the activation function, represents the convolutional layer, Indicates channel splicing, and Represent channel attention map and spatial attention map respectively;
[0023] S3-3. Apply channel and spatial attention to the original input feature map F to obtain the final attention map. The final attention map is expressed as:
[0024]
[0025] Among them, represents element-wise multiplication, and A is the final attention map that combines channel and spatial attention;
[0026] S3-4, a cross-layer connection path is established between student networks, and by adopting a cross-layer feature fusion strategy, effective integration is achieved between student features at different layers. Three types of convolutional layers are created for each input feature layer: , , . It is a horizontal convolution, which unifies the number of channels of input feature maps of different scales to a fixed number of output channels. Used to smooth the result after feature fusion. Used to enhance low-level features. The feature fusion step can be expressed as:
[0027]
[0028] in, Indicates upsampling using bilinear interpolation, Indicates The final attention map of the layer, Represents the final fusion result; It is a horizontal convolution, which unifies the number of channels of input feature maps of different scales to a fixed number of output channels; Used to smooth the result after feature fusion. Used to enhance low-level features;
[0029] S3-5, use the lowest layer feature map as the first horizontal convolution, perform feature fusion from bottom to top, and before fusion, further process the original input feature map of each layer through the convolution layer that enhances the low-level features to enhance the low-level features. Combine the upsampled high-level features, horizontal convolution features and enhanced low-level features, and then apply Smoothing. The final feature alignment process can be expressed as:
[0030]
[0031] in, Indicates students, represents the teacher, and uses mean squared error (MSE) loss for feature alignment; represents the feature alignment loss, which is used to quantify the difference between the student model and the teacher model at the feature level. Indicates that the student model is The final fusion result of the layer, Indicates that the teacher model is The attention map of the layer, The student model is The feature map of the layer, The teacher model is The feature map of the layer.
[0032] Furthermore, step S4 introduces prototype contrast learning based on prototype alignment, which specifically includes the following steps:
[0033] S4-1. Using prototypes to assist local adaptive aggregation and personalized training, the prototype of category c is represented as:
[0034]
[0035] in, Indicates Client local dataset The category c subset contains all samples corresponding to label c; Local Datasets The label set in represents the feature representation corresponding to the k-th sample of category c in the i-th client;
[0036] S4-2. By collecting local data sets The local prototype of the i-th client is represented as:
[0037]
[0038] S4-3. Given a category c, the server receives prototypes from a group of clients with category c; after aggregating the prototypes, the global prototype of category c is obtained, which is specifically expressed as:
[0039]
[0040] in, Indicates that the category The client collection of
[0041] S4-4. Global Model Distributed Using Servers The local dataset on client i Generate feature embeddings on Denoted as data set D i For sample-label pairs, there is feature embedding :
[0042]
[0043] in, Indicates the number of images on client i;
[0044] S4-5. Based on the global prototype received from the server Embed the output features Alignment, , Indicates the global prototype of all categories of all clients in the tth iteration; Global prototype; the specific expression is as follows:
[0045]
[0046] in, is labeled as category c, represents the global prototype of category c; represents the prototype alignment loss of the i-th client in the t+1th iteration; the target It is called mean squared error loss, which aims to minimize the squared Euclidean distance between features and class prototypes;
[0047] S4-6. Learning representations by comparing positive and negative pairs helps to achieve personalization by using positive and negative pairs between local and global representations. Positive and negative representation pairs: Intuitively, for the feature representation labeled as category c , the global prototype The same labels in are regarded as positive samples, while other labels are regarded as negative samples. Given that the global prototype corresponds to the category |C| of all clients, a positive pair is constructed. and negative pair :
[0048]
[0049]
[0050]
[0051] Among them, c represents the category label, Represented as the client's index, where , refers to the different clients participating in federated learning; N represents the number of clients; Represents the index of a data point in a specific client, Indicates the number of data points in the client; Representation and Other category labels of different categories; Indicates except Global prototypes of categories other than their own; Indicates Client The labels of the data points, No. Client data points, C represents all label categories;
[0052] S4-7. Calculate the global prototype contrast loss using positive and negative pairs , the calculation expression of the global prototype contrast loss is as follows:
[0053]
[0054]
[0055] Among them, τ is a temperature hyperparameter used to adjust the degree of attention paid to positive and negative samples, and is the cosine similarity; Indicates that in the t+1th iteration, for the Client datasets Data points in The expected value of , Indicates Client data samples;
[0056] S4-8. Calculate the final objective function of the t-th iteration local update as:
[0057]
[0058] in, ; represents the cross entropy loss function; , represents the classifier function; No. Client No. Feature distillation loss in round iterations; , , is a hyperparameter that balances various losses.
[0059] The present invention adopts the above technical solution. In each round of federated learning communication, when the client receives the global feature extractor and starts the local model update, attention-guided feature distillation is performed; then after the client completes the local update of the feature extractor and classifier, and before uploading the local model update to the server, a prototype comparison and alignment mechanism is performed, which achieves the enhancement of the generalization ability of the model while retaining personalized historical knowledge, improving the generalization ability of the model while maintaining personalized accuracy. The present invention solves the conflicts between privacy protection, data heterogeneity, personalization and generalization in federated learning through attention-guided feature distillation technology and prototype comparison and alignment mechanism. The technical solution of the present invention is suitable for scenarios where global models need to be trained collaboratively among multiple clients, especially in applications where data privacy protection and communication efficiency are critical. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments;
[0061] Figure 1 It is a flowchart of the federated learning method based on attention-guided feature distillation and prototype comparison alignment of the present invention;
[0062] Figure 2 Schematic diagram of the system framework of the federated learning method based on attention-guided feature distillation and prototype contrast alignment of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0064] like Figure 1 As shown in or 2, the present invention discloses a federated learning method based on attention-guided feature distillation and prototype contrast alignment, comprising the steps of:
[0065] S1. Initialize the global feature extractor. The server first initializes a global feature extractor. The global feature extractor is used to receive local model updates from each client and perform model aggregation;
[0066] S2, client local training. In each communication round, the server sends the global feature extractor to the selected client. Each client initializes its local feature extractor using the global feature extractor sent by the server and trains it using local private data;
[0067] S3. During the local model training process, each client uses an attention-guided feature distillation method to retain and utilize historical local knowledge. Specifically, the client uses the local feature extractor of the previous round as the teacher model and the global feature extractor of the current round as the student model, and guides the learning of the student model through the attention mechanism, thereby effectively retaining personalized historical knowledge;
[0068] S4. After feature distillation, the client performs prototype comparison and alignment operations. This includes adjusting the local prototype according to the global prototype received from the server, and optimizing the generalization ability of the model through prototype comparison learning loss and prototype alignment loss. This step aims to enhance the model's ability to distinguish between different client data by dynamically adjusting the relationship between global prototypes and local prototypes;
[0069] S5. The client uploads the updated local model parameters and the calculated local prototype to the server. The server then aggregates the model parameters and prototypes uploaded by all clients to update the global model parameters and global prototype in preparation for the next round of federated learning process.
[0070] S6. The server sends the updated global model parameters and global prototype back to the client, and the client uses these global parameters to initialize a new round of local model training. This process is repeated until the model performance reaches a satisfactory level or the predetermined communication round is reached.
[0071] Furthermore, the method is based on two metrics: attention-guided feature distillation and prototype contrast alignment mechanism.
[0072] Furthermore, in step S3, the feature distillation based on attention guidance is used to achieve higher performance and stronger generalization ability, thereby reducing the loss of personality knowledge caused by model updating, which specifically includes the following steps:
[0073] S3-1. Use the local model from the previous round As the teacher model, the global model of the current round as a student model;
[0074] S3-2. Use the convolutional block attention module to generate channel and spatial attention maps to accurately guide features and improve the effectiveness of the distillation process. The channel and spatial attention modules are expressed as:
[0075]
[0076]
[0077] in, represents the feature map extracted in the block, Presentation layer, It is the feature map after channel attention modulation, and Represent global average pooling and global maximum pooling respectively, and Respectively represent the global average pooling and maximum pooling along the channel direction, represents a multilayer perceptron, represents the activation function, represents the convolutional layer, Indicates channel splicing, and Represent channel attention map and spatial attention map respectively;
[0078] S3-3. Apply channel and spatial attention to the original input feature map F to obtain the final attention map. The final attention map is expressed as:
[0079]
[0080] Among them, represents element-wise multiplication, and A is the final attention map that combines channel and spatial attention;
[0081] S3-4, a cross-layer connection path is established between student networks, and by adopting a cross-layer feature fusion strategy, effective integration is achieved between student features at different layers. Three types of convolutional layers are created for each input feature layer: , , . It is a horizontal convolution, which unifies the number of channels of input feature maps of different scales to a fixed number of output channels. Used to smooth the result after feature fusion. Used to enhance low-level features. The feature fusion step can be expressed as:
[0082]
[0083] in, Indicates upsampling using bilinear interpolation, Indicates The final attention map of the layer, Represents the final fusion result; It is a horizontal convolution, which unifies the number of channels of input feature maps of different scales to a fixed number of output channels; Used to smooth the result after feature fusion. Used to enhance low-level features;
[0084] S3-5, use the lowest layer feature map as the first horizontal convolution, perform feature fusion from bottom to top, and before fusion, further process the original input feature map of each layer through the convolution layer that enhances the low-level features to enhance the low-level features. Combine the upsampled high-level features, horizontal convolution features and enhanced low-level features, and then apply Smoothing. The final feature alignment process can be expressed as:
[0085]
[0086] in, Indicates students, represents the teacher, and uses mean squared error (MSE) loss for feature alignment; represents the feature alignment loss, which is used to quantify the difference between the student model and the teacher model at the feature level. Indicates that the student model is The final fusion result of the layer, Indicates that the teacher model is The attention map of the layer, The student model is The feature map of the layer, The teacher model is The feature map of the layer.
[0087] Furthermore, step S4 introduces prototype contrast learning based on prototype alignment, which specifically includes the following steps:
[0088] S4-1. Using prototypes to assist local adaptive aggregation and personalized training, the prototype of category c is represented as:
[0089]
[0090] in, Indicates Client local dataset The category c subset contains all samples corresponding to label c; Dataset local The label set in represents the feature representation corresponding to the k-th sample of category c in the i-th client;
[0091] S4-2. By collecting local data sets The local prototype of the i-th client is represented as:
[0092]
[0093] S4-3. Given a category c, the server receives prototypes from a group of clients with category c; after aggregating the prototypes, the global prototype of category c is obtained, which is specifically expressed as:
[0094]
[0095] in, Indicates that the category The client collection of
[0096] S4-4. Global Model Distributed Using Servers The local dataset on client i Generate feature embeddings on Denoted as data set D i For sample-label pairs, there is feature embedding :
[0097]
[0098] in, Indicates the number of images on client i;
[0099] S4-5. Based on the global prototype received from the server Embed the output features Alignment, , Indicates the global prototype of all categories of all clients in the tth iteration; Global prototype; the specific expression is as follows:
[0100]
[0101] in, is labeled as category c, represents the global prototype of category c; represents the prototype alignment loss of the i-th client in the t+1th iteration; the target It is called mean squared error loss, which aims to minimize the squared Euclidean distance between features and class prototypes;
[0102] S4-6. Learning representations by comparing positive and negative pairs helps to achieve personalization by using positive and negative pairs between local and global representations. Positive and negative representation pairs: Intuitively, for the feature representation labeled as category c , the global prototype The same labels in are regarded as positive samples, while other labels are regarded as negative samples. Given that the global prototype corresponds to the category |C| of all clients, a positive pair is constructed. and negative pair :
[0103]
[0104]
[0105]
[0106] Among them, c represents the category label, Represented as the client's index, where Related to the i mentioned above, it refers to the different clients participating in federated learning, and N represents the number of clients; Represents the index of a data point in a specific client, Indicates the number of data points in the client; Representation and Other category labels of different categories; Indicates except Global prototypes of categories other than their own; Indicates Client The labels of the data points, No. Client data points, C represents all label categories;
[0107] S4-7. Calculate the global prototype contrast loss using positive and negative pairs , the calculation expression of the global prototype contrast loss is as follows:
[0108]
[0109]
[0110] Among them, τ is a temperature hyperparameter used to adjust the degree of attention paid to positive and negative samples, and is the cosine similarity; Indicates that in the t+1th iteration, for the Client datasets Data points in The expected value of , Indicates Client data samples;
[0111] S4-8. Calculate the final objective function of the t-th iteration local update as:
[0112]
[0113] in, ; represents the cross entropy loss function; , represents the classifier function; No. Client No. Feature distillation loss in round iterations; , , is a hyperparameter that balances various losses.
[0114] The beneficial effects of the present invention are as follows: the present invention provides a personalized federated learning method based on attention-guided feature distillation and prototype comparison and alignment mechanism, in which, in each round of federated learning communication, when the client receives the global feature extractor and starts updating the local model, attention-guided feature distillation is performed; then, after the client completes the local update of the feature extractor and classifier, and before uploading the local model update to the server, a prototype comparison and alignment mechanism is performed, thereby enhancing the generalization ability of the model while retaining personalized historical knowledge, improving the generalization ability of the model, and maintaining personalized accuracy.
[0115] Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians of the art without making creative work are within the scope of protection of the present application.
Claims
1. A federated learning method based on attention-guided feature distillation and prototype contrast alignment, characterized by: It includes the steps of: S1. Initialize global feature extractor: The server first initializes a global feature extractor, which is used to receive local model updates from each client and perform model aggregation; S2, client local training: In each communication round, the server sends the global feature extractor to the selected client. Each client initializes the local feature extractor using the global feature extractor sent by the server and trains it using local private data; S3. During the local model training process, each client uses an attention-guided feature distillation method to retain and utilize historical local knowledge. Step S3 specifically includes the following steps: S3-1. Use the local model from the previous round As the teacher model, the global model of the current round as a student model; S3-2. Use the convolutional block attention module to generate channel and spatial attention maps; the channel and spatial attention maps are expressed as: in, represents the feature map extracted in the block, Presentation layer, It is the feature map after channel attention modulation, and Represent global average pooling and global maximum pooling respectively, and Respectively represent the global average pooling and maximum pooling along the channel direction, represents a multilayer perceptron, represents the activation function, represents the convolutional layer, Indicates channel splicing, and Represent the channel attention map and spatial attention map respectively; S3-3. Apply channel and spatial attention to the original input feature map F to obtain the final attention map. The final attention map is expressed as: Among them, represents element-wise multiplication, and A is the final attention map that combines channel and spatial attention; S3-4. Establish cross-layer connection paths between student networks and create three types of convolutional layers for each input feature layer. , , , the final fusion result is obtained by adopting the cross-layer feature fusion strategy to achieve effective integration between student features at different layers; the feature fusion steps are expressed as: in, Indicates upsampling using bilinear interpolation, Indicates The final attention map of the layer, Represents the final fusion result; It is a horizontal convolution, which unifies the number of channels of input feature maps of different scales to a fixed number of output channels; Used to smooth the result after feature fusion. Used to enhance low-level features; S3-5, use the lowest layer feature map as the first horizontal convolution to perform feature fusion from bottom to top; combine the upsampled high-level features, horizontal convolution features and enhanced low-level features, and then apply Smoothing is performed to obtain the final feature alignment. The specific expression is: in, Indicates students, represents the teacher, and uses the mean square error MSE loss for feature alignment; represents the feature alignment loss, which is used to quantify the difference between the student model and the teacher model at the feature level. Indicates that the student model is The final fusion result of the layer, Indicates that the teacher model is The attention map of the layer, Indicates that the student model is The feature map of the layer, Indicates that the teacher model is Feature map of the layer; S4. After feature distillation, the client performs prototype comparison and alignment operations to obtain updated local model parameters; that is, the local prototype is adjusted according to the global prototype received from the server, and the generalization ability of the model is optimized through prototype comparison learning loss and prototype alignment loss; S5. The client uploads the updated local model parameters and the calculated local prototype to the server. The server aggregates the model parameters and prototypes uploaded by all clients to update the global model parameters and global prototype to prepare for the next round of federated learning process. S6. The server sends the updated global model parameters and global prototype back to the client, and the client uses the updated global parameters to initialize a new round of local model training; the iteration is performed until the model performance reaches a satisfactory level or the predetermined communication round is reached.
2. The federated learning method based on attention-guided feature distillation and prototype contrast alignment according to claim 1, characterized in that: In step S3, the client uses the local feature extractor of the previous round as the teacher model and the global feature extractor of the current round as the student model, and guides the learning of the student model through the attention mechanism.
3. The federated learning method based on attention-guided feature distillation and prototype contrast alignment according to claim 1, characterized in that: In S3-5, before feature fusion, the original input feature map of each layer is processed by a convolutional layer that enhances low-level features to enhance low-level features.
4. The method for federated learning based on attention-guided feature distillation and prototype contrast alignment according to claim 1, characterized in that: In step S4, the relationship between the global prototype and the local prototype is dynamically adjusted to enhance the model's ability to distinguish between different client data.
5. The method for federated learning based on attention-guided feature distillation and prototype contrast alignment according to claim 1, characterized in that: Step S4 specifically includes the following steps: S4-1. Using prototypes to assist local adaptive aggregation and personalized training, the prototype of the category label c is represented as: ; in, Indicates Client local dataset The category label c subset contains all samples corresponding to category label c; Dataset The label set in represents the feature representation corresponding to the k-th sample with category label c in the i-th client; S4-2. By collecting local data sets The local prototype of the i-th client is represented as: S4-3. Given a category label c, the server receives prototypes from a group of clients with category label c; after aggregating the prototypes, the global prototype of category label c is obtained, which is specifically expressed as: in, Indicates The category label at the round iteration is The global prototype of represents the set of clients containing category label c; S4-4. Global Model Distributed Using Servers The local dataset on client i Generate feature embeddings on Denoted as data set D i For sample-label pairs, we have feature embedding : ; in, Represents a local dataset, Indicates the number of images on client i. Represents K rows The feature space size of the column; S4-5. Based on the global prototype received from the server Embed the output features Alignment, , It represents the global prototype of all category labels c of all clients in the tth iteration; the specific expression is as follows: ; in, is labeled with category label c; represents the prototype alignment loss of the i-th client in the t+1th iteration; S4-6. Feature representation for the category label c , the global prototype The same labels in are considered as positive samples, while other labels are considered as negative samples; given that the global prototype corresponds to the number of all category labels of all clients |C|, construct a positive pair and negative pair : Among them, c represents the category label, Represented as the client's index, where , N represents the number of clients; Represents the index of a data point in a specific client, Indicates the number of data points in the client; Representation and Other category labels of different categories; Indicates except Global prototypes of categories other than their own; Indicates Client The labels of the data points, The generated embedded data Client data points, C represents all category labels; S4-7. Calculate the global prototype contrast loss using positive and negative pairs , the calculation expression of the global prototype contrast loss is as follows: Among them, τ is a temperature hyperparameter used to adjust the degree of attention paid to positive and negative samples, and is the cosine similarity; Indicates that in the t+1th iteration, for the Client datasets Data points in The expected value of Indicates Client No. The data set at the time of round iteration, Indicates of clients data samples; S4-8. Calculate the final objective function of the t-th iteration local update as: in, Expressed as Client No. Cross entropy loss in round iterations; represents the cross entropy loss function; Expressed as Client No. The classifier parameters in the round iteration, represents the classifier function; Indicates Client No. Feature distillation loss in round iterations; , , is a hyperparameter that balances various losses.
Citation Information
Patent Citations
Personalized federal learning method based on decoupling knowledge distillation
CN117152480A
Personalized federal learning method based on prototype alignment learning
CN118674015A