Personalized federal learning method and system based on domain invariant text representation and intra-domain global prior

By introducing a personalized federated learning method with domain invariant text representation and a global prior in the domain, the problems of global goal instability and data heterogeneity in personalized federated learning are solved, and the stability and fairness of the model are improved.

CN120508883AActive Publication Date: 2025-08-19ZHEJIANG UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511000078.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-19
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

The existing personalized federated learning method is susceptible to dominant clients, noise data interferes with, and global targets in heterogeneous scenarios of cross-client data distribution, resulting in degradation in model performance and insufficient generalization capabilities.

Method used

A personalized federated learning method based on domain invariant text representation and global prior in the domain is adopted. By introducing text embedding as global representation and prior in the domain, a stable and unbiased global feature representation is generated, and the local model is optimized using graphic and text alignment loss and prior in the domain.

Benefits of technology

Effectively narrow the intra-class distance, expand the inter-class distance, alleviate the problem of data heterogeneity, improve model performance and stability, and enhance the fairness and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508883A_ABST
    Figure CN120508883A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized federal learning method and system based on domain invariant text representation and intra-domain global prior, and the method comprises the steps: introducing text embedding as global representation, and generating stable and unbiased global feature representation with domain invariance through a category label described by a natural language; through the mutual guidance of the local features and the text embedding vectors, the intra-class distance can be effectively reduced and the inter-class distance can be effectively expanded, so that the problem of data isomerism is relieved. Meanwhile, the introduced intra-domain priori module generates global sample priori by aggregating observable data embedding, and can help a local model to better understand global data distribution. Compared with a traditional global model or prototype generated by depending on image features, the method is based on text embedding and intra-domain global prior, does not depend on local data quality, naturally has denoising and generalization capabilities, and effectively relieves the problem that a global target changes along with training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distributed machine learning technology, and relates to a personalized federated learning method and system based on domain-invariant text representation and global prior within the domain. It is a personalized federated learning method combined with global prior within the domain, and is particularly suitable for optimizing domain-invariant text representation in scenarios with heterogeneous data distribution across clients. Background Art

[0002] With the increasing demand for data privacy protection, federated learning (FL), a distributed machine learning technology that enables collaborative modeling without centralized data, has attracted increasing attention. This technology allows multiple devices or organizations to jointly train a global model without sharing the original data. It is widely used in scenarios with extremely high data privacy requirements, such as healthcare, finance, and mobile devices.

[0003] Traditional federated learning methods, such as FedAvg, build a global model by aggregating local model parameters through weighted averaging, enabling collaborative learning across devices or organizations. However, because the data distribution of each client often exhibits significant statistical heterogeneity (non-IID), directly aggregating local models often leads to degraded global model performance and poor performance on some clients (Kairouz P, McMahan HB, Avent B, et al. Advances and open problems infederated learning[J]. Foundations and trends® in machine learning, 2021, 14(1–2): 1-210).

[0004] To address this issue, personalized federated learning (pFL) has become a research hotspot in recent years. This approach aims to learn an optimal personalized model for each client based on shared global information. Existing personalized federated learning methods primarily use two approaches to guide local model learning: one uses the global model as a regularizer to constrain the local model (e.g., pFedMe and Ditto); the other uses global prototypes to align local feature distributions (e.g., FedProto and FedPHP).

[0005] However, these existing methods based on global models or prototypes often have the following shortcomings (e.g. Figure 1(a), (b)): 1. Global information is easily affected by the dominant client: the global model or prototype is usually an aggregation of models of clients with different distributions, which is prone to favoring dominant clients with large data volume or good quality, thus becoming an unfair convergence target; 2. High risk of misleading local models: low-quality or noisy local models will affect the quality of global information after uploading, and then mislead the training of other clients; 3. The convergence target keeps changing: since each round of aggregation is based on the latest local model, the distribution of the global model or prototype keeps changing during the collaborative training process, making the local model always optimized towards a dynamic target, which is difficult to converge stably.

[0006] These issues limit the effectiveness and generalization of personalized federated learning methods in non-IID scenarios. Therefore, a global guidance mechanism that can provide stable, unbiased, and domain-invariant guidance is urgently needed to improve local model performance and enhance system stability and fairness. Summary of the Invention

[0007] The present invention addresses the shortcomings of existing technologies by proposing a personalized federated learning method and system based on domain-invariant text representations and global priors within the domain. This method introduces text embeddings as a global representation, leveraging category labels described in natural language to generate a stable, unbiased, and domain-invariant global feature representation. Compared to traditional global models or prototypes that rely on image feature generation, text embeddings are independent of local data quality and inherently possess denoising and generalization capabilities, effectively alleviating the problem of global objective variation during training.

[0008] The technical solution adopted in the present invention is as follows:

[0009] A personalized federated learning method based on domain-invariant text representation and domain-wide global priors, including the introduction of domain-invariant text representation and domain-wide priors into the process of training local data on each client to obtain a local model. Specifically:

[0010] The training process uses bimodal input of images and text collections. The input image is extracted through the image encoder participating in the training to obtain image features, and the text collection is extracted through the text encoder participating in the training and the frozen text encoder to obtain training text embedding and frozen text embedding respectively. In this process, the contrast loss of image-text alignment is used to optimize the matching between image features and text embedding features. The frozen text encoder receives global parameters sent by the server and locks the update.

[0011] The training text is embedded in the domain prior module to calculate the domain prior. The domain prior and image features are added and input into the trained classifier. At the same time, the image features are input into the frozen classifier separately. The results of the trained classifier and the frozen classifier are added and output as the prediction result. This process is supervised by the task loss.

[0012] The frozen text encoder and frozen classifier are respectively the global text encoder and global classifier downloaded from the server after the previous iteration step, and are locked for update in the current iteration step.

[0013] In the above technical solution, further, the contrast loss of the image-text alignment It consists of two symmetric log-negative probability terms, where and Represent the normalized similarity probabilities based on frozen text embedding and training text embedding respectively. middle is the ratio of the two parts, where the numerator is , represents the image features Frozen text embedding with similar tags The negative exponential similarity between ; the denominator is composed of repeated numerators ( ) and the negative sample summation term, the negative sample summation term ( ) represents the image features and all heterogeneous frozen text embeddings The second term is the sum of the similarities of middle A structure symmetrical to the first one, but with different parameters, using image features that do not participate in gradient updates , text embedding uses trainable text embedding and . Its molecular parts are: , indicating image features that do not participate in the update Training text embeddings with the same labels as image features The negative exponential similarity of ; its denominator is composed of repeated numerators ( ) and the negative sample summation term, the negative sample summation term ( ) indicates image features that do not participate in the update With all heterogeneous training text embeddings The sum of the similarities. That is:

[0014] .

[0015] in, is the image feature, is the frozen text embedding of the same label as the image feature, is the frozen text embedding with different class labels from the image features, is the image feature that does not participate in the update, To train text embeddings with the same labels as image features, To train text embeddings with different class labels than image features, is the similarity function, expressed as: .

[0016] Furthermore, the intra-domain prior module calculates the intra-domain prior using the data distribution of each classification of the training text embedding and training set statistics, specifically:

[0017] Set the global tag Each label in the Each client generates a corresponding set of text embedding vectors: ,in Indicates the Client corresponds to Text embedding representation of the class;

[0018] By combining the proportion of samples of each category in the local data of the client in the number of local data categories, the text embedding representation in the text embedding vector set is weighted averaged to obtain the domain prior of the client.

[0019] Furthermore, the domain prior is:

[0020] ;

[0021] in Is the indicator function, used to judge the sample Tags Belongs to category , is the number of categories present in the local data.

[0022] Furthermore, the task loss is , which calculates the one-hot encoding vector of the true label The cross entropy between the prediction output and the model is used to measure the degree of match between the two. The model prediction output is obtained by the sum of two parts: the first part is the classifier result during the training process , the second part is the result of a frozen classifier The outputs of these classifiers are passed through the softmax function Normalization is performed to generate the final probability distribution. That is:

[0023]

[0024] in, is the one-hot encoded vector of the true label, represents the softmax function, is the result of the trained classifier, The result of the frozen classifier.

[0025] Furthermore, the goal of local training for each client is to minimize the sum of the contrast loss and task loss of image-text alignment.

[0026] Furthermore, after the current iteration step is completed, all clients upload the trained local models to the server, including the text encoder, image encoder and trained classifier involved in the training. The server calculates the corresponding average models respectively to obtain the global model, including the global text encoder, global image encoder and global classifier, and transmits it back to each client. The client initializes the text encoder and frozen text encoder involved in the training as the global text encoder, the image encoder as the global image encoder, and the classifier involved in the training and the frozen classifier as the global image classifier. During the local training process, the parameters of the frozen text encoder and frozen classifier are locked and not updated, and a loop iteration is performed.

[0027] The present invention also provides a personalized federated learning system based on domain-invariant text representation and domain-wide global priors, for implementing the method described in any one of the above items.

[0028] The present invention further provides an electronic device, comprising:

[0029] one or more processors;

[0030] a memory for storing one or more programs;

[0031] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above methods.

[0032] The present invention also provides a computer-readable storage medium storing computer-executable instructions, which are used to implement any of the above methods when executed.

[0033] The beneficial effects of the present invention are:

[0034] The method of the present invention converts data labels into text descriptions and uses text embedding vectors as global representations. These embedding vectors are unbiased and will not favor any specific client. Through the mutual guidance of local features and text embedding vectors, the intra-class distance can be effectively narrowed and the inter-class distance can be widened, thereby alleviating the problem of data heterogeneity. The text embedding distribution remains stable during the collaborative learning process, providing a consistent target for model optimization. This stability, coupled with the noise-resistant and domain-invariant properties of text embedding, makes it an ideal global representation. At the same time, the introduced intra-domain prior module generates a global sample prior by aggregating observable data embeddings, which can help local models better understand the global data distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Comparison of the optimization process and distribution of different global representations: (a) The global prototype representation is susceptible to noise data and its distribution changes significantly during collaborative training; (b) The global model representation also produces a continuously changing distribution during collaborative learning; (c) When text embedding is used as a global representation, it is not only unaffected by noise data but also maintains stable distribution characteristics during training.

[0036] Figure 2 The training structure framework (left) and communication mechanism (right) in the method of the present invention.

[0037] Figure 3 It is a flow chart of the intra-domain prior module in the method of the present invention.

[0038] Figure 4 1 is a curve showing the test accuracy variation of different methods during the communication round in the embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.

[0040] The present invention provides a personalized federated learning method based on domain-invariant text representation and domain global prior. It is a domain-invariant text representation optimization method that is particularly suitable for scenarios with heterogeneous cross-client data distribution and combines domain global prior. It is referred to as FedDTR. This method is based on text embedding as a global representation and has unique advantages over other different global representation methods, such as Figure 1As shown in the figure, (a) the global prototype representation is easily affected by noisy data and its distribution changes significantly during the collaborative training process; (b) the global model representation also produces a continuously changing distribution during the collaborative learning process; (c) when text embedding is used as a global representation, it is not only unaffected by noisy data, but also maintains a stable distribution during the training process. The following is a detailed description of the technical implementation of the FedDTR framework (its training structure and communication process are shown in Figure 2). Figure 2 shown).

[0041] To address the shortcomings of existing methods, this paper innovatively introduces two core components: a domain-invariant text representation and a domain-specific prior module. Text embeddings guide local feature extraction through a dual mechanism: they cluster features of similar samples more closely together, while also widening the distance between features of different categories. Furthermore, these text embeddings dynamically adjust based on sample characteristics, optimizing the mapping from text space to image space.

[0042] The domain prior module integrates local samples and text embeddings to provide domain-specific global prior knowledge to each client. This design has two advantages: 1) it enhances the model's ability to understand the global data distribution; 2) it effectively suppresses the overfitting of personalized models on local data. With classification header The present invention adopts a parameter diversion strategy: upload the trained classification head parameters to the server and download the global classification head parameters to the frozen head. , in order to achieve efficient fusion of global information.

[0043] Domain-invariant text representation:

[0044] This section focuses on how to build a fair and stable global convergence target through unbiased and domain-invariant text embedding. FedDTR uses bimodal input: ,in Indicates the The input image of the client, For a text collection (such as Corresponding to the sentence "This is a cat", the category label is described using natural language). The processing flow is as follows:

[0045] 1. Feature extraction stage: Image features: through the encoder generate . Text features: text set Converted into a vector set after tokenization (C is the number of categories, d is the vector dimension), and then through the text encoder Output embedding vector .

[0046] 2. Joint Optimization Objective: Designing a Contrastive Loss Function Achieve dual optimization: align image features with similar text embeddings (narrowing intra-class distances) and map text embeddings to image space (enhancing category representation). The similarity calculation is defined as .

[0047] 3. Stability enhancement mechanism: To address the problem of text embedding fluctuations during training, a frozen text encoder is introduced. : Receive global parameters sent by the server and lock the update. Generate consistent global text embedding across clients .

[0048]

[0049] in is the image feature that does not participate in the gradient update.

[0050] In-domain prior module:

[0051] In the federated learning process, the data distribution of each client is usually independent and inconsistent, that is, the data of each client comes from a different data domain. clients, and the label set of their training data is denoted as , where the global label set is , and the client local dataset is , the corresponding label set is .

[0052] In traditional personalized federated learning methods, the task of each local model can be expressed as , that is, the model only performs classification predictions on the label categories contained locally. In order to achieve cross-client collaborative training, the local model task needs to be expanded to , that is, making predictions in the entire global label space. However, this extension will cause the classification structure of the model to change, making the local model unable to focus on its original subtask, affecting the training effect.

[0053] To solve the above problems, the present invention proposes a task guidance mechanism based on intra-domain priors. By introducing intra-domain priors, it helps the local model understand the overall data distribution, alleviates the impact of changes in task structure, and thus improves personalized modeling capabilities.

[0054] The specific technical solutions are as follows:

[0055] 1. Introduction of text embedding representation: First, the global label set Each class of labels in is converted into a standardized text description to generate a corresponding set of text embedding vectors: .in Indicates the Client corresponds to Text embedding representation of classes.

[0056] 2. Construction of domain prior representation: This paper combines the sample ratio of each category in the local data and performs weighted average on the text embedding to obtain the domain prior representation of the client. , the calculation formula is as follows:

[0057]

[0058] in Is an indicator function used to determine whether the sample belongs to a category , is the number of categories that exist locally. It represents the global prior of the category distribution in the client domain. By fusing this prior with image features and inputting it into the model, it helps the local model obtain global information from a specific perspective and improves its adaptability to the overall task.

[0059] Target optimization:

[0060] 1. Dual classification head structure design: In order to retain local personalization capabilities while using global knowledge to guide training, this paper divides the model classifier into two parts: the local head and frozen header : Local Header Receive local training updates; freeze head The global classifier parameters downloaded to the server are retained and not updated during training. Both participate in predictions during inference, achieving a balance between personalization and versatility through fusion output.

[0061] 2. Local optimization objective function design: The overall loss function consists of two parts: the contrast loss of image-text alignment , used to optimize the matching between image features and text embedding; task loss , used for training the model classifier, is defined as follows:

[0062] in represents the softmax function, is the one-hot encoded vector of the true label.

[0063] The final goal of local training is to minimize the following comprehensive loss: .

[0064] 3. Communication and Collaboration Strategy: In each round of communication, the server and client exchange text encoder, image encoder, and classifier parameters, and uniformly freeze the models for global alignment across clients to ensure global information consistency.

[0065] According to a specific embodiment of the present invention, the training process and reasoning process of the method of the present invention are specifically as follows:

[0066] Training process: When a batch of training data (images and corresponding labels) is input, the image features are first obtained by the image encoder, and then the corresponding labels are converted into natural language text descriptions. The resulting text set is first encoded into a text vector set using word2vec. The text vector set is simultaneously input into the text encoder participating in the training and the text encoder with frozen parameters to obtain the training text embedding and the frozen text embedding. Image features and frozen text embedding are based on , so that the image features are brought closer to the text embedding of the corresponding label and pushed away from the text embedding of other labels. Train text embeddings and image features so that the corresponding training text embeddings are brought closer to the image features and other text embeddings are pushed away from the image features.

[0067] The domain prior module calculates the domain prior using the training text embedding and the data distribution of each category (percentage of each class) of the training set statistics. The domain prior and image features are added one by one and input into the training classifier. The image features are input into the frozen classifier separately. The results of the two classifiers are added together to form the final result, which is determined by the objective function. Supervised training.

[0068] After each training session, the client uploads the trained text encoder, image encoder, and image classifier to the server. The server calculates the average model for each model as the global model and returns the global text encoder, global image encoder, and global image classifier to the client. The client initializes the trained text encoder and frozen text encoder to the global text encoder, the image encoder to the global image encoder, and the trained image classifier and frozen image classifier to the global image classifier. This process continues in an iterative loop.

[0069] Inference process: An image is input and passed through a trained image encoder to obtain image features. A text vector set is passed through a trained text encoder to obtain text embeddings. The text embeddings are passed through the domain prior module to calculate the domain prior. The domain prior and image features are added one by one and input into the trained classifier. The image features are separately input into the frozen classifier. The results of the two classifiers are added together to obtain the final prediction result.

[0070] Experimental verification

[0071] To verify the performance of the proposed method, FedDTR, in terms of effectiveness, scalability, stability, and convergence speed, this example conducted systematic comparative experiments with 13 currently mainstream personalized federated learning methods based on multiple public image classification and natural language processing tasks, and designed ablation experiments to analyze the contributions of each proposed sub-module.

[0072] 1. Experimental Setup

[0073] (1) Comparison methods. The selected comparison methods cover the following categories: traditional federated learning methods: FedAvg, FedProx; global and local structure decoupling methods: FedPer, FedRoD, FedRep; personalized methods based on global model reference: Ditto, pFedMe; personalized methods based on global prototype reference: FedProto, FedPHP; other personalized methods: Per-FedAvg, FedFomo, FedAMP, FedALA. The proposed method FedDTR is compared with the above methods in all task scenarios using the same training rounds and communication strategy.

[0074] (2) Dataset and task division. The computer vision task uses five public image classification datasets: MNIST, Fashion-MNIST (FMNIST), CIFAR-10, CIFAR-100, and Tiny-ImageNet; the natural language processing task uses two text classification datasets: AG News and Amazon Review.

[0075] (3) Model structure and learning rate setting. MNIST, FMNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet use a 4-layer convolutional neural network (CNN) as the base model; to test scalability, ResNet-18 is also used as the deep network on Tiny-ImageNet; for AG News and Amazon Review, fastText and a 3-layer multi-layer perceptron (MLP) are used respectively; the local learning rate The settings are: 0.005 for CNN and MLP structures, and 0.1 for ResNet-18 and fastText.

[0076] (4) Heterogeneous settings. To simulate a real federated learning environment, two typical non-IID data settings are designed in the experiment: Pathological Setting: The client contains only a small number of categories, the number of samples is uneven and there is no overlap; MNIST / CIFAR-10 / CIFAR-100 are only assigned 2 / 2 / 10 categories respectively; Practical Setting: Dirichlet distribution sampling is used to control the uneven distribution of data categories on the client; for category , whose samples are distributed to clients The probability of , default parameters The datasets involved strictly control the number of samples and the degree of label imbalance to test the robustness of the model to heterogeneous scenarios.

[0077] (5) Training strategy and operating environment. The total number of clients is set to 20, and all clients participate in each round by default (participation rate (J = 1.0)); each client uses 75% of its local data for training and 25% for testing; the local training batch size is 10, and the number of local iterations per round is set to 1; the total number of communication rounds is 2000; all methods are implemented based on PyTorch 1.7 and run on an Ubuntu 16.04 server with the following configurations: dual-core Intel Xeon Silver 4210 CPUs; 256GB of memory; and eight NVIDIA 2080 Ti graphics cards.

[0078] 2. Experimental Results

[0079] (1) To fully verify the adaptability and superior performance of the proposed method FedDTR in different heterogeneous environments, experiments were conducted in both pathological and practical settings. In the pathological setting, three datasets, MNIST, CIFAR-10, and CIFAR-100, were selected, and each client had only a small number of non-overlapping category labels. In the practical setting, five datasets, MNIST, CIFAR-10, CIFAR-100, Tiny-ImageNet, and AG News, were selected. Dirichlet distribution was used to control the uneven distribution of labels, simulating a more realistic federated environment.

[0080] Under the above two settings, FedDTR was systematically compared with 13 mainstream personalized federated learning methods. The experimental results show that the proposed method FedDTR achieved optimal performance in all tasks and scenarios. In image and text classification tasks, the accuracy was significantly higher than that of traditional methods and existing personalized strategies, showing good convergence, stability and generalization capabilities. In basic visual tasks such as MNIST and CIFAR-10, FedDTR outperformed traditional methods such as FedAvg and FedProx; in complex category tasks such as CIFAR-100 and Tiny-ImageNet, FedDTR significantly surpassed personalized methods such as FedPer and FedProto; in the AG News text classification task, FedDTR performed best in both accuracy and stability.

[0081] The above results fully demonstrate that the present invention has excellent cross-scenario adaptability and can effectively solve key technical problems such as performance degradation and optimization instability in existing personalized federated learning methods. The specific results are shown in the following table:

[0082] To further verify the accuracy and convergence of the proposed method on different tasks, we conducted experiments on the AmazonReview text classification task and the Fashion-MNIST (FMNIST) image classification task, using the same parameter settings. During the experiment, the changes in the test accuracy of each method during the communication round were recorded, and the convergence curve was plotted, as shown in the figure below. Figure 4 shown.

[0083] Results show that, on both the AmazonReview and FMNIST tasks, FedDTR consistently maintains the highest test accuracy throughout training, with fast convergence and minimal fluctuation, demonstrating excellent stability. In contrast, the FedProto method exhibits the slowest convergence and lowest final accuracy in both tasks, indicating its limited adaptability to non-IID data. These experimental results further demonstrate that the proposed method exhibits consistent and superior performance across diverse task types, effectively improving the robustness and convergence efficiency of personalized federated learning systems.

[0084] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0085] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0086] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0088] The embodiments described above are merely some preferred embodiments of the present invention and are not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.

Claims

1. A personalized federated learning method based on domain-invariant text representation and domain-wide global priors, characterized by: This involves introducing domain-invariant text representation and domain prior modules in the process of training local data on each client to obtain a local model. Specifically: The training process uses bimodal input of images and text collections. The input image is extracted through the image encoder participating in the training to obtain image features, and the text collection is extracted through the text encoder participating in the training and the frozen text encoder to obtain training text embedding and frozen text embedding respectively. In this process, the contrast loss of image-text alignment is used to optimize the matching between image features and text embedding features. The frozen text encoder receives global parameters sent by the server and locks the update. The training text is embedded in the domain prior module to calculate the domain prior. The domain prior and image features are added and input into the trained classifier. At the same time, the image features are input into the frozen classifier separately. The results of the trained classifier and the frozen classifier are added and output as the prediction result. This process is supervised by the task loss. The frozen text encoder and frozen classifier are respectively the global text encoder and global classifier downloaded from the server after the previous iteration step, and are locked for update in the current iteration step.

2. The personalized federated learning method based on domain-invariant text representation and domain-wide prior according to claim 1 is characterized in that: The contrast loss of the image-text alignment consists of two symmetrical log-negative probability terms, which represent the negative logarithms of the normalized similarity probabilities based on the frozen text embedding and the training text embedding, respectively. The normalized similarity probability based on the frozen text embedding is the ratio of the two parts, where the numerator of the two parts is: image feature Frozen text embedding with similar tags Negative exponential similarity between , the denominator is The sum of the negative sample sum terms, the negative sample sum term represents the image feature Freeze text embedding with all heterogeneous The normalized similarity probability based on the training text embedding adopts the same structure, but with different parameters, and uses image features that do not participate in gradient updates. , text embedding uses trainable text embedding and , whose numerator is: image features that do not participate in the update Training text embeddings with similar labels Negative exponential similarity , the denominator is and the sum of the corresponding negative sample summation items, where the corresponding negative sample summation items are image features that do not participate in the update With all heterogeneous training text embeddings The sum of similarities.

3. The personalized federated learning method based on domain-invariant text representation and domain-wide prior according to claim 1 is characterized in that: The domain prior module uses the training text embedding and the data distribution of each classification in the training set statistics to calculate the domain prior, specifically: Set the global tag Each label in the Each client generates a corresponding set of text embedding vectors: ,in Indicates the Client corresponds to Text embedding representation of the class; By combining the proportion of samples of each category in the local data of the client in the number of local data categories, the text embedding representation in the text embedding vector set is weighted averaged to obtain the domain prior of the client.

4. The personalized federated learning method based on domain-invariant text representation and domain-wide prior according to claim 3 is characterized in that: The domain prior is: ; in Is the indicator function, used to judge the sample Tags Belongs to category , is the number of categories present in the local data.

5. The personalized federated learning method based on domain-invariant text representation and domain-wide prior according to claim 1 is characterized in that: The task loss is calculated by calculating the one-hot encoding vector of the true label The cross entropy between the predicted output and the model is used to measure the degree of match between the two. The predicted output of the model is obtained by the sum of two parts: the first part is the classifier result during training. , the second part is the result of freezing the classifier , and the outputs of these classifiers are normalized by the softmax function to generate the final probability distribution.

6. The personalized federated learning method based on domain-invariant text representation and domain-wide prior according to claim 1, characterized in that: The goal of local training for each client is to minimize the sum of the contrast loss and task loss of image-text alignment.

7. The personalized federated learning method based on domain-invariant text representation and domain-wide prior according to claim 1 is characterized in that: After the current iteration step is completed, all clients upload the trained local models to the server, including the text encoder, image encoder and trained classifier involved in the training. The server calculates the corresponding average models respectively to obtain the global model, including the global text encoder, global image encoder and global classifier, and transmits it back to each client. The client initializes the text encoder and frozen text encoder involved in the training as the global text encoder, the image encoder as the global image encoder, and the classifier involved in the training and the frozen classifier as the global image classifier. During the local training process, the parameters of the frozen text encoder and frozen classifier are locked and not updated, and the loop iteration is performed.

8. A personalized federated learning system based on domain-invariant text representation and domain-wide global priors, characterized by: Used to implement the method according to any one of claims 1 to 7.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing computer-executable instructions, wherein the instructions are used to implement the method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Image classification model training method and system based on enhanced federal domain generalization

    CN115731424A

  • Partial prompt learning-based lifelong target re-identification method

    CN118864825A

  • Federal domain generalization method based on trainable prototype

    CN120032203A

  • Personalized federal learning method for visual language model, terminal and medium

    CN120069007A

  • Modal heterogeneous federated learning privacy protection method based on cross-modal prototype

    CN120180463A