Personalized Federated Learning Method, Terminal and Medium for Vision-Language Model
By constructing client graphs and iterative clustering, combining LoRA modules and parameter regularization, the Byzantine attack and communication overhead problems in federated learning systems are solved, and efficient training and robustness improvement of visual language models in heterogeneous environments are achieved.
Patent Information
- Application Number
- CN202510525367.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Federated learning systems face the problems of Byzantine attacks and communication overhead when training visual language models, especially in heterogeneous client environments, traditional methods are difficult to effectively resist malicious attacks and reduce communication overhead.
Using a personalized federated learning method, by constructing a client graph, iterative clustering is performed to minimize structural entropy, the visual language model is fine-tuned using the LoRA module, and parameter regularization is introduced in the local update stage, combining low-rank adaptation and efficient parameter updates to form a global model.
Effectively resist Byzantine attacks, reduce communication overhead, improve the performance and robustness of the model in a heterogeneous environment, and realize efficient visual language model training.
Smart Images

Figure CN120069007B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and specifically to a personalized federated learning method, terminal, and medium for visual language models. Background Art
[0002] Visual language models have extraordinary capabilities when performing various image classification tasks. The performance of the model depends to a large extent on the training with a large amount of data, which is usually stored on different devices of different organizations. Some privacy data protection regulations have become a bottleneck for further model development. To solve this problem, researchers have introduced federated learning (FL) during the model training process. Federated learning is a distributed deep learning framework that allows multiple devices (called clients in federated learning) to jointly train a central federated learner. Following the principle of "data does not move, model moves", local data is used to jointly train a centralized federated model. This method makes full use of the computing resources of the clients and ensures a certain degree of privacy of the original data during the training process.
[0003] Although federated learning effectively protects the privacy of the original data during the training process, its practice of relying on a large amount of privacy-sensitive data for model development will inevitably attract malicious attackers. For example, federated learning faces the challenge of Byzantine attacks. Malicious participants (called Byzantine clients) may upload arbitrary local updates, damaging the performance of the global model. Therefore, it is urgent to solve this problem.
[0004] In addition, the federated learning system faces communication overhead. Local clients will encounter computing and storage limitations. Distributed clients and services must frequently exchange model gradients or parameters, usually on the scale of millions or billions. The communication bandwidth of the FL system is limited, which will cause serious delays during these frequent exchanges. Local clients often lack the computing resources to fine-tune the entire model, and storing different instances for various tasks will further consume memory resources. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides a personalized federated learning method, terminal, and medium for visual language models, which can effectively resist Byzantine attacks while reducing communication overhead.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The present invention discloses a personalized federated learning method for visual language models, including the following steps:
[0008] S1. Model the clients in the federated learning system as nodes in an undirected graph, and calculate the edge weights between nodes by the cosine similarity of client parameters, thereby constructing a multi-modal client graph;
[0009] S2. Divide the client graph into clusters and perform iterative clustering to minimize the structural entropy of the client graph; wherein, the structural entropy is jointly determined by intra-cluster homogeneity and inter-cluster separation;
[0010] S3. Fine-tune the vision-language model for the clients within each cluster through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization during the local update stage of the clients to obtain the model upload parameters of the clients;
[0011] S4. According to the results of iterative clustering, the servers within each cluster receive the model upload parameters of each client within the cluster, aggregate them to form a global model, and distribute it to the clients;
[0012] Among them, during the clustering process in S2, malicious clients and normal clients are fully assigned to different clusters by globally minimizing the structural entropy; after clustering, the clusters with the number of clients within the cluster less than the preset threshold are directly discarded, so that malicious clients become ineffective.
[0013] As a further improvement of the above solution, in step S2, the calculation formula for the structural entropy of the client graph is:
[0014] ;
[0015] In the formula, is the structural entropy of the client graph in the t th round; is the total number of clusters into which the client graph is divided in each round;
[0016] L is the total number of client nodes in each cluster
[0017] ; is the local degree of each node
[0018] in , defined as the sum of the edge weights connecting i to other nodes in the same cluster; i ;
[0019] is the cluster quantity, used to quantify the total connection strength within , ;
[0020] is the global graph quantity, used to quantify The cumulative strength of all connections in ;
[0021] is the trimmed edge, defined as the sum of the connection weights connecting and other cluster nodes.
[0022] As a further improvement to the above solution, in step S2, the greedy algorithm is used for iterative clustering;
[0023] Among them, in the initial state, each client forms a separate cluster, that is L = N ; L is the total number of clusters into which the client graph is divided in each round, N is the total number of clients in the federated learning system;
[0024] Iteratively merge the cluster pair that can minimize the structural entropy, that is, obtain the maximum entropy reduction , is the overall structural entropy of the client graph before merging, is the overall structural entropy of the client graph after merging;
[0025] When the entropy reduction is lower than the preset threshold, stop.
[0026] As a further improvement to the above solution, in step S3, low-rank adaptation is performed through the global model of each client within the cluster to fine-tune the vision-language model; among them, in the t round, the global model i of client is calculated as:
[0027] ;
[0028] In the formula, represents the global model of client t in the previous round, that is i -1; represents the original parameters of client i ; represents low-rank adaptation; represents the deployed LoRA module;
[0029] For each cluster c after clustering, its global model is calculated as:
[0030] ;
[0031] In the formula, i , j represent the cluster cAny two clients within are nodes;
[0032] is a cluster c clients within i parameters;
[0033] is a cluster c clients within j parameters.
[0034] As a further improvement to the above solution, in the client local update phase, each client k based on its own global model and local data minimizes the following objective function:
[0035] ;
[0036] In the formula, is the local objective function of client k ; is the local loss function of client k used to reflect the matching degree between the client local model w and the client local data; is an L2 regularization term used to penalize the large deviation between the local model and the global model ; is the regularization coefficient used to control the penalty strength for deviating from the global model to prevent overfitting of local data .
[0037] As a further improvement to the above solution, in step S1, the client graph is represented as ; where node represents the client, and the weight calculation formula for the edge W between nodes is , , , N is the total number of clients in the federated learning system; and respectively represent the parameters of client i and client j .
[0038] As a further improvement to the above solution, in step S3, the LoRa module is selectively deployed in the text encoder of the vision - language model for fine - tuning.
[0039] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the personalized federated learning method for the vision language model as described above are implemented.
[0040] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the personalized federated learning method for the vision language model as described above are implemented.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] 1. The personalized federated learning method for the vision language model disclosed by the present invention combines the LoRA technology in the parameter efficient fine-tuning method with personalized joint learning. By using LoRA on the vision language model, only a small part of the model parameters needs to be updated to fine-tune it, while most of the pre-trained weights are retained, thereby reducing communication overhead.
[0043] In a system with inherent partitions or large differences in data distribution among clients, the traditional client-server joint learning architecture is not the best choice. Instead, a multi-model method of training FL models for each client group is more suitable, which requires the server to adopt an aggregation strategy that takes into account client heterogeneity and individual characteristics. The present invention proposes a personalized global aggregation method, which can construct a client graph on the server and minimize its two-dimensional structure entropy. During the global aggregation process, the present invention weights the aggregation according to the similarity between clients. This similarity-based weighting strategy helps to mitigate the impact of data heterogeneity on the global model, thereby obtaining the best aggregation strategy and reducing global data heterogeneity at the same time. Experiments show that the present invention can effectively resist Byzantine attacks in joint learning.
[0044] 2. In the local update stage of the client, the present invention introduces a client parameter regularization method. Each client updates the model according to the global model and its own data distribution, and punishes the updates that deviate significantly from the global model to maintain consistency.
[0045] 3. The computer terminal and computer-readable storage medium disclosed by the present invention can produce the same beneficial effects as the above method by applying the above method, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of the personalized federated learning method for the vision language model in Embodiment 1 of the present invention.
[0047] Figure 2Schematic diagram of the deployment of the LoRA module in Embodiment 1 of the present invention.
[0048] Figure 3 Flowchart of minimizing structural entropy in Embodiment 1 of the present invention.
[0049] Figure 4 Visualization diagram of the Dirichlet distribution (α = 1) on the non-independent and identically distributed Flowers dataset in Embodiment 1 of the present invention.
[0050] Figure 5 Visualization diagram of the pathological distribution on the non-independent and identically distributed CIFAR-10 dataset in Embodiment 1 of the present invention.
[0051] Figure 6 Transformer architecture diagram with LoRA-enhanced text encoder in Embodiment 1 of the present invention.
[0052] Figure 7 Comparison chart of the attack experiment results of the method of the present invention and the FedAvg method in Embodiment 1 of the present invention.
[0053] Figure 8 Accuracy graph of various datasets affected by different α in Embodiment 1 of the present invention.
[0054] Figure 9 Schematic diagram of the computer terminal structure in Embodiment 2 of the present invention. Detailed implementation manners
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment 1
[0057] This embodiment provides a personalized federated learning method for a vision-language model. This method is applied to the training of the vision-language model, combining the low-rank adaptation technology in the parameter-efficient fine-tuning method with personalized federated learning.
[0058] Please refer to Figure 1 , and the federated learning method includes the following steps, namely S1 to S4.
[0059] S1. Model the clients in the federated learning system as nodes in an undirected graph, and calculate the edge weights between the nodes by the cosine similarity of the client parameters, so as to construct a multi-modal client graph.
[0060] In step S1, the client graph is represented as ; among which, the node represents the client, and the weight calculation formula of the edge W between nodes is , , , N , where and respectively represent the parameters of client i and client j . The client parameters only include the parameters of the text encoder, and the parameters can be regarded as unimodal data with multimodal associations.
[0061] S2. Divide the client graph into clusters and perform iterative clustering to minimize the structural entropy of the client graph, aiming to minimize the structural disorder of the client graph. Among them, the structural entropy is jointly determined by intra-cluster homogeneity and inter-cluster separation.
[0062] In step S2, the calculation formula of the structural entropy of the client graph is:
[0063] ;
[0064] In the formula, is the structural entropy of the client graph in the t -th round ;
[0065] L is the total number of clusters into which the client graph is divided in each round;
[0066] is the total number of client nodes in each cluster ;
[0067] is the local degree of each node i in i , defined as the sum of the edge weights connecting
[0068] to other nodes in the same cluster; is the cluster quantity, used to quantify ;
[0069] is the global graph quantity, used to quantify the cumulative strength of all connections in ;
[0070] In addition to intra-cluster dynamics, each cluster is also connected by cut edges Interact with other clusters to define the cut edges Is for connection And the sum of the connection weights to other cluster nodes. These variables together define the structural entropy, where the local degree and the cluster volume determine the homogeneity within the cluster, while the cut edges and the global volume affect the separation between clusters. These metrics are optimized by iteratively merging clusters, which gradually reduces the uncertainty in the hierarchical structure of the graph.
[0071] In step S2, a greedy algorithm is used for iterative clustering;
[0072] Wherein, in the initial state, each client forms an independent cluster, that is L = N ; L Is the total number of clusters into which the client graph is partitioned in each round, N Is the total number of clients in the federated learning system;
[0073] Iteratively merge the cluster pair that can minimize the structural entropy, that is, obtain the largest entropy reduction , Is the structural entropy of the overall client graph before merging, Is the structural entropy of the overall client graph after merging;
[0074] When the entropy reduction Is lower than the preset threshold, stop. In this embodiment, the preset threshold is set to 0.05 according to experience.
[0075] In the scenario of fine-tuning visual language models, traditional federated learning frameworks struggle to handle client heterogeneity. These systems enforce strict global model consistency and cannot adapt to clients with diverse data distributions or specific task requirements. For systems with inherent partitions or significant differences in data distribution among clients, the traditional client-server federated learning architecture for training a shared global model is not the optimal choice. A multi-model approach that trains a federated learning model for each homogeneous group of clients is more appropriate, and the server requires an aggregation strategy that can account for client heterogeneity and their individual characteristics. Clustering similar clients can improve the collaborative training effect. Therefore, a multi-model approach is more effective, that is, training a federated learning model for each homogeneous group of clients. Clustering Federated Learning (CFL) partitions clients into several clusters, and each cluster is represented by a model centroid.
[0076] Although clustering-based federated learning methods provide a certain solution through client grouping, they ignore the subtle relationships between clients that are critical for effective knowledge transfer. Static client grouping further ignores the evolution and underlying relationships of model behavior and fails to detect the development of complementary feature extractors between clients with different raw data during training. Traditional k-means clustering cannot dynamically adjust the number of clusters and relies only on mean parameter vector information, thus ignoring multimodal interaction relationships. Therefore, the present invention addresses these limitations through structural entropy-guided aggregation, a new personalized federated learning framework for visual language models that can optimize dynamic client clustering.
[0077] S3. Fine-tune the visual language model for each client in each cluster through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization during the client local update phase to obtain the client model upload parameters (i.e. the updated parameters in the previous article). and ).
[0078] See also Figure 2 , clients are clustered (Cluster 1 to Cluster M) based on model or task similarity. Each cluster shares a common LoRA adapter. Each client uses frozen pre-trained image / text encodings paired with a trainable LoRA adapter.
[0079] In this paper, CLIP (Contrastive Language-Image Pre-Training) is a multimodal pre-training neural network consisting of two core components: an image encoder I(·) and a text encoder T(·). These encoders convert images and text into vector representations and project them into a shared embedding space. The training objective of CLIP is to maximize the cosine similarity between the embeddings of matching image-text pairs while minimizing the similarity of mismatched pairs. Parameter regularization is employed during local training to align client-specific LoRA updates with the stability of the global model.
[0080] For inference, CLIP computes the normalized embedding of the image f i = I ( x i ) and text prompts t k = T ( c k ), such as "a picture of [category]". x The model will calculate I ( x ) and all candidate text embeddings The cosine similarity between them, scaled by a temperature parameter, the predicted probability of the category i is as follows: y i is as follows:
[0081] ;
[0082] In the formula, T represents the text feature vector; T ( t i ) is the text feature vector of the i th category; K represents the total number of categories; τ is the temperature parameter, used to scale the value of the cosine similarity.
[0083] By contrastive learning to align multimodal representations, CLIP achieves robust zero-shot generalization without retraining for specific tasks.
[0084] The LoRA model incrementally updates the pre-trained weights through the product of two small matrices A and B, based on the intrinsic rank of the downstream task. In this embodiment, CLIP is used as the backbone vision-language model, and its text encoder is fine-tuned through low-rank adaptation (LoRA). For the Q, K, V, and O matrices of the text encoder ( W q , W k , W v , W o ∈R d×d ), this embodiment injects trainable low-rank matrices A and B to approximate the weight update, where R is the set of real numbers, the dimension of the multi-head attention matrix in the Transformer is d×d, and the LoRa module can be selectively deployed on the Transformer to optimize the transmission efficiency. The specific deployment method can refer to the existing research: Maxime Zanella and Ismail Ben Ayed. 2024. Low-rank few-shot adaptation of vision-language models. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 1593–1603.
[0085] The above mechanism establishes a robust framework for the fine-tuning of vision-language models, which is formalized as follows:
[0086] ;
[0087] Where the scaling factor α regulates the contribution of the low-rank update . Here, B ∈ R d×r , A ∈ R r×k . The dimension of the original weight matrix of the local model is d × k. The rank r is a hyperparameter used to control the inner dimension of matrices A and B. The dimension of matrix A is d × r, and the dimension of matrix B is r × k. The size of the trainable parameters is reduced from d × k to r × (d + k). r is much smaller than the smaller value of d and k, thus retaining the original weight W while only adjusting A and B, achieving efficient parameter updates.
[0088] In LoRA fine-tuning, the scaling factor α and the rank r act synergistically to regulate the adjustment of the adapter to the pre-trained model weights. The ratio α / r scales the amount of weight update , where α globally amplifies the influence of the adapter, while r limits the dimension of the low-rank factorization, thus controlling the number of trainable parameters. Adjusting the ratio α / r can balance between retaining the knowledge of the original model and adapting to specific tasks. For example, increasing α can enhance the update amplitude of the adapter, which is beneficial for complex tasks that require significant weight adjustments; while decreasing α can reduce the update amount, thus avoiding overfitting in the case of less data. The key is that this reduction in parameters does not affect performance because the low-rank matrix still has the ability to capture the key changes in the input data.
[0089] In step S3, low-rank adaptation is performed through the global model of each client within the cluster for visual language model fine-tuning; where, in the t -th round, the global model i of client is calculated as:
[0090] ;
[0091] Where represents the global model of client t in the previous round, i.e., i -1; represents the original parameters of client i ; represents the low-rank adaptation; represents the deployed LoRA module;
[0092] For each cluster after clustering ends c , its global model is calculated by the formula:
[0093] ;
[0094] In the formula, i , j represent any two clients (i.e., nodes) within the cluster c ;
[0095] is the parameter of the client c within the cluster i ;
[0096] is the parameter of the client c within the cluster j .
[0097] In the local update stage of the client, each client k minimizes the following objective function based on its own global model and local data :
[0098] ;
[0099] In the formula, is the local objective function of the client k ; is the local loss function of the client k , which is used to reflect the matching degree between the local model w of the client and the local data of the client; is an L2 regularization term, which is used to punish the large deviation between the local model and the global model ; is the regularization coefficient (hyperparameter), which is used to control the penalty strength for deviating from the global model . A larger will force it to be more strictly consistent with the global model, ensuring that the client update is consistent with the global model and preventing overfitting of the local data . By punishing the large deviation, it can stabilize the federated learning process and reduce client drift. Otherwise, the local model may deviate significantly from the global model
[0100] It's important to note that the cluster's global model is the model sent to the client in the next round. The complete federated learning process is as follows: 1. The client receives the cluster's global model, trains it using local data, and then uses parameter regularization and LoRa to generate updated client parameters. The client then uploads its parameters, and the server aggregates them to form the cluster's global model. This completes the round.
[0101] like Figure 3 As shown, the federated learning framework of the present invention facilitates semantically based knowledge transfer while preserving client-specific patterns, overcoming the limitations of traditional federated learning and static collaborative federated learning methods. The structural entropy formula, an information-theoretic regularization term, automatically balances the model's personalization and generalization capabilities throughout the training process.
[0102] S4. Based on the iterative clustering results of S2, the server in each cluster receives the model parameters uploaded by each client in the cluster, aggregates them into a global model, and sends it to the clients. During the clustering process, the global structural entropy is minimized to fully separate malicious clients from normal clients into different clusters. After the clustering is completed, clusters with fewer than a preset threshold of clients are directly discarded and do not participate in the global aggregation, thereby rendering the malicious clients ineffective.
[0103] In this embodiment, the effectiveness of the above method is also demonstrated through experiments, and the specific contents of the experiments are as follows:
[0104] 1 Experimental Setup
[0105] 1.1 Dataset
[0106] This example uses six datasets: Digit5, Flowers102, Caltech101, Descriptive Texture Dataset (DTD), Oxford-IIIT Pet, and CIFAR-10. Flowers102, Caltech101, DTD, and Oxford-IIIT Pet are datasets for label-skewed scenes, while Digit5 is a dataset for feature-shifted scenes.
[0107] Digit5 is a multi-domain dataset for digit classification, consisting of five domains: MNIST, SVHN, USPS, MNIST-M, and SynthDigits. Each domain contains images of handwritten or synthesized digits from 0 to 9, with different image styles and distributions. For example, MNIST contains grayscale handwritten digits, while SVHN contains color house numbers from street view images. The feature offset between different domains makes it challenging to generalize the model between different domains. This example randomly selects 25,000 training samples and 9,000 test samples in each domain.
[0108] Flowers102 is a fine-grained image classification dataset containing 102 flower categories. This dataset has approximately 8,189 images, with each category containing between 40 and 258 images. These images were taken under different conditions, including different lighting, angles, and backgrounds. Label bias stems from the uneven distribution of images across categories, making it suitable for federated learning scenarios where client data may be imbalanced.
[0109] Caltech101 is a classic image classification dataset containing 101 object categories, plus an additional background category. Each category has approximately 40 to 800 images, for a total of approximately 9,000 images. This dataset covers a wide variety of objects, such as animals, vehicles, and everyday items. Due to the unequal number of images in each category, label bias is quite evident. The Describable Textures Dataset (DTD) is a dataset for label bias scenarios focused on texture recognition, containing 47 texture categories, such as stripes, squares, and checkerboards. Each category has 120 images, with image sizes ranging from 300x300 to 640x640.
[0110] Oxford-IIIT Pet is a pet image classification dataset containing 37 pet categories, for a total of 7,349 images. The number of images in each category varies, from dozens to hundreds. This dataset contains images of cats and dogs, and annotates both species and their breeds.
[0111] CIFAR-10 is a standard benchmark dataset in computer vision research, containing 60,000 color images covering 10 object categories. As a standard evaluation benchmark for image classification models, it is widely used in research on feature learning. Each 32×32 pixel image belongs to one of the following categories: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck.
[0112] 1.2 Data Distribution
[0113] In the experimental setup section, to simulate a heterogeneous environment, this embodiment uses the Dirichlet distribution and the pathological distribution. Each of these distributions has unique characteristics that can effectively simulate the non-independent and identically distributed (non-IID) nature of real-world data.
[0114] The Dirichlet Distribution is a commonly used multivariate probability distribution for generating probability vectors. It is a multi-dimensional generalization of the Beta distribution and is suitable for modeling categorical distributions. In federated learning, the Dirichlet Distribution is used to simulate the class distribution of client data. By adjusting the parameter α of the Dirichlet Distribution, the proportion of different class data that each client draws from the global dataset can be controlled, thus simulating different degrees of data heterogeneity. In this embodiment, α is set to 0.1. Figure 4 Shows the resulting label statistical heterogeneity on the Flowers dataset. This partitioning is non-independent and identically distributed, which is evident from the absence of certain classes in each client.
[0115] This embodiment uses the Pathological Distribution to simulate extreme or specific scenarios. This method is used to test the robustness and adaptability of the algorithm in the case of extremely unbalanced data distribution. It simulates the situation of uneven data distribution in the real world, such as the differences in patient data in hospitals in different regions, or different data patterns on user devices. For example, if there are a total of ten classes, Client 1 only has classes 1, 2, and 3, Client 2 only has classes 4, 5, and 6, and Client 3 only has classes 7, 8, 9, and 0. This embodiment divides the dataset into three non-overlapping class subsets, and each client is only assigned a part of them. By restricting each client to only use a specific subset of the data, the pathological distribution ensures that there are significant differences in data characteristics among clients. Figure 5 Shows the resulting label distribution heterogeneity on CIFAR-10.
[0116] 1.3 Baselines
[0117] By applying LoRA to more network layers and vectors, the performance of the model can be further improved. However, this improvement comes at the cost of increased computational burden, because adding more weight matrices introduces additional parameters. LoRA provides a method to balance performance and parameter efficiency, enabling effective fine-tuning of the model while maintaining its original size. Therefore, in this embodiment, LoRA is only applied to the QKVO vectors of the 12-layer transformers in the text encoder, while keeping the image encoder unchanged. The specific structure is as Figure 6 shown. For a fair and comprehensive evaluation, this embodiment selects FedCLIP as the baseline. FedCLIP inserts an adapter on top of CLIP to identify the part of the features suitable for a specific task.
[0118] 1.4 Implementation Details
[0119] In the experiment, this embodiment adopted a dynamic learning rate, set the number of local training rounds for each client to 50 rounds, with 10 clients participating in each round of sampling, and a batch size of 128 to balance computational efficiency and model stability. The clients can have different numbers of training samples. For the Adam optimizer, the learning rate was set to 5e-5, the momentum was set to 1e-6, and the weight decay was set to 0.2 to promote stable convergence and reduce overfitting. Additionally, when using the LoRA method, this embodiment set the rank r to 2 and the scaling factor α to 32 to achieve fast adaptation and generalization of the model using low-rank approximation. These meticulous parameter selections aim to ensure the accuracy of the experimental results and optimize the model performance.
[0120] 1.5 Privacy issues
[0121] Byzantine attacks pose a significant threat to federated learning systems because they not only disrupt the convergence of the model but also reduce the accuracy and usability of the model. In the non-independent and identically distributed (non-IID) data setting, there are significant differences in the local data distributions of the clients, which makes it more difficult for federated learning systems to distinguish between benign and malicious model updates. Attackers can take advantage of this by subtly manipulating the updates to conceal their intentions, thereby reducing the effectiveness of traditional robust aggregation rules in identifying and filtering malicious updates. The sign-flipping (SF) attack is a simple Byzantine attack method where malicious clients disrupt the training of the global model by reversing the sign of their model updates.
[0122] This SF attack disrupts federated learning by reversing the direction of the gradient update, turning gradient descent into gradient ascent to maximize the model loss and prevent convergence. Different from poisoning attacks that modify training data or model parameters, the SF attack directly exploits the federated framework to flip the gradient sign during local training. This attack has been proven effective in both FedSGD (Federated Stochastic Gradient Descent algorithm) and FedAvg (Federated Averaging algorithm): in FedSGD, the attacker submits reversed gradients. Different from the cumulative damage caused by sign flipping in the repeated local iterations in FedAvg, the simplicity of the SF attack stems from avoiding complex data or model operations, making it particularly suitable for federated settings where clients compute updates independently. The SF attack has extremely low requirements for the capabilities of the adversary. The attacker only needs to control one or more compromised clients without colluding with other clients or obtaining external updates. The SF attack requires very little knowledge of the adversary. The attacker does not need to know the global model architecture, the data distributions of other clients, or the aggregation rules, enabling the adoption of a "black-box" strategy.
[0123] Although this method is simple, it can effectively disrupt the training process, especially when the federated learning system has weak detection and defense mechanisms against malicious updates. Implementing the sign-flipping attack does not require in-depth knowledge of the global model or data, so it is a low-cost and easy-to-execute attack strategy.
[0124] 2 Experimental Results
[0125] The experimental results were obtained under the above experimental settings. This embodiment uses the following research questions as guidelines for the experiment:
[0126] Research Question 1: Compared with the FedCLIP algorithm (Lu W, Hu X, Wang J, et al. Fedclip: Fast generalization and personalization for clip in federated learning[J]. arXiv preprint arXiv:2302.13485, 2023.), can the method of the present invention achieve better performance on multiple datasets under Dirichlet distribution-based data heterogeneity?
[0127] Research Question 2: How does the method of the present invention perform compared to FedCLIP in an extreme pathological data distribution scenario?
[0128] Research Question 3: How does the scaling factor α in LoRA affect the performance of the method of the present invention on different datasets?
[0129] Research Question 4: Is the method of the present invention more robust to Byzantine attacks compared to the classical FedAvg algorithm?
[0130] 2.1 Research Question 1: Performance Comparison
[0131] The experiments in this embodiment compared the proposed method with FedCLIP on six benchmark datasets: CIFAR-10, Caltech101, Flowers, DTD, Oxford-IIIT Pet, and Digit5. This embodiment focused on the federated fine-tuning of vision-language models. Table 1 summarizes the experimental results, showing that the present invention has always been significantly better than FedCLIP, with an average accuracy improvement of 9.04%. The present invention exceeded FedCLIP on all datasets, especially showing significant improvements on Caltech101, Flowers, and DTD. These results indicate that integrating LoRA into the QKVO matrix of the text encoder is effective, enabling the model to adapt to task-specific language patterns while retaining the pre-trained visual representations of CLIP. Even on relatively simple tasks like Digit5, the present invention achieved a 9.67% improvement, highlighting its robustness in handling cross-domain challenges.
[0132] Table 1: Comparison of experimental results between FedCLIP and the method of the present invention under Dirichlet distribution data
[0133] ;
[0134] The above method of FedCLIP trains a single adapter on top of CLIP, which limits its ability to adapt to heterogeneous tasks, as reflected in the suboptimal performance on complex datasets such as DTD and Flowers. In contrast, the method of the present invention performs low-rank updates on the core attention mechanism (QKVO) of the text encoder, enabling targeted adjustments without adding excessive parameters. This strategy reduces the computational overhead while retaining the generalization ability of the original model. By only deploying low-rank adaptation (LoRA) in the text encoder, the method of the present invention achieves a strategic balance between adaptation efficiency and computational cost.
[0135] There is a significant performance gap between the method of the present invention and FedCLIP, which reveals the limitations of the simple adapter-based method in the federated learning environment. The global adapter of FedCLIP operates as a monolithic component independent of the intrinsic architecture of CLIP and is difficult to adapt to different client-specific tasks. In contrast, the method of the present invention distributes the adaptation process to the attention mechanism of the text encoder through low-rank adaptation, thus achieving localized and efficient updates. By embedding LoRA into the QKVO matrix of the text encoder, the method of the present invention achieves a principled balance between task-specific adaptation and retention of pre-trained knowledge. This method not only improves the state-of-the-art performance but also provides a scalable framework for existing federated learning deployments.
[0136] 2.2 Research Question 2: Pathological Distribution Analysis
[0137] In this embodiment, the method of the present invention and FedCLIP were evaluated under pathological and Dirichlet data distributions. Tables 2 and 1 show the comparison results. In the pathological setting (Table 2), the clients accessed mutually exclusive subsets of classes, and the method of the present invention performed excellently on all datasets. On Caltech101, the method of the present invention was 9.31% higher than FedCLIP, 27.89% higher on Flowers, 14.96% higher on DTD, 7.08% higher on Oxford-IIIT Pet, with an average improvement of 14.81%. The performance gap on the Flowers dataset was the largest: the accuracy of FedCLIP was only 68.27%, while the method of the present invention reached 96.16%. These results indicate that in the case of extreme data scarcity and class isolation, traditional adapter methods like FedCLIP cannot retain transferable knowledge. The present invention achieves parameter-efficient adaptation by strategically integrating LoRA into the text encoder QKVO vectors of CLIP while retaining important pre-trained features.
[0138] Table 2: Comparison of experimental results between FedCLIP and the method of the present invention under pathological distribution data
[0139] ;
[0140] In the case of Dirichlet distribution data with partial class overlap (Table 1), the method of the present invention still maintained an advantage, but the magnitude of the advantage decreased. The advantage in average accuracy decreased from 14.81% in the pathological setting to 9.04% in the Dirichlet scenario. This comparison shows that although benefiting from class overlap in the Dirichlet setting, the shallow adapter architecture of FedCLIP is still fundamentally limited in cross-client knowledge integration and cannot compare with the deep parameter-efficient strategy of the present invention.
[0141] 2.3 Research Question 3: Robustness Analysis
[0142] In this embodiment, the personalized federated learning framework of the present invention was compared and evaluated with the baseline method in scenarios with and without attacks. Figure 7 The results of the method of the present invention in resisting sign-flipping attacks on different datasets are shown.
[0143] In Dirichlet distribution data (30% malicious clients), during the collaborative attack, the accuracy of traditional FedAvg drops significantly, dropping to 88.79% on Caltech101, 40.57% on DTD, 65.70% on Flowers, and 85.16% on Oxford-IIIT Pet, on average 25.3% lower than the method of the present invention. The greedy clustering strategy of the present invention isolates malicious clients by splitting the network into two subgraphs, thus effectively excluding the smaller attacker cluster (30% of the nodes) from global aggregation. This defense mechanism improves the DTD texture classification accuracy by 66.6% compared to the compromised FedAvg (67.60% vs. 40.57%), and is particularly effective for visual language tasks.
[0144] Comparison with the non-attack scenario (Table 1) shows that the method of the present invention causes only a minimal accuracy loss while maintaining security. Between the benign and attack conditions, the average accuracy drops only 0.83% (from 87.98% to 87.15%), 9.04% higher than the non-attack baseline of FedCLIP, while maintaining robustness. There are three key mechanisms that enable the system to have this anti-attack ability: ① The graph-based clustering of the present invention isolates malicious clients through two-dimensional similarity embedding minimization; ② The framework separates clients with different update patterns (e.g., gradient sign flipping) to prevent model contamination; ③ Different from the simple averaging of FedAvg, two-dimensional similarity embedding minimization can quantify the credibility of clients without prior knowledge of the attack. The system dynamically partitions clients according to the update similarity pattern, actively adapting to changing threats such as sign flipping and other gradient manipulation attacks. This method achieves an attack detection accuracy of 92.4% (from excluding 30% of malicious clients), and the performance overhead is less than 1%.
[0145] 2.4 Research Question 4: Ablation Study
[0146] Please refer to Figure 8 , this embodiment analyzes how different scaling factors (α) affect the performance of LoRA-based fine-tuning. The research covers four datasets: Caltech101, Flowers, DTD, and Oxford-IIIT Pet. The ablation study shows that α has a significant impact on the accuracy of the model, indicating that it is necessary to balance adaptability and generalization ability by carefully selecting α during the fine-tuning process.
[0147] The optimal choice of α depends on the complexity of the dataset. For the Caltech101 dataset, the highest accuracy can be achieved when α = 32, while smaller α values (e.g., α = 4) will reduce performance. It is worth noting that overly large α values (such as α = 64) will also reduce accuracy, with diminishing returns. Similarly, for the Flowers dataset, peak accuracy is reached at α = 32, and accuracy drops significantly at α = 8 and α = 4, indicating that smaller α values may not be able to adapt to the specific details of the task, while larger values risk overfitting or making the model unstable. In contrast, the DTD dataset performs well at α = 64, meaning that tasks with higher intrinsic variability (such as texture recognition in DTD) may benefit from stronger updates to capture diverse patterns. For Oxford-IIIT Pet, α = 32 performs better than other values, but even a moderate reduction (e.g., α = 16) will significantly affect accuracy, highlighting the dependence of this dataset on precise parameter tuning.
[0148] Larger α values (such as 64) can enhance adaptability to complex tasks such as DTD, but there is a risk of overfitting on datasets with lower diversity such as Oxford-IIIT Pet. Smaller α values (such as 4 - 8) can preserve pre-trained features but cannot adapt to the specific nuances of tasks such as Flowers and Caltech101. However, α = 32 can achieve near-optimal results on most datasets, striking a balance between task-specific adaptability and the generalization ability of the pre-trained model.
[0149] The empirical results of this embodiment show that choosing the right α is crucial for LoRA-based fine-tuning. A moderate α = 32 can maintain a balance between adaptability and generalization ability, achieving robust performance on different datasets. This finding highlights the role of α as an adjustable lever, helping to achieve coordination between model stability and task-specific learning, especially in a federated learning environment where heterogeneous data distributions require flexible and controllable updates.
[0150] Example 2
[0151] This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the personalized federated learning method for a vision-language model as described in Example 1.
[0152] As Figure 9As shown in the figure, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to the at least one processor 101. In this embodiment, the specific connection medium between the processor 101 and the memory 102 is not limited. Figure 9 In the figure, it is taken as an example that the processor 101 and the memory 102 are connected through a bus 100. The bus 100 is Figure 9 shown as a thick line in the figure. The connection manners between other components are only for illustrative purposes and are not to be taken as limitations. The bus 100 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 9 it is shown as only one thick line in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 101 can also be called a controller, and there is no limitation on the name.
[0153] In this embodiment, the memory 102 stores instructions executable by the at least one processor 101. By executing the instructions stored in the memory 102, the at least one processor 101 can execute the foregoing method.
[0154] Among them, the processor 101 is the control center of the device, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 102 and calling the data stored in the memory 102, various functions of the device and process data, so as to monitor the device as a whole.
[0155] In a possible design, the processor 101 may include one or more processing units. The processor 101 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the foregoing modem processor may not be integrated into the processor 101. In some embodiments, the processor 101 and the memory 102 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.
[0156] The processor 101 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the personalized federated learning method for the vision language model disclosed in conjunction with Embodiment 1 can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor 101.
[0157] The memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 102 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, and so on. The memory 102 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 102 in this embodiment may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0158] By programming the design of the processor 101, the code corresponding to the security verification method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 1 the steps of the personalized federated learning method for the vision language model as shown. How to program the design of the processor 101 is a well-known technology to those skilled in the art and will not be elaborated herein.
[0159] Embodiment 3
[0160] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the personalized federated learning method for the vision language model as described in Embodiment 1.
[0161] The computer-readable storage medium may include flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system and various application software installed in the computer device. In addition, the memory may also be used to temporarily store various data that have been output or will be output.
[0162] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A personalized federated learning method for visual language models, characterized in that, Including the following steps: S1. Model the clients in the federated learning system as nodes in an undirected graph, and calculate the edge weights between nodes by the cosine similarity of client parameters, thereby constructing a multi-modal client graph; S2. Divide the client graph into clusters and perform iterative clustering to minimize the structural entropy of the client graph; wherein, the structural entropy is jointly determined by intra-cluster homogeneity and inter-cluster separation; S3. Fine-tune the vision-language model for each client in the cluster through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization during the local update phase of the client to obtain the model upload parameters of the client; S4. According to the results of iterative clustering, the server in each cluster receives the model upload parameters of each client in the cluster, aggregates them to form a global model and distributes it to the clients; Among them, in the clustering process of S2, the malicious clients and normal clients are fully divided into different clusters by globally minimizing the structural entropy; after the clustering ends, the clusters with the number of clients within the cluster less than the preset threshold are directly discarded, so that the malicious clients become ineffective; in step S3, low-rank adaptation is performed through the global model of each client within the cluster to fine-tune the vision-language model; among them, in the t-th round, the global model of client i is calculated as follows: Wherein, represents the global model of client i in the previous round, i.e., round t - 1; ΔW i represents the original parameters of client i; represents low-rank adaptation; M LoRA represents the deployed LoRA module; For each cluster c after clustering is completed, its global model is calculated by the following formula: where i and j represent any two clients, i.e., nodes, within cluster c; W c,i is the parameter of client i within cluster c; W c,j is the parameter of client j within cluster c.
2. The personalized federated learning method for visual language models according to claim 1, wherein In step S2, the calculation formula for the structural entropy of the client graph is: Wherein, is the structural entropy of the client graph G in the t-th round; L is the total number of clusters into which the client graph is divided in each round; n t ; l is the total number of client nodes in each cluster X l ; is the local degree of each node i in X l , defined as the sum of the edge weights connecting i to other nodes in its same cluster; vol(X l ) is the cluster volume, used to quantify the total connection strength inside X l ; vol(G t ) is the global graph volume, used to quantify the cumulative strength of all connections in G t ; g l is the cut edge, defined as the sum of the connection weights connecting X l and nodes in other clusters.
3. The personalized federated learning method for a vision-language model according to claim 1, wherein In step S2, the greedy algorithm is used for iterative clustering; Among them, in the initial state, each client forms an independent cluster, that is, L = N; L is the total number of clusters into which the client graph is divided in each round, and N is the total number of clients in the federated learning system; Iteratively merge the cluster pairs that can minimize the structural entropy, that is, obtain the maximum entropy reduction is the overall client graph structure entropy before merging, is the overall client graph structure entropy after merging; Stop when the entropy reduction is lower than the preset threshold.
4. The personalized federated learning method for a vision-language model according to claim 1, wherein In the client-side local update phase, each client k minimizes the following objective function based on its own global model w t and local data as follows: where h k (w; w t ) is the local objective function of client k; F k (w) is the local loss function of client k, which is used to reflect the matching degree between the local model w of the client and the local data of the client; is an L2 regularization term, which is used to penalize the large deviation between the local model w and the global model w t ; μ is the regularization coefficient, which is used to control the penalty strength for deviating from the global model w t to prevent overfitting of local data 5. The personalized federated learning method for visual language models according to claim 1, characterized in that In step S1, the client graph is represented as G t =(V, W); where the node V represents the client, and the weight calculation formula of the edge W between nodes is N is the total number of clients in the federated learning system; φ i and φ j represent the parameters of client i and client j respectively.
6. The personalized federated learning method for visual language models according to claim 1, characterized in that, In step S3, the LoRa module is selectively deployed in the text encoder of the vision-language model for fine-tuning.
7. A computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the personalized federated learning method for the vision-language model as described in any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the personalized federated learning method for the vision-language model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Robust federated learning abnormal client detection method based on spectral clustering
CN116862023A