Personalized federal learning method for visual language model, terminal and medium
By constructing client graphs and iterative clustering, separating malicious clients, and fine-tuning them using LoRA technology in visual language models, the problem of large Byzantine attacks and communication overhead in federated learning is solved, and effective model fine-tuning and resource conservation are achieved.
Patent Information
- Application Number
- CN202510525367.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Federated learning faces the problems of Byzantine attacks and communication overhead during training, especially in the development of visual language models. Malicious clients may upload arbitrary local updates, destroying global model performance. At the same time, frequent model gradients or parameter exchanges lead to severe latency and memory consumption.
By modeling the client in the federated learning system as nodes in the undirected graph, computing the edge weights between nodes as cosine similarity of client parameters, building a multimodal client graph, and iterative clustering is performed to minimize structural entropy and separate malicious clients. Clients in each cluster fine-tune the visual language model through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization during the local update stage.
Effectively resist Byzantine attacks and reduce communication overhead. Through LoRA technology, only a small part of the model parameters can be updated to fine-tune the visual language model, retaining most of the pre-training weights, and reducing the consumption of computing and storage resources.
Smart Images

Figure CN120069007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and in particular to a personalized federated learning method, terminal and medium for visual language models. Background Technique
[0002] Visual language models have extraordinary capabilities when performing various image classification tasks. The performance of the model largely depends on the training of a large amount of data, which is usually stored on different devices of different organizations. Some privacy data protection regulations have become a bottleneck for further model development. To solve this problem, researchers have introduced federated learning (FL) during the model training process. Federated learning is a distributed deep learning framework that allows multiple devices (referred to as clients in federated learning) to jointly train a central federated learner. Following the principle of "data does not move, model moves", local data is used to jointly train a centralized federated model. This method makes full use of the computing resources of the clients and ensures a certain degree of privacy of the original data during the training process.
[0003] Although federated learning effectively protects the privacy of the original data during the training process, its practice of relying on a large amount of privacy-sensitive data for model development will inevitably attract malicious attackers. For example, federated learning faces the challenge of Byzantine attacks. Malicious participants (referred to as Byzantine clients) may upload arbitrary local updates, damaging the performance of the global model, so it needs to be solved urgently.
[0004] In addition, the federated learning system faces communication overhead. Local clients will encounter computing and storage limitations. Distributed clients and services must frequently exchange model gradients or parameters, usually reaching the scale of millions or billions. The communication bandwidth of the FL system is limited, which will cause serious delays during these frequent exchanges. Local clients often lack the computing resources to fine-tune the entire model, and storing different instances for various tasks will further consume memory resources. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides a personalized federated learning method, terminal and medium for visual language models, which can effectively resist Byzantine attacks while reducing communication overhead.
[0006] To achieve the above object, the present invention provides the following technical solutions: The present invention discloses a personalized federated learning method for visual language models, including the following steps: S1. Model the clients in the federated learning system as nodes in an undirected graph, and the edge weights between the nodes are calculated by the cosine similarity of the client parameters, so as to construct a multimodal client graph; S2. Divide the client graph into clusters and perform iterative clustering to minimize the structural entropy of the client graph; wherein, the structural entropy is jointly determined by intra-cluster homogeneity and inter-cluster separation; S3. Fine-tune the visual language model for each client within a cluster through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization during the client local update phase to obtain the model upload parameters of the client; S4. According to the results of iterative clustering, the server within each cluster receives the model upload parameters of each client within the cluster, aggregates them to form a global model, and distributes it to the clients; Among them, during the clustering process in S2, malicious clients and normal clients are fully separated into different clusters by globally minimizing the structural entropy; after clustering, clusters with the number of clients within the cluster less than the preset threshold are directly discarded, so that malicious clients become ineffective.
[0007] As a further improvement of the above solution, in step S2, the calculation formula for the structural entropy of the client graph is: ; In the formula, is the structural entropy of the client graph in the t th round; ; L is the total number of clusters into which the client graph is divided in each round; is the total number of client nodes in each cluster ; is the local degree of each node in i , defined as the sum of the edge weights connecting i to other nodes in the same cluster; is the cluster quantity, used to quantify the total connection strength within , ; is the global graph quantity, used to quantify the cumulative strength of all connections in , ; is the cut edge, defined as the total connection weight connecting to nodes in other clusters.
[0008] As a further improvement of the above solution, in step S2, the greedy algorithm is used for iterative clustering; Among them, in the initial state, each client forms an independent cluster, that is L = N ; Lis the total number of clusters into which the client graph is divided in each round, N is the total number of clients in the federated learning system; Iteratively merge the cluster pairs that can minimize the structural entropy, that is, obtain the maximum entropy reduction , is the overall structural entropy of the client graph before merging, is the overall structural entropy of the client graph after merging; When the entropy reduction is lower than the preset threshold, stop.
[0009] As a further improvement of the above solution, in step S3, low-rank adaptation is performed through the global model of each client within the cluster to fine-tune the vision-language model; among them, in the t round, the global model of client i is calculated as follows: ; ; In the formula, represents the global model of the client in the previous round, that is, t -1 round; i global model; represents the original parameters of client i ; represents low-rank adaptation; represents the deployed LoRA module; For each cluster c after clustering, its global model is calculated as follows: ; In the formula, i , j represent any two clients, that is, nodes, within cluster c ; is the parameter of client c within cluster i ; is the parameter of client c within cluster j ;
[0010] As a further improvement of the above solution, in the client local update stage, each client k minimizes the following objective function based on its own global model and local data : ; In the formula, is the client kThe local objective function; is the client k The local loss function, which is used to reflect the w matching degree between the client's local model is an L2 regularization term, which is used to penalize the local model and the global model for a large deviation; is the regularization coefficient, which is used to control the penalty strength for deviating from the global model to prevent overfitting of local data .
[0011] As a further improvement of the above solution, in step S1, the client graph is represented as ; where the node represents the client, and the weight calculation formula of the edge W between nodes is , , , N is the total number of clients in the federated learning system; and respectively represent the parameters of the client i and the client j .
[0012] As a further improvement of the above solution, in step S3, the LoRa module is selectively deployed in the text encoder of the vision - language model for fine - tuning.
[0013] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the personalized federated learning method for the vision - language model as described above are implemented.
[0014] The present invention also discloses a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the personalized federated learning method for the vision - language model as described above are implemented.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. The personalized federated learning method for the vision - language model disclosed by the present invention combines the LoRA technology in the parameter - efficient fine - tuning method with personalized joint learning. By using LoRA on the vision - language model, only a small part of the model parameters needs to be updated to fine - tune it, while most of the pre - trained weights are retained, thereby reducing the communication overhead; In a system where the client has inherent partitions or significant differences in data distribution, the traditional client-server federated learning architecture is not the best choice. Instead, a multi-model approach that trains FL models for each customer group is more suitable, which requires the server to adopt an aggregation strategy that takes into account customer heterogeneity and individual characteristics. The present invention proposes a personalized global aggregation method that constructs a customer graph on the server and minimizes its two-dimensional structural entropy. During the global aggregation process, the present invention weights the aggregation based on the similarity between clients. This similarity-based weighting strategy helps to mitigate the impact of data heterogeneity on the global model, thereby obtaining the optimal aggregation strategy while reducing global data heterogeneity. Experiments show that the present invention can effectively resist Byzantine attacks in federated learning.
[0016] 2. In the local update stage of the client in the method of the present invention, a client parameter regularization method is introduced. Each client updates the model according to the global model and its own data distribution, and penalizes updates that deviate significantly from the global model to maintain consistency.
[0017] 3. The computer terminal and computer-readable storage medium disclosed in the present invention can produce the same beneficial effects as the above method by applying the above method, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of the personalized federated learning method for the vision-language model in Embodiment 1 of the present invention.
[0019] Figure 2 It is a deployment schematic diagram of the LoRA module in Embodiment 1 of the present invention.
[0020] Figure 3 It is a flowchart of minimizing the structural entropy in Embodiment 1 of the present invention.
[0021] Figure 4 It is a visualization diagram of the Dirichlet distribution (α = 1) on the non-i.i.d. Flowers dataset in Embodiment 1 of the present invention.
[0022] Figure 5 It is a visualization diagram of the pathological distribution on the non-i.i.d. CIFAR-10 dataset in Embodiment 1 of the present invention.
[0023] Figure 6 It is a Transformer architecture diagram with a LoRA-enhanced text encoder in Embodiment 1 of the present invention.
[0024] Figure 7 It is a comparison diagram of the attack experiment results between the method of the present invention and the FedAvg method in Embodiment 1 of the present invention.
[0025] Figure 8 This is the accuracy rate graph of various data sets in Embodiment 1 of the present invention under the influence of different α values.
[0026] Figure 9 This is a schematic diagram of the computer terminal structure in Embodiment 2 of the present invention. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] Embodiment 1 This embodiment provides a personalized federated learning method for a vision-language model. This method is applied to the training of the vision-language model, and combines the low-rank adaptation technology in the parameter efficient fine-tuning method with personalized federated learning.
[0029] Please refer to Figure 1 , the federated learning method includes the following steps, namely S1~S4.
[0030] S1. Model the clients in the federated learning system as nodes in an undirected graph, and calculate the edge weights between the nodes by the cosine similarity of the client parameters, so as to construct a multi-modal client graph.
[0031] In step S1, the client graph is represented as ; where the node represents a client, and the weight calculation formula of the edge W between the nodes is , , , N , and respectively represent the parameters of client i and client j . The client parameters only include the parameters of the text encoder, and the parameters can be regarded as single-modal data with multi-modal associations.
[0032] S2. Divide the client graph into clusters and perform iterative clustering to minimize the structural entropy of the client graph, with the aim of minimizing the structural chaos of the client graph. Among them, the structural entropy is jointly determined by intra-cluster homogeneity and inter-cluster separation.
[0033] In step S2, the calculation formula of the structural entropy of the client graph is: ; In the formula, is the structural entropy of the t round client graph ; L is the total number of clusters into which the client graph is divided in each round; is the total number of client nodes in each cluster ; is the local degree of each node in i , defined as the sum of the edge weights connecting i to other nodes in its same cluster; is the cluster quantity, used to quantify the total connection strength inside, reflecting its internal cohesion, ; is the global graph quantity, used to quantify the cumulative strength of all connections in ; In addition to the intra-cluster dynamics, each cluster also interacts with other clusters through the cut edges . The cut edge is defined as the sum of the connection weights connecting and the nodes of other clusters. These variables together define the structural entropy, where the local degree and the cluster volume determine the homogeneity within the cluster, while the cut edge and the global volume affect the separation between clusters. By iteratively merging clusters to optimize these metrics, this approach gradually reduces the uncertainty in the hierarchical structure of the graph.
[0034] In step S2, a greedy algorithm is used for iterative clustering; Among them, in the initial state, each client forms an independent cluster, that is L = N ; L is the total number of clusters into which the client graph is divided in each round, N is the total number of clients in the federated learning system; Iteratively merge the cluster pair that can minimize the structural entropy, that is, obtain the maximum entropy reduction , is the structural entropy of the overall client graph before merging, is the structural entropy of the overall client graph after merging; When the entropy reduction is lower than the preset threshold, it stops. In this embodiment, the preset threshold is set to 0.05 according to experience.
[0035] In the scenario of visual language model fine-tuning, traditional federated learning frameworks struggle to handle client heterogeneity. These systems enforce strict global model consistency and are unable to adapt to clients with diverse data distributions or specific task requirements. For systems with inherent partitions or significant differences in data distribution among clients, the traditional client-server federated learning architecture for training a shared global model is not the optimal choice. A multi-model approach that trains a federated learning model for each homogeneous client group is more appropriate, and the server requires an aggregation strategy that can take into account client heterogeneity and their individual characteristics. Clustering similar clients can improve the collaborative training effect. Therefore, a multi-model approach is better, that is, training a federated learning model for each homogeneous client group. Clustering Federated Learning (CFL) divides clients into several clusters, and each cluster is represented by a model centroid.
[0036] Although clustering-based federated learning methods provide a certain solution through client grouping, they ignore the subtle relationships between clients that are crucial for effective knowledge transfer. Static client grouping further ignores the evolution of model behavior and potential relationships, and cannot detect the development of complementary feature extractors between clients with different original data during the training process. Traditional k-means clustering cannot dynamically adjust the number of clusters and only relies on the mean parameter vector information, thus ignoring multimodal interaction relationships. Therefore, the present invention addresses these limitations through structure entropy-guided aggregation, which is a novel personalized federated learning framework for visual language models that can optimize dynamic client clustering.
[0037] S3. Fine-tune the visual language model for the clients within each cluster through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization during the client local update phase to obtain the model upload parameters of the clients (i.e., the updated and ).
[0038] Please refer to Figure 2 , and cluster the clients according to the similarity of the models or tasks (cluster 1 to cluster M). Each cluster shares a common LoRA adapter. Each client uses a frozen pre-trained image / text encoding in combination with a trainable LoRA adapter.
[0039] In the present invention, CLIP (Contrastive Language-Image Pre-Training) is a multi-modal pre-trained neural network that includes two core components: an image encoder I(·) and a text encoder T(·). These encoders convert images and text into vector representations and project them into a shared embedding space. The training objective of CLIP is to maximize the cosine similarity between the embeddings of matching image-text pairs while reducing the similarity of non-matching pairs. Local training uses parameter regularization to align the client-specific LoRA updates with the stability of the global model.
[0040] For inference, CLIP computes the normalized embedding of an image f i = I ( x i ) and a text prompt t k = T ( c k ), such as "a picture of [category]". To predict the category of a test image x , the model computes the I ( x ) and the cosine similarity between all candidate text embeddings , scaled by a temperature parameter. The predicted probability i of category y i is as follows: ; In the formula, T represents the text feature vector; T ( t i ) is the text feature vector of the i th category; K represents the total number of categories; τ is the temperature parameter used to scale the value of the cosine similarity.
[0041] By aligning the multi-modal representations through contrastive learning, CLIP achieves robust zero-shot generalization without the need for task-specific re-training.
[0042] The LoRA model incrementally updates the pre-trained weights through the product of two small matrices A and B, based on the intrinsic rank of the downstream task. In this embodiment, CLIP is used as the backbone vision-language model, and its text encoder is fine-tuned through low-rank adaptation (LoRA). For the Q, K, V, and O matrices of the text encoder ( W q , Wk , W v , W o ∈R d×d ), in this embodiment, trainable low-rank matrices A and B are injected to approximate the weight update, where R is the set of real numbers, the dimension of the multi-head attention matrix in the Transformer is d×d, and the LoRa module can be selectively deployed on the Transformer to optimize the transmission efficiency. The specific deployment method can refer to the existing research: Maxime Zanella and Ismail Ben Ayed. 2024. Low-rank few-shot adaptation of vision-language models. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 1593–1603.
[0043] The above mechanism establishes a robust framework for the fine-tuning of vision-language models, which is formalized as follows: ; In the formula, the scaling factor α regulates the contribution of the low-rank update . Here, B ∈ R d×r , A ∈ R r×k , and the dimension of the original weight matrix of the local model is d×k. The rank r is a hyperparameter used to control the inner dimension of matrices A and B. The dimension of matrix A is d×r, and the dimension of matrix B is r×k. The size of the trainable parameters is reduced from d×k to r×(d + k). r is much smaller than the smaller value of d and k, so that while only adjusting A and B, the original weight W is retained, realizing efficient parameter update.
[0044] In LoRA fine-tuning, the scaling factor α and the rank r act together to regulate the adjustment of the adapter to the weights of the pre-trained model. The ratio α / r scales the amount of weight update , where α globally amplifies the influence of the adapter, while r limits the dimension of the low-rank decomposition, thus controlling the number of trainable parameters. Adjusting the ratio α / r can balance between retaining the knowledge of the original model and adapting to specific tasks. For example, increasing αIt can enhance the update amplitude of the adapter, which is beneficial for complex tasks that require significant weight adjustment; while reducing α can reduce the update amount, thereby avoiding overfitting in the case of less data. The key is that this reduction in parameters does not affect performance because the low-rank matrix still has the ability to capture the key changes in the input data.
[0045] In step S3, low-rank adaptation is performed through the global model of each client within the cluster to fine-tune the vision-language model; among them, in the t round, the client i global model is calculated as: ; In the formula, represents the global model of the client in the previous round, i.e., t -1 round; i global model; represents the original parameters of the client i ; represents low-rank adaptation; represents the deployed LoRA module; For each cluster after clustering c , its global model is calculated as: ; In the formula, i , j represent any two clients within the cluster, i.e., nodes; c c is the parameter of the client i within the cluster ; c is the parameter of the client j within the cluster ;
[0046] In the local update stage of the client, each client k minimizes the following objective function based on its own global model and local data : ;
[0047] In the formula, is the local objective function of the client k ; is the local loss function of the client k , which is used to reflect the matching degree between the local model w of the client and the local data of the client; is an L2 regularization term used to penalize the local model and the global model for large deviations; is the regularization coefficient (hyperparameter) used to control the penalty strength for deviating from the global model A larger will enforce stricter consistency with the global model, ensuring that client updates are consistent with the global model and preventing overfitting to local data . By penalizing large deviations, it can stabilize the federated learning process and mitigate client drift, otherwise the local model may deviate significantly from the global model.
[0048] It should be noted that the global model of the cluster is the model sent to the client in the next round. The complete federated learning process is as follows: 1. The client receives the global model of the cluster, trains it using local data, and then forms the updated client parameters using the parameter regularization method and LoRA. After that, the client uploads the client parameters, and the server aggregates them to form the global model of the cluster. This concludes one round.
[0049] As Figure 3 shown, the federated learning framework of the present invention can promote knowledge transfer based on semantics while retaining client-specific patterns, overcoming the limitations of traditional federated learning and static collaborative federated learning methods. The structural entropy formula, as an information-theoretic regularization term, automatically balances the personalization and generalization capabilities of the model throughout the training process.
[0050] S4. According to the results of iterative clustering in S2, the servers within each cluster receive the model upload parameters of each client within the cluster, aggregate them to form a global model, and send it to the client; among them, during the clustering process, the malicious clients and normal clients are fully separated into different clusters by globally minimizing the structural entropy; after clustering, the clusters with the number of clients within the cluster less than the preset threshold are directly discarded and do not participate in the global aggregation, so that the malicious clients become ineffective.
[0051] In this embodiment, the effectiveness of the above method is also demonstrated through experiments. The specific content of the experiment is as follows: 1 Experimental settings 1.1 Datasets This embodiment uses six datasets: Digit5, Flowers102, Caltech101, Describable Textures Dataset (DTD), Oxford-IIIT Pet, and CIFAR-10. Among them, Flowers102, Caltech101, DTD, and Oxford-IIIT Pet belong to the label-skewed scenario datasets, while Digit5 belongs to the feature-offset scenario dataset.
[0052] Digit5 is a multi-domain dataset for digit classification, containing five domains: MNIST, SVHN, USPS, MNIST-M, and SynthDigits. Each domain contains handwritten or synthetic digit images from 0 to 9, with different image styles and distributions. For example, MINIST contains grayscale handwritten digits, while SVHN contains colored house numbers from street view images. The feature shift between different domains makes it challenging for models to generalize across domains. In this example, 25,000 training samples and 9,000 test samples are randomly selected from each domain.
[0053] Flowers102 is a fine-grained image classification dataset containing 102 flower categories. The dataset contains approximately 8,189 images, with each category having between 40 and 258 images. These images are taken under different conditions, including different lighting, angles, and backgrounds. The label bias stems from the uneven distribution of images across categories, making it suitable for federated learning scenarios where client data may be imbalanced.
[0054] Caltech101 is a classic image classification dataset containing 101 object categories and an additional background category. Each category has approximately 40 to 800 images, for a total of approximately 9,000 images. The dataset covers a wide variety of objects such as animals, vehicles, and everyday items. Due to the unequal number of images per category, the label bias is significant. The Describable Textures Dataset (DTD) is a dataset for label bias scenarios focused on texture recognition, containing 47 texture categories such as stripes, squares, and checkerboards. Each category has 120 images, with image sizes ranging from 300x300 to 640x640.
[0055] Oxford-IIIT Pet is a pet image classification dataset containing 37 pet categories and a total of 7,349 images. The number of images per category varies, from dozens to hundreds. The dataset contains images of cats and dogs, and both species and their breeds are labeled.
[0056] CIFAR-10 is a standard benchmark dataset in computer vision research, containing 60,000 color images covering 10 object categories. As a standard evaluation benchmark for image classification models, it is widely used in research on feature learning. Each 32×32 pixel image belongs to one of the following categories: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck.
[0057] 1.2 Data Distribution In the experimental setup section, to simulate a heterogeneous environment, this embodiment adopts the Dirichlet distribution and the pathological distribution. Each of these two distributions has unique characteristics and can effectively simulate the non-independent and identically distributed (non-IID) nature of real-world data.
[0058] The Dirichlet Distribution is a commonly used multivariate probability distribution for generating probability vectors. It is a multi-dimensional generalization of the Beta distribution and is suitable for modeling categorical distributions. In federated learning, the Dirichlet distribution is used to simulate the class distribution of client data. By adjusting the parameter α of the Dirichlet distribution, the proportion of different class data drawn by each client from the global dataset can be controlled, thereby simulating different degrees of data heterogeneity. In this embodiment, α is set to 0.1. Figure 4 Shows the resulting label statistical heterogeneity on the Flowers dataset. This partitioning is non-independent and identically distributed, which is evident from the fact that certain classes are missing from each client.
[0059] This embodiment uses the Pathological Distribution to simulate extreme or specific scenarios. This method is used to test the robustness and adaptability of the algorithm under extremely unbalanced data distributions. It simulates the situation of uneven data distribution in the real world, such as the differences in patient data from hospitals in different regions, or different data patterns on user devices. For example, if there are a total of ten classes, Client 1 only has classes 1, 2, and 3, Client 2 only has classes 4, 5, and 6, and Client 3 only has classes 7, 8, 9, and 0. This embodiment divides the dataset into three non-overlapping class subsets, and each client is only assigned a part of them. By restricting each client to only use a specific subset of the data, the pathological distribution ensures that there are significant differences in data characteristics among clients. Figure 5 Shows the resulting label distribution heterogeneity on CIFAR-10.
[0060] 1.3 Baseline By applying LoRA to more network layers and vectors, the performance of the model can be further improved. However, this improvement comes at the cost of increased computational burden because adding more weight matrices introduces additional parameters. LoRA provides a method to balance performance and parameter efficiency, enabling effective fine-tuning of the model while maintaining its original size. Therefore, in this embodiment, LoRA is only applied to the QKVO vectors of the 12-layer transformers in the text encoder, while keeping the image encoder unchanged. The specific structure is as Figure 6 shown. For a fair and comprehensive evaluation, this embodiment selects FedCLIP as the baseline. FedCLIP inserts an adapter on top of CLIP to identify the part of the features suitable for a specific task.
[0061] 1.4 Implementation Details In the experiment, this embodiment adopted a dynamic learning rate and set the number of local training rounds for each client to 50 rounds. A total of 10 clients participated in each round of sampling, and the batch size was 128 to balance computational efficiency and model stability. The clients can have different numbers of training samples. For the Adam optimizer, the learning rate was set to 5e-5, the momentum was set to 1e-6, and the weight decay was set to 0.2 to promote stable convergence and reduce overfitting. In addition, when using the LoRA method, this embodiment set the rank r to 2 and the scaling factor α to 32 to achieve fast adaptation and generalization of the model using low-rank approximation. These meticulous parameter selections are aimed at ensuring the accuracy of the experimental results and optimizing the model performance.
[0062] 1.5 Privacy Issues Byzantine attacks pose a significant threat to federated learning systems because they not only disrupt the convergence of the model but also reduce the accuracy and usability of the model. In the non-independent and identically distributed (non-IID) data setting, there are significant differences in the local data distributions of the clients, which makes it more difficult for federated learning systems to distinguish between benign and malicious model updates. Attackers can take advantage of this by subtly manipulating the updates to conceal their intentions, thereby reducing the effectiveness of traditional robust aggregation rules in identifying and filtering malicious updates. The sign-flipping (SF) attack is a simple Byzantine attack method where malicious clients disrupt the training of the global model by reversing the sign of their model updates.
[0063] The SF attack disrupts federated learning by reversing the direction of the gradient update, turning gradient descent into gradient ascent to maximize the model loss and prevent convergence. Different from poisoning attacks that modify training data or model parameters, the SF attack directly exploits the federated framework to flip the gradient sign during local training. This attack has been proven effective in both FedSGD (Federated Stochastic Gradient Descent algorithm) and FedAvg (Federated Averaging algorithm): In FedSGD, the attacker submits reversed gradients. Different from the cumulative damage caused by sign flipping in repeated local iterations in FedAvg, the simplicity of the SF attack stems from avoiding complex data or model operations, making it particularly suitable for federated settings where clients compute updates independently. The SF attack has extremely low requirements for the capabilities of adversaries. An attacker only needs to control one or more compromised clients without colluding with other clients or obtaining external updates. The SF attack requires very little knowledge of the adversary. An attacker does not need to know the global model architecture, the data distributions of other clients, or the aggregation rules, enabling the adoption of a "black-box" strategy.
[0064] Although this method is simple, it can effectively disrupt the training process, especially when the detection and defense mechanisms of the federated learning system against malicious updates are weak. Implementing the sign-flipping attack does not require in-depth knowledge of the global model or data, so it is a low-cost and easy-to-execute attack strategy.
[0065] 2 Experimental Results The experimental results were obtained under the above experimental settings. This embodiment uses the following research questions as the guidelines for the experiment: Research Question 1: Compared with the FedCLIP algorithm (Lu W, Hu X, Wang J, et al. Fedclip: Fast generalization and personalization for clip in federated learning[J]. arXiv preprint arXiv:2302.13485, 2023.), can the method of the present invention achieve better performance on multiple datasets under the data heterogeneity of the Dirichlet distribution? Research Question 2: How does the method of the present invention perform compared with FedCLIP in the scenario of extreme pathological data distribution? Research Question 3: How does the scaling factor α in LoRA affect the performance of the method of the present invention on different datasets? Research Question 4: Compared with the classical FedAvg algorithm, is the method of the present invention more robust against Byzantine attacks? 2.1 Research Question 1: Performance Comparison The experiment of this embodiment compared the proposed method with FedCLIP on six benchmark datasets: CIFAR-10, Caltech101, Flowers, DTD, Oxford-IIIT Pet, and Digit5. This embodiment focused on the federated fine-tuning of visual language models. Table 1 summarizes the experimental results, showing that the present invention has always been significantly better than FedCLIP, with an average accuracy improvement of 9.04%. The present invention exceeded FedCLIP on all datasets, especially with significant improvements on Caltech101, Flowers, and DTD. These results indicate that integrating LoRA into the QKVO matrix of the text encoder is effective, enabling the model to adapt to task-specific language patterns while retaining the pre-trained visual representation of CLIP. Even on relatively simple tasks like Digit5, the present invention achieved a 9.67% improvement, highlighting its robustness in addressing cross-domain challenges.
[0066] Table 1: Comparison of Experimental Results between FedCLIP and the Method of the Present Invention under Dirichlet Distribution Data ; The above-mentioned FedCLIP method trains a single adapter on top of CLIP, which limits its ability to adapt to heterogeneous tasks, as reflected in its suboptimal performance on complex datasets such as DTD and Flowers. In contrast, the method of the present invention performs a low-rank update on the core attention mechanism (QKVO) of the text encoder, enabling targeted adjustment without adding excessive parameters. This strategy reduces computational overhead while retaining the generalization ability of the original model. By only deploying low-rank adaptation (LoRA) in the text encoder, the method of the present invention achieves a strategic balance between adaptation efficiency and computational cost.
[0067] There is a significant performance gap between the method of the present invention and FedCLIP, which reveals the limitations of simple adapter-based methods in a federated learning environment. The global adapter of FedCLIP operates as a monolithic component independent of the intrinsic architecture of CLIP and is difficult to adapt to different client-specific tasks. In contrast, the method of the present invention distributes the adaptation process into the attention mechanism of the text encoder through low-rank adaptation, thereby achieving localized and efficient updates. By embedding LoRA into the QKVO matrix of the text encoder, the method of the present invention achieves a principled balance between task-specific adaptation and retention of pre-trained knowledge. This method not only improves the state-of-the-art performance but also provides a scalable framework for existing federated learning deployments.
[0068] 2.2 Research Question 2: Pathological Distribution Analysis This embodiment evaluates the method of the present invention and FedCLIP under pathological and Dirichlet data distributions. Tables 2 and 1 show the comparison results. In the pathological setting (Table 2), the clients access mutually exclusive class subsets, and the method of the present invention performs excellently on all datasets. On Caltech101, the method of the present invention is 9.31% higher than FedCLIP, 27.89% higher on Flowers, 14.96% higher on DTD, and 7.08% higher on Oxford-IIIT Pet, with an average improvement of 14.81%. The performance gap on the Flowers dataset is the largest: the accuracy of FedCLIP is only 68.27%, while the method of the present invention reaches 96.16%. These results indicate that in the case of extreme data scarcity and class isolation, traditional adapter methods like FedCLIP cannot retain transferable knowledge. The present invention achieves parameter-efficient adaptation while retaining important pre-trained features by strategically integrating LoRA into the QKVO vectors of the CLIP text encoder.
[0069] Table 2: Comparison of Experimental Results between FedCLIP and the Method of the Present Invention under Pathological Distribution Data ; In the case of Dirichlet distribution data with partial class overlap (Table 1), the method of the present invention still maintains an advantage, but the magnitude of the advantage decreases. The advantage in average accuracy drops from 14.81% in the pathological setting to 9.04% in the Dirichlet scenario. This comparison shows that although benefiting from class overlap in the Dirichlet setting, the FedCLIP shallow adapter architecture is still fundamentally limited in cross-client knowledge integration and cannot compare with the deep parameter efficient strategy of the present invention.
[0070] 2.3 Research Question 3: Robustness Analysis In this embodiment, the personalized federated learning framework of the present invention is compared and evaluated with the benchmark method in scenarios with and without attacks. Figure 7 The results of the method of the present invention in resisting sign-flipping attacks on different datasets are shown.
[0071] In Dirichlet distribution data (30% malicious clients), the traditional FedAvg shows a significant drop in accuracy during the collaborative attack, dropping to 88.79% on Caltech101, 40.57% on DTD, 65.70% on Flowers, and 85.16% on Oxford-IIIT Pet, on average 25.3% lower than the method of the present invention. The greedy clustering strategy of the present invention isolates malicious clients by splitting the network into two subgraphs, effectively excluding smaller attacker clusters (30% of the nodes) from global aggregation. This defense mechanism improves the DTD texture classification accuracy by 66.6% compared to the compromised FedAvg (67.60% vs. 40.57%) and is particularly effective for vision-language tasks.
[0072] Comparison with the non - attack scenario (Table 1) shows that the method of the present invention causes only a minimal loss of accuracy while maintaining security. Between benign and attack conditions, the average accuracy drops only 0.83% (from 87.98% to 87.15%), which is 9.04% higher than the non - attack baseline of FedCLIP while maintaining robustness. There are three key mechanisms that enable the system to have this anti - attack ability: ① The graph - based clustering of the present invention isolates malicious clients through two - dimensional similarity embedding minimization; ② The framework separates clients with different update patterns (e.g., gradient sign flipping) to prevent model contamination; ③ Different from the simple averaging of FedAvg, two - dimensional similarity embedding minimization can quantify the credibility of clients without prior knowledge of the attack. The system dynamically partitions clients according to the update similarity pattern, actively adapting to changing threats such as sign flipping and other gradient manipulation attacks. This method achieves an attack detection accuracy of 92.4% (derived from excluding 30% of malicious clients), and the performance overhead is less than 1% 2.4 Research Question 4: Ablation Study Please refer to Figure 8 , in this embodiment, it is analyzed how different scaling factors (α) affect the performance of LoRA - based fine - tuning. The research covers four datasets: Caltech101, Flowers, DTD, and Oxford - IIIT Pet. The ablation study shows that α has a significant impact on the accuracy of the model, indicating that it is necessary to balance adaptability and generalization ability by carefully selecting α during the fine - tuning process.
[0073] The optimal choice of α depends on the complexity of the dataset. For the Caltech101 dataset, the highest accuracy can be achieved when α = 32, while smaller α values (e.g., α = 4) will reduce performance. It is worth noting that too large an α value (such as α = 64) will also reduce accuracy, with diminishing returns. Similarly, for the Flowers dataset, the peak accuracy is reached at α = 32, and the accuracy drops significantly at α = 8 and α = 4, indicating that smaller α values may not be able to adapt to the specific details of the task, while larger values risk overfitting or making the model unstable. In contrast, the DTD dataset performs well at α = 64, meaning that tasks with higher intrinsic variability (such as texture recognition in DTD) may benefit from stronger updates to capture diverse patterns. For Oxford - IIIT Pet, α = 32 performs better than other values, but even a moderate reduction (e.g., α = 16) will significantly affect the accuracy, highlighting the dependence of this dataset on precise parameter tuning.
[0074] Larger α values (e.g., 64) can enhance the adaptability to complex tasks such as DTD, but there is a risk of overfitting on datasets with lower diversity such as Oxford-IIIT Pet. Smaller α values (e.g., 4 - 8) can preserve the pre-trained features but cannot adapt to the nuances of specific tasks in tasks such as Flowers and Caltech101. However, α = 32 achieves near-optimal results on most datasets, striking a balance between task-specific adaptability and the generalization ability of the pre-trained model.
[0075] The empirical results of this embodiment show that choosing the appropriate α is crucial for LoRA-based fine-tuning. A moderate α = 32 can maintain a balance between adaptability and generalization ability, achieving robust performance on different datasets. This finding highlights the role of α as an adjustable lever, facilitating coordination between model stability and task-specific learning, especially in the context of federated learning where heterogeneous data distributions require flexible and controllable updates.
[0076] Embodiment 2 This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the personalized federated learning method for the vision language model as described in Embodiment 1 are implemented.
[0077] As Figure 9 shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to at least one processor 101. In this embodiment, the specific connection medium between the processor 101 and the memory 102 is not limited. Figure 9 Here, it is taken as an example that the processor 101 and the memory 102 are connected through a bus 100. The bus 100 is represented by a thick line in Figure 9 Here. The connection manners between other components are only illustrative and not to be taken as a limitation. The bus 100 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 it is only represented by a thick line in here, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 101 can also be called a controller, and there is no limitation on the name.
[0078] In this embodiment, the memory 102 stores instructions executable by at least one processor 101. By executing the instructions stored in the memory 102, at least one processor 101 can execute the foregoing method.
[0079] Among them, the processor 101 is the control center of the device, which can connect various parts of the entire control device through various interfaces and circuits. By running or executing the instructions stored in the memory 102 and calling the data stored in the memory 102, various functions of the device and process data, so as to monitor the device as a whole.
[0080] In a possible design, the processor 101 may include one or more processing units. The processor 101 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 101 either. In some embodiments, the processor 101 and the memory 102 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.
[0081] The processor 101 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, which can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the personalized federated learning method for the vision language model disclosed in conjunction with Embodiment 1 can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor 101.
[0082] The memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 102 can include at least one type of storage medium. For example, it can include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disc, and so on. The memory 102 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 102 in this embodiment can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0083] By programming the design of the processor 101, the code corresponding to the security verification method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 1 the steps of the personalized federated learning method for the vision language model shown. How to program the design of the processor 101 is a well-known technology to those skilled in the art and will not be elaborated herein.
[0084] Embodiment 3 This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the personalized federated learning method for the vision language model as described in Embodiment 1.
[0085] The computer-readable storage medium may include flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory may also be used to temporarily store various data that have been output or will be output.
[0086] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A personalized federated learning method for visual language models, characterized by: The following steps are involved: S1. The clients in the federated learning system are modeled as nodes in an undirected graph. The edge weights between nodes are calculated by the cosine similarity of the client parameters, thereby constructing a multimodal client graph. S2. Divide the client graph into clusters and perform iterative clustering to minimize the structural entropy of the client graph; wherein the structural entropy is jointly determined by intra-cluster homogeneity and inter-cluster separation; S3. Fine-tune the visual language model of the client in each cluster through the LoRA module, and constrain the deviation between the local model and the global model through parameter regularization in the local update stage of the client to obtain the model upload parameters of the client; S4. According to the result of iterative clustering, the server in each cluster receives the model parameters uploaded by each client in the cluster, aggregates them into a global model and sends it to the client; Among them, in the clustering process of S2, the global structural entropy is minimized to fully divide malicious clients and normal clients into different clusters; after the clustering is completed, the clusters with the number of clients less than the preset threshold are directly discarded, thereby making the malicious clients ineffective.
2. The personalized federated learning method for the visual language model according to claim 1, characterized in that: In step S2, the calculation formula of the structural entropy of the client graph is: ; In the formula, For the t Wheel Client Diagram The structural entropy of L is the total number of clusters that the client graph is divided into in each round; For each cluster The total number of client nodes in; for Each node i The local degree is defined as the connection i The sum of the edge weights of other nodes in the same cluster; is the cluster quantity, used to quantify The total internal connection strength, ; is the global graph quantity, used for quantification The cumulative strength of all connections in ; is a cut edge, defined as a connection The sum of the connection weights with other cluster nodes.
3. The personalized federated learning method for a visual language model according to claim 1, characterized in that: In step S2, a greedy algorithm is used for iterative clustering; In the initial state, each client forms an independent cluster, that is, L = N ; L is the total number of clusters that the client graph is divided into in each round, N is the total number of clients in the federated learning system; Iteratively merge the cluster pairs that can reduce the structural entropy to the greatest extent, that is, obtain the maximum entropy reduction , is the overall client graph structure entropy before merging, is the overall client graph structure entropy after merging; When entropy decreases Stop when it falls below a preset threshold.
4. The personalized federated learning method for visual language model according to claim 1, characterized in that: In step S3, low-rank adaptation is performed through the global model of each client in the cluster to fine-tune the visual language model; t In round, client i The global model The calculation formula is: ; In the formula, Indicates that the previous round t -1 round of clients i Global model; Represents the client i The original parameters of represents low-rank adaptation; Indicates the deployed LoRA module; For each cluster after clustering c , its global model The calculation formula is: ; In the formula, i , j Representation Cluster c Any two clients in are nodes; Cluster c Clients within i Parameters; Cluster c Clients within j Parameters.
5. The personalized federated learning method for visual language model according to claim 1, characterized in that: During the client local update phase, each client k Based on its own global model and local data Minimize the following objective function: ; In the formula, For Clients k The local objective function of For Clients k The local loss function is used to reflect the client local model w The degree of match with the client's local data; is an L2 regularization term used to penalize the local model With the global model The larger deviation of is the regularization coefficient, which is used to control the deviation from the global model The penalty intensity is to prevent overfitting of local data .
6. The personalized federated learning method for visual language model according to claim 1, characterized in that: In step S1, the client graph is represented as ; Among them, the node Represents the client, the edge between nodes W The weight calculation formula is: , , , N is the total number of clients in the federated learning system; and Respectively represent the client i and client j Parameters.
7. The personalized federated learning method for visual language model according to claim 1, characterized in that: In step S3, the LoRa module is selectively deployed in the text encoder of the visual language model for fine-tuning.
8. A computer terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the personalized federated learning method for the visual language model as described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the personalized federated learning method for a visual language model as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Robust federated learning abnormal client detection method based on spectral clustering
CN116862023A
Self-adaptive clustering federal learning method with differential privacy
CN117634594A
Rolling bearing fault diagnosis method based on dynamic clustering federated learning
CN118362309A
Method for collaborative training of ML model based on relay assisted federated learning with over-the-air computation and data privacy
WO2025046273A1
Cited By
Personalized federal learning method and system based on domain invariant text representation and intra-domain global prior
CN120508883A