Clustering-based Adaptive Fine-tuning Method for Heterogeneous Federated Foundation Models and Computer Device

Through the adaptive fine-tuning method of heterogeneous federation basic model based on clustering, the problem of FM deployment and training imbalance in resource heterogeneous scenarios is solved, efficient model training and accuracy improvement are achieved, and computing and communication overhead is reduced.

CN119646552BActive Publication Date: 2025-07-11THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411759651.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-07-11
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

In the highly resource heterogeneous scenario, traditional federated learning methods have problems with FM deployment problems, large computing and communication overhead, unbalanced training and poor performance. Especially in network connections with limited bandwidth, existing methods cannot effectively solve the challenges brought by model heterogeneity.

Method used

Adaptive fine-tuning method of heterogeneous federation basic model based on clustering is adopted, and the multi-factor heterogeneous perceptual clustering module and knowledge-aware model architecture search algorithm is used to select representative clients and optimal sub-models for each cluster, and the server model is optimized through the cluster-aware knowledge transfer module, and the client model is updated in combination with reverse knowledge distillation to achieve efficient training under resource limitations.

Benefits of technology

It effectively reduces computing and communication overhead, improves training efficiency and model adaptability, enables each client's model to meet resource limitations, improves training accuracy, and is significantly better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646552B_ABST
    Figure CN119646552B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of heterogeneous federated foundation model adaptation, and specifically relates to a clustering-based heterogeneous federated foundation model adaptive fine-tuning method and a computer device. The solution includes: through a multi-factor heterogeneous perception clustering module, a representative client is selected for each cluster, and the representative client will select a corresponding model as the cluster model according to its computing power limit; through a knowledge-aware model architecture search algorithm, an optimal sub-model based on the cluster model is searched for all clients within each cluster, and the optimal sub-model is deployed on the clients; the parameters are uploaded to the representative client, and the corresponding parameters are aggregated on the representative client, and after aggregation, they are sent to the clients within the cluster. Through a cluster-aware knowledge transfer module, the knowledge of each cluster is transferred to the server model, and through reverse knowledge distillation, the knowledge of the server model is back-transmitted and used to update the representative client of each cluster. The present invention is applicable to the adaptive fine-tuning of heterogeneous federated foundation models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of heterogeneous federated foundation model adjustment, and particularly relates to a clustering-based adaptive fine-tuning method for heterogeneous federated foundation models and a computer device. Background Art

[0002] Large models are pre-trained on a large amount of data sets, making them adaptable to a very large number of application scenarios and having broad generalization. Therefore, by fine-tuning the large model on a specific data set, the large model can be adapted to various specific tasks. However, in reality, there are many problems in fine-tuning large models, especially the leakage of private data. Most data is distributed locally and is private, which limits the data scope for fine-tuning large models. In the cloud-edge collaborative training scenario, federated learning emerges as a promising method, which can achieve collaborative training of models among multiple clients without directly exchanging private data. However, due to the increasing number of parameters in large models, usually having billions of parameters or even more, many edge clients cannot deploy or fine-tune these large models. However, traditional federation, such as FedAvg, requires clients and servers to share the same model for parameter aggregation, so traditional federation is not applicable. How to optimize heterogeneous federated models has become an urgent task.

[0003] Current research has proposed several strategies to achieve heterogeneous federation, including methods based on knowledge distillation and partial training. For example, methods based on knowledge distillation such as FedDF and DS-FL enable small models to be deployed on clients and large models to be deployed on servers, and transmit logits (maximum likelihood estimation). Logits (logical values) usually refer to the raw and unprocessed scores or scores of the output layer of the model. The small model on the client is regarded as the teacher to guide the training of the model on the server, thereby achieving heterogeneous federation of models.

[0004] However, these methods require each client to interact with the server, and the knowledge distillation method has a high time cost. For clients with very limited computing resources, it may not be possible to find a suitable smaller version of the FM (Foundation Model) for deployment, or there may not be enough versions of the FM to match the highly heterogeneous client resource constraints, resulting in resource waste, etc. Based on the partial training methods such as HeteroFL and FedRolex, sub-models of the server model will be deployed on the clients, and the corresponding parameters will be uploaded to the server for aggregation, achieving model heterogeneous federation in this way. However, HeteroFL can only fine-tune the first part of the parameters of each layer of the model and cannot fine-tune the entire model. Although FedRolex adopts the method of rolling extraction of sub-models, making all parameters able to be fine-tuned, it can only be randomly rolled extraction and cannot highlight the importance of some layers.

[0005] Therefore, the existing methods have the following problems:

[0006] 1. FM deployment problem;

[0007] 2. Collaborative training of FMs brings huge computational and communication overheads. In federated learning, frequent exchange of large model parameters or gradients can lead to significant communication and computational overheads, especially in network connections with limited bandwidth;

[0008] 3. Heterogeneous data and resource distributions lead to unbalanced training, slow convergence speed, and poor performance. FMs place higher requirements on the quality of training data, so the actual data and resource heterogeneity will have a serious impact. Summary of the Invention

[0009] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a clustering-based heterogeneous federated foundation model adaptive fine-tuning method and a computer device, which effectively solve the problem of fine-tuning FMs in a highly resource-heterogeneous scenario, enabling the models deployed on each client to meet their resource constraints and significantly reducing the computational and communication overheads.

[0010] The present invention adopts the following technical solutions to achieve the above purpose. In the first aspect, the present invention provides a clustering-based heterogeneous federated foundation model adaptive fine-tuning method, including:

[0011] S1. Through a multi-factor heterogeneous perception clustering module, clustering is performed by comprehensively considering the computing power resource constraints and data distributions of each client, and a representative client is selected for each cluster. The representative client will select a corresponding model as the cluster model according to its computing power constraints;

[0012] S2. Through the knowledge-aware model architecture search algorithm, according to the heterogeneous computing power limitations of each client, search for the optimal sub-model based on the cluster model for all clients within each cluster, and deploy the optimal sub-model on the client;

[0013] S3. Conduct local training on the clients within each cluster, then upload the parameters to the representative client, perform aggregation of the corresponding parameters on the representative client, and after aggregation, distribute them to the clients within the cluster. Repeat step S3 until the training within the cluster is completed;

[0014] S4. Through the cluster-aware knowledge transfer module, transfer the knowledge of each cluster to the server model to achieve the training of the server model;

[0015] S5. Through reverse knowledge distillation, transfer the knowledge of the server model back and update the representative client of each cluster. The updated representative client distributes the corresponding parameters to each client within the cluster.

[0016] Furthermore, step S1 specifically includes:

[0017] The multi-factor heterogeneous-aware clustering module uses the K-means algorithm, comprehensively considering the computing power resource limitations and data distribution of each client, and divides the clients with similar data distribution and computing power limitations into one cluster;

[0018] During clustering, the differential privacy method is used to add Gaussian noise to the data of each client. For client i, its data distribution is P(D i ), after adding Gaussian noise, the data feature is:

[0019] Δ f represents the sensitivity of the function, and ∈ represents a parameter measuring the strength of privacy protection;

[0020] For each cluster, according to the computing power limitations of each client within the cluster, select the client with the maximum computing power as the representative client of this cluster. The representative client selects the corresponding base model according to its own computing power limitations and deploys the selected base model on the representative client as the cluster model.

[0021] Furthermore, step S2 specifically includes:

[0022] For different computing power limitations, through the knowledge-aware model architecture search algorithm, search for the optimal sub-model of the cluster model for the clients with insufficient computing power and deploy it on the client;

[0023] The knowledge-aware model architecture search algorithm is a depth pruning algorithm based on the genetic algorithm, which will prune the entire transformer blocks. The fitness is calculated through two measurement metrics. One is the NASWOT measurement metric, and the calculation method is as follows:

[0024] S = log|K|, N A is the unit of the activation function, d H represents the Hamming distance, K represents the kernel matrix, and S represents the NASWOT measurement metric;

[0025] The other is the measurement metric of KL divergence, and the calculation method is as follows:

[0026]

[0027] where p is the logits of the original model, q is the logits of the sub-model, T is an adjustable hyperparameter to control the influence between logits, and d represents the measurement metric of KL divergence;

[0028] Then the fitness F, F = S - d;

[0029] The specific search of the knowledge-aware model architecture search algorithm includes:

[0030] Step1: Generate the structures of multiple sub-models;

[0031] Step2: Randomly select two sub-model structures and calculate the fitness of the two sub-model structures. If the fitness of the first sub-model structure is greater than that of the second sub-model structure, then the first sub-model structure is the winner and the second sub-model structure is the loser; otherwise, the second sub-model structure is the winner and the first sub-model structure is the loser;

[0032] Step3: Generate a random number. If it is less than the crossover rate, then perform crossover calculation on the sub-model structures corresponding to the winner and the loser to obtain a new structure. If the random number is less than the mutation rate, then flip the sub-model structure corresponding to the loser to obtain a new structure;

[0033] Step4: Calculate the fitness of the new structure. If the fitness of the new structure is greater than the fitness of the sub-model structure corresponding to the loser, then replace the sub-model structure corresponding to the loser with the new structure;

[0034] Step5: Repeat Step1 to Step4 until the loop is completed.

[0035] Furthermore, Step S3 specifically includes:

[0036] S301. Fine-tune using the private data of the client, and only save the parameters after fine-tuning;

[0037] S302. Upload the parameters of each client to the representative client within the cluster, and then perform parameter aggregation;

[0038] S303. Repeat steps S301 to S302 until the in-cluster training is completed.

[0039] Furthermore, step S4 specifically includes:

[0040] S401. Perform knowledge transfer through the representative client in each cluster. The weight calculation formula for each representative client is as follows:

[0041] where ω m represents the weight of each representative client, M is the number of clusters, N k is the data volume of the representative client, x i is the original data, and y i is the label;

[0042] S402. Use the unlabeled public dataset for knowledge distillation. Use the representative client in each cluster as the teacher model to generate pseudo-labels for the unlabeled data, and calculate the cross-entropy loss through the pseudo-labels and the prediction values of the server model. The calculation formula is as follows:

[0043]

[0044] where represents the unlabeled data, is the pseudo-label generated after passing through θ leader(m) ;

[0045] S403. Transmit the logits calculated by the teacher model through the public dataset to the server, and calculate the KL divergence with the logits calculated by the server model. The calculation formula is as follows:

[0046] D KL represents the KL divergence, and σ represents the activation function;

[0047] S404. Combine the cross-entropy loss and the KL loss to obtain the total loss. Optimize the server model by minimizing the loss, thereby realizing the fine-tuning of the server model. The total loss calculation formula is as follows:

[0048] α represents a hyperparameter that controls the ratio of cross - entropy loss and KL loss.

[0049] Furthermore, step S4 specifically includes:

[0050] S501. After optimizing the server model, use the knowledge distillation method to update the representative model of each cluster through an unlabeled public dataset, and at the same time, the representative model also saves the parameters.

[0051] S502. After the representative model of each cluster is updated, distribute the parameters according to the sub - model structure of each client within the cluster, distribute the parameters corresponding to the sub - model structure, and update the models of each client.

[0052] S503. Execute step S1 until the fine - tuning ends.

[0053] In a second aspect, the present invention provides a computer device, including a memory, where the memory stores program instructions. When the program instructions run, they execute the above - mentioned adaptive fine - tuning method for a clustering - based heterogeneous federated basic model.

[0054] The beneficial effects of the present invention are as follows:

[0055] The present invention utilizes the partial training (PT) method and the knowledge distillation (KD) method, effectively solves the problem of fine - tuning FMs in a highly resource - heterogeneous scenario, enables the models deployed on each client to meet their resource limitations, and significantly reduces the computational overhead and communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flowchart of the adaptive fine - tuning method for a clustering - based heterogeneous federated basic model provided by an embodiment of the present invention;

[0057] Figure 2 is a schematic diagram of aggregating sub - model parameters within a cluster provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0059] The present invention provides an adaptive fine - tuning method for a clustering - based heterogeneous federated basic model, as Figure 1 shown, specifically including:

[0060] S1. Through the MHAC (Multi-Factor Heterogeneous Aware Clustering) module, comprehensively consider the computing power resources and data distribution of each client to perform clustering. Select a representative client (leader node) for each cluster, and the representative client will select a suitable model according to its computing power limit, which is called the cluster model.

[0061] S2. Through the KAMAS (Knowledge-Aware Model Architecture Search) algorithm, according to the heterogeneous computing power limits of each client, search for the optimal sub-model based on the cluster model for all clients within each cluster, and deploy the optimal sub-model on the client.

[0062] S3. After the clients within each cluster complete local training (efficient fine-tuning based on LoRA), they will upload the LoRA parameters to the representative client (leader node). The corresponding parameters will be aggregated on the representative client and then sent to the clients within the cluster. Repeat the above steps until the training within the cluster is completed.

[0063] S4. Through the CAKT (Cluster-Aware Knowledge Transfer) module, transfer the knowledge of each cluster to the server model to achieve the training of the server model.

[0064] S5. Through reverse knowledge distillation, transfer the knowledge of the server model back and update the representative client (leader node) of each cluster. The updated leader node will send the corresponding LoRA parameters to each client within the cluster.

[0065] Specifically, step S1 specifically includes:

[0066] S101: The MHAC module uses the K-means algorithm for clustering, comprehensively considering the computing power and data distribution on the clients. To protect data privacy, we use the differential privacy method to add Gaussian noise to the data of each client. For client i, with data distribution P(D i ), after adding noise, the data feature is Δ f represents the sensitivity of the function, ∈ represents a parameter measuring the strength of privacy protection; C(D i ) represents the computing power limit of the client, and the K-means algorithm comprehensively considers M(D i ) and C(Di ) After that, clients with similar data distributions and computing power limitations are grouped into a cluster.

[0067] S102: For each cluster, according to the computing power limitations of each client within the cluster, the client with the highest computing power is located as the representative client (leader node) of this cluster. The representative client can select a suitable Foundation model (FM) according to its own computing power limitations, such as clip-base, clip-large models, etc. The selected model will be deployed on the leader node as the cluster model.

[0068] Specifically, step S2 includes:

[0069] S201: In highly heterogeneous clients, even within the same cluster, the computing power limitations of each client may be different. Therefore, not every client can deploy the same model (cluster model) as the representative client (leader node) within the cluster. Thus, for different computing power limitations, through the knowledge-aware model architecture search algorithm, this invention searches for the optimal sub model of the cluster model for clients with insufficient computing power and deploys it on this client.

[0070] S202: The KAMAS algorithm is a depth pruning algorithm based on the genetic algorithm, which will perform whole pruning on transformer blocks (layers). There are two measurement metrics, one is the Neural Architecture Search without Training (NASWOT) score and the Kullback-Leibler (KL) Divergence score.

[0071] S = log|K|, which is the calculation method of the NASWOT score, where N A are the units of the activation function, d H represents the Hamming distance. The NASWOT score uses the initial activation pattern of the activation units in the untrained network to predict its final performance. The construction of the kernel matrix K is completed by calculating the Hamming distance between the binary encodings representing the activation states of the input data points (c1, c2,..., c N ) in the linear region of the network. The final NASWOT score S is derived from the logarithm of the absolute value of the determinant of K.

[0072] The other is the measurement metric of KL divergence, and the calculation method is as follows:

[0073]

[0074] Where p are the logits of the original model, q are the logits of the sub - model, and T is an adjustable hyperparameter to control the influence between the logits. Finally, considering two metrics comprehensively, the fitness F can be calculated as F = s - d.

[0075] The specific steps of the search algorithm are as follows:

[0076] Step1: First, generate the structures of multiple sub - models (since the model based on Transformer is composed of multiple Transformer blocks stacked together, the sub - model structure is the expression of which Transformer blocks are selected, for example, [1, 0, 0, 1, 1, 0, 1...], where 1 represents selection and 0 represents non - selection).

[0077] Step2: Randomly select two sub - model structures A and B, calculate the fitness, FA and FB. If FA > FB, then A is the winner and B is the loser, and vice versa.

[0078] Step3: Generate a random number. If it is less than the crossover rate, then perform crossover calculation on the structures of the winner and the loser to obtain a new structure. If the random number is less than the mutation rate, then flip the structure of the loser to obtain a new structure.

[0079] Step4: Recalculate the fitness of the updated structure. If the new fitness is greater than the fitness of the loser, then replace the loser structure in the population with this new structure.

[0080] Step5: Repeat Step1 to Step4 until the loop is completed.

[0081] Specifically, step S3 specifically includes:

[0082] S301: The client performs efficient fine - tuning (using LoRA) with its own private data, and only save the LoRA parameters after fine - tuning.

[0083] S302: Upload the LoRA parameters of each client to the representative client (leader node) within the cluster, and then complete parameter aggregation here, as Figure 2 shown, and redistribute them to each client.

[0084] S303: Repeat steps S301 to S302 until the in - cluster training is completed.

[0085] Specifically, step S4 specifically includes:

[0086] S401: Knowledge transfer is carried out through the representative clients (leader nodes) within each cluster. Since the capabilities provided by each representative client are different, in order to ensure that the server model obtains more accurate knowledge, the present invention utilizes the accurately trained ones within each cluster to design a cluster-aware knowledge transfer module. The weight calculation formula for each representative client is as follows:

[0087] where M is the number of clusters, N k is the data volume of the representative client, x i is the original data, y i is the label;

[0088] S402: Use an unlabeled public dataset for knowledge distillation. Utilize the representative clients (leader nodes) within each cluster as the teacher model to generate pseudo-labels for the unlabeled data. Calculate the cross-entropy loss through the pseudo-labels and the predicted values of the server model. The calculation formula is as follows

[0089]

[0090] where, represents the unlabeled data, is the pseudo-label generated after leader(m) θ;

[0091] S403: Transmit the logits calculated by the teacher model through the public dataset to the server, and calculate the KL divergence with the logits calculated by the server model. The calculation formula is as follows:

[0092] D KL represents the KL divergence, and σ represents the activation function;

[0093] S404: Combine the cross-entropy loss and the KL loss to obtain the total Loss. Finally, optimize the server model by minimizing the Loss, thereby realizing the fine-tuning of the server model. The calculation formula for the total Loss is as follows:

[0094] α represents a hyperparameter that controls the ratio of the cross-entropy loss and the KL loss.

[0095] Specifically, step S5 includes:

[0096] S501: After updating the server model, use the knowledge distillation method to update the representative model (leader node) of each cluster through the untagged public dataset. At the same time, the model of the leader node only needs to save the LoRA parameters.

[0097] S502: After the leader node of each cluster is updated, distribute the LoRA parameters according to the sub-model structure of each client within the cluster, distribute the parameters corresponding to the sub-model structure, and update the models of each client.

[0098] S503: After all clients are updated, it means that a round of overall training is completed. Restart step S1 until the fine-tuning ends.

[0099] Compared with the traditional solution, the performance of the present invention significantly exceeds the existing solutions. In a large number of experiments, the present invention has achieved significant improvement. Specifically, in the CIFAR10, CIFAR100, and Tiny-ImageNet datasets, FedCAMS has improved the image classification accuracy by 3-10% compared with other baseline methods. At the same time, compared with the baseline method based on partial training, the communication overhead is significantly reduced, and the communication overhead is almost negligible. By using the LoRA efficient fine-tuning method, compared with full-scale fine-tuning, the trainable parameters are only about 1% of the original, and the computational overhead is also significantly reduced.

[0100] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention should be within the protection scope of the appended claims of the present invention.

Claims

1. A clustering-based adaptive fine-tuning method for heterogeneous federated foundation models, characterized in that Including: S1. Through the multi-factor heterogeneous perception clustering module, comprehensively consider the computing power resource limitations and data distribution of each client for clustering, select a representative client for each cluster, and the representative client will select a corresponding model as the cluster model according to its own computing power limitations; S2. Through the knowledge-aware model architecture search algorithm, according to the heterogeneous computing power limitations of each client, search for the optimal sub-model based on the cluster model for all clients within each cluster, and deploy the optimal sub-model on the client; S3. Conduct local training on the clients within each cluster, then upload the parameters to the representative client, aggregate the corresponding parameters on the representative client, and distribute them to the clients within the cluster after aggregation. Repeat step S3 until the training within the cluster is completed; S4. Through the cluster-aware knowledge transfer module, transfer the knowledge of each cluster to the server model to realize the training of the server model; S5. Through reverse knowledge distillation, transfer the knowledge of the server model back and update the representative client of each cluster. The updated representative client will distribute the corresponding parameters to each client within the cluster.

2. The method for adaptive fine-tuning of a heterogeneous federated base model based on clustering according to claim 1, wherein Step S1 specifically includes: The multi-factor heterogeneous perception clustering module uses the K-means algorithm, comprehensively considers the computing power resource limitations and data distribution of each client, and divides the clients with similar data distribution and computing power limitations into one cluster; When clustering, the differential privacy method is adopted to add Gaussian noise to the data of each client. For client i, its data distribution is P(D i ), after adding Gaussian noise, the data feature is: Δ f represents the sensitivity of the function, and ∈ represents a parameter that measures the strength of privacy protection; For each cluster, according to the computing power limitations of each client within the cluster, select the client with the highest computing power as the representative client of this cluster. The representative client selects the corresponding basic model according to its own computing power limitations, and deploys the selected basic model on the representative client as the cluster model.

3. The method for adaptive fine-tuning of a heterogeneous federated foundation model based on clustering according to claim 1, wherein Step S2 specifically includes: For different computing power limitations, through the knowledge-aware model architecture search algorithm, search for the optimal sub-model of the cluster model for the clients with insufficient computing power and deploy it on the client; The knowledge-aware model architecture search algorithm is a depth pruning algorithm based on the genetic algorithm, which will prune the entire transformer blocks, and calculate the fitness through two measurement indicators. One is the measurement indicator of NASWOT, and the calculation method is as follows: S = log|K|, N A is the unit of the activation function, d H represents the Hamming distance, K represents the kernel matrix, and S represents the NASWOT metric; The other is the measurement indicator of KL divergence, and the calculation method is as follows: Where p is the logits of the original model, q is the logits of the sub-model, T is an adjustable hyperparameter to control the influence between logits, and d represents the measurement indicator of KL divergence; Then the fitness F, F = S - d; The specific search of the knowledge-aware model architecture search algorithm includes: Step1: Generate the structures of multiple sub-models; Step2: Randomly select two sub-model structures and calculate the fitness of the two sub-model structures. If the fitness of the first sub-model structure is greater than that of the second sub-model structure, the first sub-model structure is the winner and the second sub-model structure is the loser; otherwise, the second sub-model structure is the winner and the first sub-model structure is the loser; Step 3: Generate a random number. If it is less than the crossover rate, perform crossover calculation on the sub-model structures corresponding to the winner and the loser to obtain a new structure. If the random number is less than the mutation rate, flip the sub-model structure corresponding to the loser to obtain a new structure; Step 4: Calculate the fitness of the new structure. If the fitness of the new structure is greater than the fitness of the sub-model structure corresponding to the loser, replace the sub-model structure corresponding to the loser with the new structure; Step 5: Repeat Step 1 to Step 4 until the loop is completed.

4. The method for adaptive fine-tuning of a heterogeneous federated foundation model based on clustering according to claim 1, wherein Step S3 specifically includes: S301. Fine-tune using the private data of the client, and only save the parameters after fine-tuning; S302. Upload the parameters of each client to the representative client within the cluster, and then perform parameter aggregation; S303. Repeat steps S301 to S302 until the in-cluster training is completed.

5. The method for adaptively fine-tuning a heterogeneous federated basic model based on clustering according to claim 1, characterized in that, Step S4 specifically includes: S401. Perform knowledge transfer through the representative client in each cluster. The weight calculation formula for each representative client is as follows: where ω m represents the weight of each representative client, M is the number of clusters, and N k is the amount of data of the representative client, x i is the original data, and y i is the label; S402. Use the unlabeled public dataset for knowledge distillation. Use the representative client in each cluster as the teacher model to generate pseudo-labels for the unlabeled data, and calculate the cross-entropy loss through the pseudo-labels and the predicted values of the server model. The calculation formula is as follows: Among them, represents unlabeled data, is the pseudo-label generated after θ leader(m) ; S403. Transmit the logits calculated by the teacher model through the public dataset to the server, and calculate the KL divergence with the logits calculated by the server model. The calculation formula is as follows: D KL is the representation of the KL divergence, and σ represents the activation function; S404. Combine the cross-entropy loss and the KL loss to obtain the total Loss, and optimize the server model by minimizing the Loss, thereby realizing the fine-tuning of the server model. The total Loss calculation formula is as follows: α represents a hyperparameter that controls the ratio of the cross-entropy loss and the KL loss.

6. The method for adaptive fine-tuning of a heterogeneous federated foundation model based on clustering according to claim 1, wherein Step S4 specifically includes: S501. After optimizing the server model, use the knowledge distillation method to update the representative model of each cluster through the unlabeled public dataset, and the representative model also saves the parameters; S502. After the representative model of each cluster is updated, distribute the parameters according to the sub-model structure of each client within the cluster, distribute the parameters corresponding to the sub-model structure, and update the models of each client; S503. Execute Step 1 until the fine-tuning ends.

7. A computer device, including a memory, the memory storing program instructions, characterized in that, When the program instructions run, execute the clustering-based heterogeneous federated basic model adaptive fine-tuning method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image classification method based on federal knowledge distillation and ensemble learning

    CN117523291A

  • Self-adaptive clustering federal learning method with differential privacy

    CN117634594A