Method for performing adaptive fine tuning on LLM based on federated learning

By adopting an adaptive fine-tuning method in the federated learning framework, dynamically adjusting the client model configuration and optimizing the model hierarchy sequence, the problems of insufficient data feature utilization and lagging client in LLM fine-tuning are solved, and efficient and personalized model adaptation and performance improvement are achieved.

CN120124778APending Publication Date: 2025-06-10CHONGQING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510288252.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When fine-tuning large language models (LLM), the existing technology lacks the ability to deeply analyze and utilize client-specific data characteristics, resulting in insufficient adaptability of the model to each client task and overall performance decline.

Method used

Adaptive fine-tuning method based on federated learning is adopted to dynamically adjust the client model configuration through the two-way connection between the server and multiple clients, eliminate the contribution of the lagging client, and optimize the model hierarchy order through dynamic rotation strategy of semantic similarity to improve the personalized adaptation performance of the model.

Benefits of technology

It significantly reduces the computing and communication overhead, improves the personalized adaptation performance of the model, optimizes the impact of lagging clients on global model training, improves the convergence efficiency and stability of the system, and reduces semantic conflicts caused by hierarchical differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124778A_ABST
    Figure CN120124778A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and particularly discloses a method for performing adaptive fine tuning on LLM based on federated learning, which optimizes a global network model based on federated learning, and enables model design to be highly matched with task requirements by analyzing client performance feedback and adaptively and dynamically adjusting client model configuration by a server. The calculation and communication overhead is obviously reduced, and meanwhile, the personalized adaptation performance is improved; the server performs fine-grained evaluation on client feedback in each iteration, rejects lagged clients which are not trained enough, and supplements a part of missing model layers of the clients by using a historical optimal weight; during model aggregation, a dynamic rotation strategy based on semantic similarity is adopted, and semantic conflicts caused by hierarchy differences are reduced by intelligently adjusting the sequence of client model layers. According to the method, the training efficiency and accuracy of the model are remarkably improved through self-adaptive adjustment of the client model, synchronization of the server-client model and a dynamic rotation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method for adaptively fine-tuning an LLM based on federated learning. Background Art

[0002] In recent years, large language models (LLMs) have made breakthroughs in natural language processing (NLP) tasks. These models are first pre-trained on large-scale corpora and then fine-tuned for specific tasks, thus significantly improving performance. LLMs have reached the state-of-the-art level in various applications, including natural language understanding, sentiment classification, and question answering systems. The success of LLMs is largely attributed to their ability to capture and utilize rich language information during training. However, training these large models from scratch requires a large amount of computing resources and storage space. For most organizations and individuals, this is impractical. In addition, centralized fine-tuning requires aggregating all data on a single device or server. It raises serious data privacy issues and complicates data protection. Therefore, distributed fine-tuning of pre-trained models for specific tasks has become the main approach. It can utilize the computing resources of edge devices while ensuring the confidentiality of device data. Federated Learning (FL), as a privacy-preserving distributed machine learning method, has become an effective solution. Each client in FL will regularly calculate and send model parameters or gradients to the central server. The server aggregates these updates to generate a new global model.

[0003] LLMs have complex architectures and high computing requirements, which pose many challenges to personalized adjustment. The current methods generally have the following two limitations:

[0004] 1. Insufficient exploration and utilization of data features

[0005] Existing methods lack the ability to deeply analyze and fully utilize client-specific data features. LLMs require highly complex feature extraction and modeling, while existing methods usually adopt a unified fixed strategy during personalized adjustment, ignoring the significant differences between the data distributions and features of different clients. This deficiency directly limits the model's adaptability to client tasks, resulting in a decline in overall performance.

[0006] 2. Lagging client problem

[0007] In distributed training, due to differences in communication bandwidth and computing resources, some clients may not be able to fully train their local models. These lagging clients provide limited information during the global model aggregation process and may even introduce noise, reducing the convergence efficiency and stability of the global model.

[0008] 3. Weight conflicts among different clients

[0009] In the federated learning framework, clients achieve local optimization through feature sharing. However, there are often significant differences in the model structures among different clients, and direct parameter aggregation may lead to semantic conflicts, thus affecting the performance of the global model. Summary of the Invention

[0010] The present invention provides a method for adaptively fine-tuning an LLM based on federated learning, and the technical problem to be solved is: how to perform high-performance fine-tuning on the LLM model.

[0011] To solve the above technical problems, the present invention provides a method for adaptively fine-tuning an LLM based on federated learning, including the steps of:

[0012] Establish a two-way connection between the server and N clients, where the server is configured with a global network model of the LLM model, the LLM model is the large language model, and the initial global model parameters of the global network model are w 0 ;

[0013] Construct a dataset and evenly distribute the dataset to N ≥ 3 clients;

[0014] Enter the iteration. In the t-th iteration, the server prunes the first L 1t , L 2t to L Nt model layers and the corresponding global model parameters from the current global network model and distributes the sub-models to N clients. After the N clients train the received sub-models using their own local datasets, they send the new local model parameters and the best F1 score to the server; the server hierarchically aggregates the local model parameters sent by each client to obtain new global model parameters w t ; The server also calculates the average F1 score of all clients based on the best F1 score sent by each client, and determines whether the capacity of each client has reached saturation according to the average F1 score of the current iteration, the average F1 score of the previous iteration, the best F1 score of each client, and the resource constraints. If it has reached saturation, the sub-model structure of the corresponding client remains unchanged and the new global model parameters w t are sent to the corresponding client. If it is not saturated, one more layer is cut from the global network model as a new sub-model and the new global model parameters w tSend it to the corresponding client and enter the next iteration;

[0015] End after T iterations, and the server obtains the global model parameters w of its global network model T 。

[0016] Further, determining whether the capacity of each client reaches saturation according to the average F1 score of the current iteration, the average F1 score of the previous iteration, the best F1 score of each client, and resource constraints is specifically as follows:

[0017] If the difference between the F1 score of the client in this iteration and the average F1 score of this iteration is less than or equal to the difference between the average F1 score of this iteration and the average F1 score of the previous iteration, and the ratio of the difference between the F1 score of the client in this iteration and the average F1 score of this iteration to the difference between the average F1 score of this iteration and the average F1 score of the previous iteration is less than or equal to the preset value δ, and the number of local model layers of the client does not reach the maximum number of encoder layers that the client can support, it is considered that the capacity of the client is not saturated.

[0018] Further, in the initial training iteration, the initial global model parameters of the first α layers of the global network model are used to initialize the parameters of the local network models of each client participating in the training.

[0019] Further, determine the size of α for each client according to the resource constraints, task types, and initial global network parameters of different clients.

[0020] Further, for the new global model parameters w t , the global model parameters of its l-th layer θ l represents the global model parameters of the l-th layer before this aggregation, |S l | represents the number of clients in the set S l , represents the set S l the parameter update amount of the l-th layer local model of the i-th client in, and the set S l includes the indices of all clients with l-th layer model parameters.

[0021] Further, after the server receives the parameter updates of all clients, it calculates the training completion degree of each client, and applies a weight α l to the parameter update amount of each layer of the local model of the i-th client in the set S i , and the training completion degree is defined as the ratio of the number of local training rounds to the preset maximum number of rounds.

[0022] Further, clients with a training completion rate lower than the threshold θ are marked as lagging clients, and the weights of the lagging clients are replaced with the historical best global weights of this layer.

[0023] Further, the server hierarchically aggregates the local model parameters sent by each client to obtain new global model parameters w t , specifically including the steps:

[0024] Calculate the semantic similarity between each layer of the client model and each layer of the global model;

[0025] Based on the calculated semantic similarity, find the global model layer that is semantically most similar to each layer of each client model, and dynamically adjust the hierarchical order of the client models to make it fit the structure of the global model;

[0026] Aggregate the client model parameters participating in the training by the method of weighted average to obtain new global model parameters w t .

[0027] Further, the dataset uses the Mental Health Sentiment Analysis MHSA dataset;

[0028] The global network model uses the LLAMA3 model.

[0029] Further, the local model of each client is trained using the cross-entropy loss function, the network parameters are optimized using the AdamW optimizer, and the learning rate is dynamically adjusted through the ExponentialLR scheduler.

[0030] A method for adaptively fine-tuning an LLM based on federated learning provided by the present invention has the beneficial effects that:

[0031] During the training process, the server adaptively and dynamically adjusts the client model configuration by analyzing the client performance feedback, making the model design highly match the task requirements, significantly reducing the computational and communication overhead, and at the same time improving the personalized adaptation performance;

[0032] The server conducts a fine-grained evaluation of the client feedback in each iteration, eliminates the contributions of under-trained lagging clients, thereby optimizing the global model; and for the missing model layers of some clients, the server uses the historical best global weights for supplementation;

[0033] During model aggregation, the present invention proposes a dynamic rotation strategy based on semantic similarity. This strategy intelligently adjusts the order of the client model layers to better align with the global model, reducing semantic conflicts caused by hierarchical differences, especially applicable to model differences in resource-constrained clients. By calculating the semantic similarity between each client model layer and the global model layer, the order of the client model layers is intelligently adjusted to optimize the aggregation effect of the global model. It can not only effectively reduce semantic conflicts but also improve the accuracy and stability of the model in the case of large hierarchical differences.

[0034] Overall, the present invention significantly improves the training efficiency and accuracy of the model through client model adaptive adjustment, server-client model synchronization, and dynamic rotation strategy. Compared with existing methods, it effectively reduces computational and communication costs; significantly enhances the personalization and adaptability of model performance; optimizes the impact of lagging clients on the global model training, improving the system convergence efficiency and stability; and reduces semantic conflicts caused by hierarchical differences through the dynamic rotation strategy based on semantic similarity. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is an example diagram of a method for adaptively fine-tuning an LLM based on federated learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0036] The embodiments of the present invention will be specifically described below in conjunction with the drawings. The examples are given only for illustrative purposes and should not be construed as limiting the present invention. The drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from its spirit and scope.

[0037] A method for adaptively fine-tuning an LLM based on federated learning provided by an embodiment of the present invention includes the steps of:

[0038] Establish a two-way connection between the server and N clients, where the server is configured with a global network model of the LLM, and the initial global model parameters of the global network model are w 0 ;

[0039] Construct a dataset and evenly distribute the dataset to N≥3 clients;

[0040] Enter the iteration. In the t-th iteration, the server prunes from the current global network model respectively, including the first L 1t , L 2t to L NtThe sub-models of a model layer and the corresponding global model parameters are distributed to N clients. After the N clients train the received sub-models using their own local datasets, they send the new local model parameters and the best F1 score to the server. The server hierarchically aggregates the local model parameters sent by each client to obtain the new global model parameter w t ; The server also calculates the average F1 score of all clients based on the best F1 scores sent by each client, and determines whether the capacity of each client has reached saturation according to the average F1 score of the current iteration, the average F1 score of the previous iteration, the best F1 scores of each client, and resource constraints. If it has reached saturation, it keeps the sub-model structure of the corresponding client unchanged and sends the new global model parameter w t to the corresponding client. If it is not saturated, it cuts one more layer from the global network model as a new sub-model and sends the new global model parameter w t to the corresponding client, and enters the next iteration;

[0041] After T iterations, it ends. The server obtains the global model parameter w of its global network model T .

[0042] The present invention considers how to extract client sub-models from two aspects. First, due to limited client resources, the model cannot be infinitely expanded. The client can only train a model that can match its device resources. Second, the model performance of the client is used to adjust the local model structure of the client. If increasing the number of model layers does not significantly improve the performance, it indicates that the current number of layers is sufficient to adapt to the client data, and the model can maintain good performance without increasing the number of layers. On the contrary, if a significant increase in the number of layers improves the model performance, it indicates that the capacity of the model exceeds the data size of the client, and the client will increase the number of layers in its model layer by layer.

[0043] Each client only trains the sub-model extracted from the global server model, and sends the updated parameters and F1 scores of the corresponding sub-model back to the server. The server aggregates and updates according to these new parameters to obtain the global model parameters participating in the next iteration, and determines whether to increase the level of the corresponding client in the next iteration according to the F1 score. In each iteration, the server extracts specific sub-models tailored to each client from the global model according to the resource limitations and local model performance of each client. Then, these extracted sub-models are distributed to each client.

[0044] The initial global network parameters of the global network model with L model layers are represented as w 0 ={θ 1 ,θ 2 ,...,θ L}, θ ldenote the initial global network parameters of its l-th layer model, where l = 1, 2, …, L. In the initial training iteration (t = 1), the initial global model parameters {θ 1 ,...,θ α} of the first α layers of the global network model are used to initialize the parameters of the local network models of each client C i (i = 1, 2, …, N). That is The client uses its local data to train the received sub-model and sends the updated parameters of its heterogeneous sub-model back to the server. The server aggregates the received parameters layer by layer to obtain the new global model parameters w 1 to participate in the second iteration.

[0045] It should be noted that in the initial iteration, the size of α for each client can be determined according to the resource constraints, task types, and initial global network parameters of different clients.

[0046] In each subsequent iteration t, each client C i calculates the best F1 score for this iteration and uploads it to the server together with the new local model update. The server dynamically adjusts the local model of client C i according to the performance of the local models trained by the client historically. Specifically, in iteration t, the server evaluates the F1 score of client C i and calculates the average F1 score AvgF1 of all clients in this iteration . According to the F1 score t of client C i and the average F1 score AvgF1 , the server dynamically adjusts the model parameters of client C t in the next iteration t + 1 i .

[0047]

[0048] where l i is the number of layers of the current model of client C i .

[0049] The above formula is: If the difference between the F1 score i of client C in this iteration and the average F1 score AvgF1 of this iteration t is less than or equal to the difference between the average F1 score AvgF1 of this iteration t and the average F1 score AvgF1 of the previous iteration t-1 , and client C in this iterationi The F1 score F1 i t and the average F1 score AvgF1 of this iteration t The difference between them and the average F1 score AvgF1 of this iteration t and the average F1 score AvgF1 of the previous iteration t-1 If the ratio of the difference is less than or equal to the preset value δ, it is considered that the capacity of the client is not saturated. Then, in the next iteration, a model layer (the (l i +1)-th layer) is added to the current submodel, and a submodel with l i +1 model layers is obtained and sent to the client C i . Otherwise, the model level is not increased, and the global model parameters of the new l i -th layer are directly sent to the client C i to participate in the next iteration.

[0050] Meanwhile, due to the resource constraints of the client C i , the number of local model layers l i of the client C i cannot exceed the maximum number of encoder layers h i that the client C i can support, that is:

[0051]

[0052] where l i is the number of layers of the current model of the client C i , h i is the maximum number of encoder layers that the client C i can accommodate, r i is the resource constraint of the client C i , and h i is calculated according to r i .

[0053] The present invention adopts a direct selective averaging scheme without client weighting to aggregate heterogeneous submodel updates sent from clients participating in training to update the global model.

[0054] When the server extracts a submodel for each client, it updates the set of clients participating in the global aggregation, where the set of clients participating in the aggregation of the global parameters of the l-th layer is denoted as S i∈{1,...,L} . Each set S l has the indices of the clients with the model parameters of the l-th encoder layer, that is:

[0055]

[0056] The above formula indicates that the set Sl Include the indices of all clients with the model parameters of the l-th layer.

[0057] In iteration t, the server first collects the latest model updates submitted by each client during this iteration and splits the model parameters of the client updates according to the encoder hierarchy. Subsequently, the server updates the parameters of the global model layer by layer according to the structure of the encoder. This process involves averaging the updates using the client model parameters containing the updates of each encoder layer. If no client updates the parameters of a specific layer, the corresponding global model parameters of that layer remain unchanged. The aggregated global model parameters of the l-th layer can be formulated as follows:

[0058]

[0059] where, θ l represents the global model parameters of the l-th layer before this aggregation, |S l | represents the number of clients in the set S l , represents the parameter update amount of the l-th layer of the local model of the i-th client in the set S l .

[0060] Furthermore, after receiving the parameter updates from all clients, the server calculates the training completion degree O i of each client. The training completion degree O i is defined as the ratio of the number of local training rounds to the preset maximum number of rounds. Clients with a value lower than the threshold θ are marked as lagging clients.

[0061] Assign different weights to the parameters of different clients according to the training completion degree:

[0062]

[0063] where, α i represents the weight of the i-th client.

[0064] Then, when performing global aggregation, apply the weight α l to the parameter update amount of each layer of the local model of the i-th client in the set S i .

[0065] For the weights of each layer of the model: Since the training completion degree of lagging clients is not high, directly use the historical best global weights of this layer to replace them and participate in global aggregation.

[0066] Calculate semantic similarity: To effectively match the hierarchy of the client model with that of the global model, it is first necessary to calculate the semantic similarity of each layer. Assume that the i-th layer of the client model is represented as and the k-th layer of the global model is represented as Then, the semantic similarity can be calculated through cosine similarity:

[0067]

[0068] Among them, represents the semantic similarity between the client layer i and the global layer k, represent the norms of the client layer i and the global layer k respectively. The similarity value obtained through calculation can reflect the semantic similarity between each layer, providing a basis for subsequent hierarchical matching.

[0069] Hierarchical matching: After calculating the semantic similarity, the next step is to perform hierarchical matching according to the principle of maximum similarity. For each client C i , the present invention selects the global model layer k with the largest semantic similarity to it * as the best matching layer. The specific matching rule is:

[0070]

[0071] Through this matching strategy, the present invention finds the global model layer with the most similar semantics to each layer of the client model, thereby dynamically adjusting the hierarchical order of the client model to make it more compatible with the structure of the global model and reducing semantic conflicts.

[0072] Dynamic FedAvg aggregation: After completing the hierarchical matching, the third step is to perform dynamic FedAvg aggregation. According to the results of the hierarchical matching, each layer of the client model will participate in the update process of the global model. In each round of training, the server updates the corresponding parameters according to the global layer matched by the client and aggregates the model parameters by the method of weighted average. Specifically, for each layer the present invention collects the parameters of this layer from all clients participating in the training and then performs weighted average on the parameters according to the weight λ of each client i :

[0073]

[0074] Among them, N is the number of clients participating in the training, and λ i is the weight of client i in this layer, which is usually dynamically adjusted according to the data volume or performance index of the client. Through this weighted average aggregation method, it can ensure that the global model maintains its architecture consistency after each round of training and at the same time absorbs the update information of the client model.

[0075] Figure 1 is an example of this method, Figure 1Examples include that the global model has four model layers, a simple task only requires one layer of structure, a normal task requires two layers of structure, a difficult task requires three layers of structure, and the model level of each client at the initial iteration is 1. In practical applications, it is designed according to the specific global model architecture, the task difficulty of the client, the resource limitations of the client, etc.

[0076] The following is an application example of this method.

[0077] Dataset: This example uses the Mental Health Sentiment Analysis (MHSA) dataset, which contains annotated text from social media platforms covering seven common mental health states. The dataset is divided into a training set, a validation set, and a test set to ensure the integrity of the test set. To simulate a real federated learning environment, 20 local clients are designed, and the dataset is evenly distributed to each client.

[0078] Training parameters: The global network model uses the LLAMA3 model. The local model is trained using the cross-entropy loss function, and the AdamW optimizer is used to optimize the network parameters. The training batch size is 64. The learning rate is dynamically adjusted by the ExponentialLR scheduler, and the gamma value is set to 0.9. The initial learning rate of the local client is dynamically adjusted according to the number of layers in the self-attention layer of the model. The local client performs 10 epochs during training, and the best parameters generated in each epoch are transmitted to the server. The server model is updated, and the updated global model is distributed back to the client. The global aggregation process iterates 10 rounds, and the model performance is gradually improved by integrating the knowledge of all clients.

[0079] To effectively verify the effectiveness of the method of the present invention, the present invention compares FedAFL with the following methods:

[0080] LLAMA3-8: A lightweight variant of LLAMA3, containing 8 layers of neural networks, suitable for applications with limited computing resources.

[0081] LLAMA3-10: A medium-depth variant of LLAMA3, containing 10 layers of neural networks, providing a strong balance between performance and efficiency.

[0082] LLAMA3-12: A deep variant of LLAMA3, containing 12 layers of neural networks, suitable for high-performance natural language processing tasks.

[0083] FEDBFPT: A model optimized for federated learning, improving resource efficiency and performance through decentralized training.

[0084] Training results: As shown in Table 1 below.

[0085] Table 1

[0086] model Communication overhead (s) Computation overhead (s) F1-Score LLAMA3-8 3143.56 1923.16 0.797 LLAMA3-10 3985.31 2141.28 0.810 LLAMA3-12 4897.08 2359.39 0.812 FEDBFPT 3588.46 2331.84 0.708 FedAFL 3221.94 2093.67 0.838

[0087] As can be seen from Table 1, in terms of communication efficiency and computational overhead, the communication overhead of FedAFL is 3221.94 seconds and the computational overhead is 2093.67 seconds, which is better than methods such as LLAMA3-10 / 12 and FEDBFPT, and is only slightly higher than LLAMA3-8. In terms of model performance, FedAFL achieves the highest F1-Score (0.838), surpassing all comparison methods. By optimizing communication, computation, and model performance, FedAFL overcomes the limitations of traditional methods and has the potential for application in edge computing and privacy-sensitive scenarios. From the training results, it can be seen that this method effectively reduces the computational overhead and communication overhead while ensuring performance.

[0088] In summary, a method for adaptive fine-tuning of LLM based on federated learning provided by the present invention realizes:

[0089] To solve the problem of client resource differences, the present invention introduces a model adaptation strategy to optimize computational and communication costs. From the perspectives of modeling and data, the present invention proposes to dynamically adjust the model layers according to the complexity of each client dataset. In the initial stage, the client initializes the local model with the weights of the first L layers of the global model and returns the local training results to the server. The server adjusts the current client model configuration by analyzing the dynamic relationship between the client model performance and the global model. This effectively solves the consistency between the client task requirements and the model design, thereby minimizing the computational and communication overhead.

[0090] To solve the "client lag" problem in the server aggregation process, the present invention proposes a heterogeneous model aggregation method that optimizes model weights within a multi-layer framework. Specifically, the server performs a fine-grained evaluation of the client feedback in each iteration, excluding the under-trained "laggard" knowledge in the aggregated model. Next, the missing layers in the client sub-model are supplemented with the historical best global weights. The hierarchical order of the sub-models is dynamically adjusted based on the principle of maximum similarity to ensure better alignment with the global model. Finally, the sub-models are fused to ensure adaptive updates while maintaining synchronization between clients. This method establishes an effective connection between clients, enhancing the stability and convergence of the system.

[0091] To solve the problem that in the federated learning framework, there are often significant differences in the model structures between different clients, and direct parameter aggregation may lead to semantic conflicts, thereby affecting the performance of the global model. The present invention proposes a dynamic rotation strategy based on semantic similarity, which intelligently adjusts the order of the client model layers to better align with the global model, reducing semantic conflicts caused by hierarchical differences, especially applicable to model differences of resource-constrained clients.

[0092] Overall, the present invention significantly improves the training efficiency and accuracy of the model through client model adaptive adjustment, server-client model synchronization, and dynamic rotation strategy. Compared with existing methods, it effectively reduces the computational and communication costs; significantly enhances the personalization and adaptability of the model performance; optimizes the impact of lagging clients on the global model training, and improves the system convergence efficiency and stability; through the dynamic rotation strategy based on semantic similarity, it reduces the semantic conflicts caused by hierarchical differences.

[0093] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for adaptive fine-tuning of LLM based on federated learning, characterized in that: Includes steps: A two-way connection is established between the server and N clients, where the server is configured with a global network model of the LLM model, which is a large language model. The initial global model parameter of the global network model is w 0 ; Build a data set and evenly distribute it to N ≥ 3 clients; Entering the iteration, in the tth iteration, the server prunes the first L from the current global network model 1t , L 2t To L Nt The model layers and the sub-models corresponding to the global model parameters are distributed to N clients. After the N clients use their own local data sets to train the received sub-models, they obtain new local model parameters and the best F1 score and send them to the server. The server aggregates the local model parameters sent by each client hierarchically to obtain the new global model parameters w t The server also calculates the average F1 score of all clients based on the best F1 score sent by each client, and determines whether the capacity of each client has reached saturation based on the average F1 score of the current iteration, the average F1 score of the previous iteration, the best F1 score of each client, and resource constraints. If saturation has been reached, the sub-model structure of the corresponding client remains unchanged and the new global model parameter w is set. t Send it to the corresponding client. If it is not saturated, cut one more layer from the global network model as a new sub-model and set the new global model parameter w t Send to the corresponding client and enter the next iteration; After T iterations, the server obtains the global model parameter w of its global network model. T .

2. The method for adaptively fine-tuning LLM based on federated learning according to claim 1, characterized in that: The method of judging whether the capacity of each client is saturated is as follows: If the difference between the client's F1 score in this iteration and the average F1 score in this iteration is less than or equal to the difference between the average F1 score in this iteration and the average F1 score in the previous iteration, and the ratio of the difference between the client's F1 score in this iteration and the average F1 score in this iteration to the difference between the average F1 score in this iteration and the average F1 score in the previous iteration is less than or equal to the preset value δ, and the number of local model layers of the client has not reached the maximum number of encoder layers that the client can support, then the capacity of the client is considered to be not saturated.

3. The method for adaptively fine-tuning LLM based on federated learning according to claim 2, characterized in that: In the initial training iteration, the initial global model parameters of the first α layers of the global network model are used to initialize the parameters of the local network model of each client participating in the training.

4. The method for adaptively fine-tuning LLM based on federated learning according to claim 3, characterized in that: The size of α for each client is determined according to the resource constraints, task types and initial global network parameters of different clients.

5. The method for adaptively fine-tuning LLM based on federated learning according to claim 4, characterized in that: For the new global model parameter w t , the global model parameters of its lth layer θ l represents the global model parameters of the lth layer before this aggregation, |S l | represents the set S l The number of clients in ,Δθ i l Represents the set S l The parameter update amount of the local model of the lth layer of the i-th client in the set S l Contains the index of all clients that have model parameters at level l.

6. The method for adaptively fine-tuning LLM based on federated learning according to claim 5, characterized in that: After receiving the parameter updates from all clients, the server calculates the training completion of each client and aggregates it into the set S l The parameter update amount of each layer of the local model of the i-th client is weighted by α i ,The training completion is defined as the ratio of the number of local training rounds to the preset maximum number of rounds.

7. The method for adaptively fine-tuning LLM based on federated learning according to claim 6, characterized in that: Clients whose training completion is lower than the threshold θ are marked as lagging clients, and the weights of the lagging clients are replaced with the historical best global weights of the layer.

8. The method for adaptively fine-tuning LLM based on federated learning according to claim 7, characterized in that: The server aggregates the local model parameters sent by each client hierarchically to obtain the new global model parameters w t , specifically including the steps: Compute the semantic similarity between each layer of the client model and each layer of the global model; Based on the calculated semantic similarity, find the global model layer with the most similar semantics for each layer of each client model, and dynamically adjust the hierarchical order of the client model to make it fit the structure of the global model; The client model parameters participating in the training are aggregated by weighted averaging to obtain the new global model parameter w t .

9. The method for adaptively fine-tuning LLM based on federated learning according to any one of claims 1 to 8, characterized in that: The dataset adopts the Mental Health Sentiment Analysis MHSA dataset; The global network model adopts the LLAMA3 model.

10. The method for adaptively fine-tuning LLM based on federated learning according to claim 9, characterized in that: The local model of each client is trained using the cross entropy loss function, the network parameters are optimized using the AdamW optimizer, and the learning rate is dynamically adjusted through the ExponentialLR scheduler.

Citation Information

Cited By

  • Big language model data processing method and device based on federal learning

    CN120597970A