Fine tuning method and system for federal large language model based on zero-order optimization

By deploying large language models in the federated learning framework, performing low-rank decomposition and zero-order fine-tuning, the problems of high memory requirements and insufficient personalized strategies during fine-tuning of large language models in the existing technology are solved, and efficient model updates and adaptability improvements are achieved.

CN119939392AInactive Publication Date: 2025-05-06SUN YAT SEN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510093795.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When fine-tuning large language models in the prior art, memory requirements are high and personalized strategies are insufficient, making it difficult to effectively implement with limited client computing capabilities.

Method used

The federated large language model fine-tuning method based on zero-order optimization is adopted. By deploying the large language model into the federated learning framework, low-rank decomposition is performed to reduce the usage of model parameters, and the model is updated using zero-order fine-tuning technology.

Benefits of technology

It significantly reduces memory requirements, improves the model's adaptability and performance on different client data, and solves the problem of high memory requirements for fine-tuning of large language models in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939392A_ABST
    Figure CN119939392A_ABST
Patent Text Reader

Abstract

The invention discloses a federal large language model fine tuning method and system based on zero-order optimization, and relates to the technical field of artificial intelligence. A large language model is deployed in a federal learning framework, and the usage amount of model parameters during training and the transmission amount in the communication process are reduced by using low-rank decomposition, so that the training efficiency is improved. And then performing zero-order fine tuning on the local model to complete updating of the model. The technical problem of high memory requirement in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for fine-tuning a federated large language model based on zero-order optimization. Background Art

[0002] With the rapid development of artificial intelligence technology, especially in the field of natural language processing, large language models have gradually become the focus of research and application. These models are trained on large-scale text data through deep learning methods, and can generate high-quality text content, understand complex semantic relationships, and support a variety of natural language tasks. At the same time, federated learning, as a distributed machine learning paradigm, aims to protect user privacy while achieving data sharing and model training across multiple devices or institutions. This combination not only promotes the development of privacy-preserving natural language processing technology, but also provides new possibilities for personalized services.

[0003] However, in actual deployment, there are significant challenges in fine-tuning large language models to adapt to specific application scenarios. On the one hand, traditional fine-tuning methods require a lot of computing resources and memory space, which makes it difficult to implement when the client has limited computing power. For example, when fine-tuning using the first-order optimization method, each participant needs to store the entire model parameters, which places high demands on memory capacity. On the other hand, in order to improve model performance, personalized adjustments are often required for the specific circumstances of different clients, but existing federated learning strategies are insufficient in this regard, usually using a unified learning rate and other settings, ignoring the differences between the participants.

[0004] In addition, although existing research has explored how to reduce communication overhead in federated learning and improve model convergence speed, it is still insufficient in solving the above problems. In particular, when facing a large-scale parameter space, how to effectively use limited client resources to complete high-quality model updates has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The present invention provides a method and system for fine-tuning a large federated language model based on zero-order optimization, which solves the technical problem of high memory requirement in the prior art.

[0006] The first aspect of the present invention provides a method for fine-tuning a federated large language model based on zero-order optimization, comprising:

[0007] Deploy the large language model in the federated learning framework to build a federated large language model;

[0008] Performing low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix;

[0009] The federated large language model is zero-order fine-tuned according to the low-rank parameter matrix.

[0010] Optionally, the federated large language model includes a central server and a plurality of clients in communication with the central server, and performing zero-order differentiation on the federated large language model according to the low-rank parameter matrix includes:

[0011] Adding random perturbations to the low-rank parameter matrix of each client and performing perturbation training to determine a zero-order gradient estimate;

[0012] Determining a global update parameter based on the zero-order gradient estimate and a preset learning rate;

[0013] The global update parameters are used to update the plurality of clients, and a preset data set is used to iteratively train the updated federated large language model;

[0014] The learning rate of each client is updated according to the training result, and the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate is jumped to execute until the preset number of iterations is met.

[0015] Optionally, it also includes:

[0016] Determining a model convergence influencing parameter from the global update parameters;

[0017] Adjust the model convergence influencing parameters according to a preset ascending gradient, and update the global update parameters;

[0018] Based on the updated global update parameters, the step of using the global update parameters to update the multiple clients and iteratively training the updated federated large language model using a preset data set is jumped to execution.

[0019] Optionally, adding random perturbations to the low-rank parameter matrix of each client and performing perturbation training to determine a zero-order gradient estimate includes:

[0020] Adding random perturbations to the low-rank parameter matrix of each of the clients;

[0021] Perform bidirectional perturbations on the low-rank parameter matrix with added random perturbations to obtain forward perturbation loss and backward perturbation loss;

[0022] Performing a difference operation using the forward disturbance loss and the backward disturbance loss to obtain a target difference;

[0023] The target difference value is multiplied by a preset disturbance coefficient to obtain a zero-order gradient estimate.

[0024] Optionally, determining a global update parameter based on the zero-order gradient estimate and a preset learning rate includes:

[0025] Performing a multiplication operation by using the zero-order gradient estimate and a preset learning rate to obtain a scaled gradient estimate;

[0026] Update low-rank parameter matrices of a plurality of the clients according to the scaled gradient estimate;

[0027] The central server aggregates multiple updated low-rank parameter matrices to determine global update parameters.

[0028] Optionally, the training result includes single-round training loss difference data, multi-round loss difference data, and model parameter update difference data, and the updating of the learning rate of each client according to the training result, and the jump execution of the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate until the preset number of iterations is met, includes:

[0029] Normalizing the single-round training loss difference data, the multi-round loss difference data, and the model parameter update difference data of each client;

[0030] The normalized single-round training loss difference data, multi-round loss difference data, and model parameter update difference data are used to determine the heterogeneity value;

[0031] Using the heterogeneity value and the preset trade-off coefficient, determine and update a new learning rate for each of the clients;

[0032] Based on the updated learning rate, jump to the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate until the preset number of iterations is met.

[0033] A second aspect of the present invention provides a federated large language model fine-tuning system based on zero-order optimization, comprising:

[0034] The framework deployment module is used to deploy the large language model in the federated learning framework and build a federated large language model;

[0035] A low-rank decomposition module, used for performing low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix;

[0036] A zero-order fine-tuning module is used to perform zero-order fine-tuning on the federated large language model according to the low-rank parameter matrix.

[0037] A third aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for fine-tuning a federated large language model based on zero-order optimization as described in any one of the above items.

[0038] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed, the method for fine-tuning a federated large language model based on zero-order optimization as described in any one of the above items is implemented.

[0039] A fifth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the federated large language model fine-tuning method based on zero-order optimization as described in any one of the above items.

[0040] It can be seen from the above technical solutions that the present invention has the following advantages:

[0041] The present invention deploys a large language model into a federated learning framework, uses low-rank decomposition to reduce the usage of model parameters during training and the amount of transmission during communication, and then performs zero-order fine-tuning on the local model to complete the model update. This solves the technical problem of high memory requirements in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0043] Figure 1 A flowchart of a method for fine-tuning a federated large language model based on zero-order optimization provided in Embodiment 1 of the present invention;

[0044] Figure 2 A flowchart of a method for fine-tuning a federated large language model based on zero-order optimization provided in Embodiment 2 of the present invention;

[0045] Figure 3 A schematic diagram of a zero-order federated learning framework for a federated large language model provided in Embodiment 2 of the present invention;

[0046] Figure 4 A structural block diagram of a federated large language model fine-tuning system based on zero-order optimization provided in Embodiment 3 of the present invention;

[0047] Figure 5 This is a structural block diagram of a computer device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0048] The embodiments of the present invention provide a method and system for fine-tuning a federated large language model based on zero-order optimization, which are used to solve the technical problem of high memory requirement in the prior art.

[0049] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] In view of the problems mentioned in the background technology, it is particularly important to develop a new method that can reduce memory consumption while maintaining efficient training effects. In recent years, zero-order optimization has attracted widespread attention as a method for optimization that does not require gradient information. It updates parameters by directly evaluating the value of the objective function, thereby avoiding the additional computational burden brought by calculating the gradient. Introducing zero-order optimization into the federated learning framework can effectively alleviate memory pressure while maintaining good convergence energy. However, how to deeply understand the behavioral characteristics of zero-order optimization in a federated learning environment at the theoretical level, and how to design a reasonable algorithm structure to give full play to its advantages, is still an open research topic.

[0051] In summary, the present invention aims to solve the problems of high memory requirement and lack of personalized strategy in the prior art by proposing a federated tuning method based on zero-order optimization, and to provide a new solution for efficient fine-tuning of large language models in federated learning.

[0052] See also Figure 1 , Figure 1 A flowchart of the steps of a method for fine-tuning a federated large language model based on zero-order optimization provided in Example 1 of the present invention.

[0053] The present invention provides a method for fine-tuning a federated large language model based on zero-order optimization, comprising:

[0054] Step 101: deploy the large language model in the federated learning framework to build a federated large language model.

[0055] In an embodiment of the present invention, a large language model is deployed to a federated learning framework to construct a federated large language model. Under the federated learning framework, the data of each client can be processed and trained locally without uploading the data to a central server. During the training process of the large language model, the data is always kept locally, and only the updated information of the model is securely exchanged and aggregated between the parties, which can effectively prevent data leakage.

[0056] Step 102: Perform low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix.

[0057] In an embodiment of the present invention, a low-rank adaptation module is used to perform low-rank decomposition on the federated large language model, and the high-dimensional weight matrix of each client in the federated large language model is decomposed into two low-rank parameter matrices. The low-rank adaptation module is used to reduce the usage of model parameters during training and the transmission amount during communication.

[0058] Step 103: Perform zero-order fine-tuning on the federated large language model according to the low-rank parameter matrix.

[0059] In the embodiment of the present invention, zero-order fine-tuning is performed on the federated large language model according to the low-rank parameter matrix. Specifically, the zero-order fine-tuning is to update the model by perturbing the parameters, thereby avoiding the extra computational burden caused by gradient calculation, thereby reducing memory requirements. This method is particularly suitable for resource-constrained client devices, and can achieve efficient model updates while maintaining low memory usage.

[0060] In the present invention, a large language model is deployed in a federated learning framework, a federated large language model is constructed, a low-rank decomposition is performed on the federated large language model, a low-rank parameter matrix is ​​obtained, and zero-order fine-tuning is performed on the federated large language model according to the low-rank parameter matrix; the present invention deploys a large language model in a federated learning framework, and uses low-rank decomposition to reduce the usage of model parameters during training and the amount of transmission during communication, and then zero-order fine-tuning is performed on the local model to complete the model update. The technical problem of high memory requirements in the prior art is solved.

[0061] See also Figure 2 , Figure 2 A flowchart of the steps of a method for fine-tuning a federated large language model based on zero-order optimization provided in Embodiment 2 of the present invention.

[0062] The present invention provides a method for fine-tuning a federated large language model based on zero-order optimization, comprising:

[0063] Step 201: deploy the large language model in the federated learning framework to build a federated large language model.

[0064] Furthermore, the federated large language model includes a central server and a plurality of clients that are communicatively connected to the central server.

[0065] It should be noted that the central server and multiple clients are connected in a star-shaped structure. The central server is located at the core, and multiple clients are directly connected to the central server. Each client has an independent communication link with the central server. The central server can easily control and coordinate the global model to ensure the training progress and model consistency of each client.

[0066] In the embodiment of the present invention, please refer to Figure 3 , deploy the large language model to the federated learning framework, which specifically involves initializing the central server and loading the pre-trained model, and then distributing the parameters of the model to each participating client. The federated large language model includes a central server and N clients, and the data in each client follows the distribution ,The aggregation goal of federated learning is to optimize the global model parameters by minimizing the weighted average of all client loss functions (the central server is not responsible for training, but only for model aggregation).

[0067] The following are the global loss function of the central server and the local loss function of the client in the federated large language model. The specific formulas for the tth communication round are as follows, where T represents the last communication round. Formula (1) is interpreted as minimizing the global loss function of the tth communication round, that is, the aggregate loss function of the last local round H of each client; Formula (2) is interpreted as using random small batch data to train the expected loss function of the kth local round of the tth communication round of the i-th client, which is equal to the actual loss function:

[0068] (1)

[0069] (2)

[0070] In the formula, represents the d-dimensional parameter of the federated large language model, represents the global loss function of the central server in the tth communication round, Indicates the number of clients. represents the total number of local rounds, represents the local loss function of the ith client in the tth communication round and the Hth local round, represents the local loss function of the kth local round of the ith client in the tth communication round, Represents a small batch data set obtained by randomly filtering the user's local data, which conforms to the distribution of local data , Representation dataset The obtained local loss function.

[0071] Step 202: Perform low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix.

[0072] In the embodiment of the present invention, a low-rank adaptation module is introduced into the federated large language model, and then the low-rank adaptation module is used to perform low-rank decomposition on the federated large language model, and the high-dimensional weight matrix of each client in the federated large language model is decomposed into two low-rank parameter matrices and , each client only updates the low-rank matrix instead of the entire model parameters during local training. Specifically, in the subsequent training process, each client uses its local mini-batch dataset right and Perform zero-order fine-tuning to update the parameters to minimize the loss function value, that is, fix the weights of the original model and only update the low-rank parameter matrix and Parameters.

[0073] It should be noted that low-rank decomposition uses the low-rank approximation principle in linear algebra, that is, a large and complex matrix can be approximated by the product of two or more much smaller matrices. In this way, the parameter matrix of the originally huge federated language model Can be In the form of and The dimension of is much smaller than that of the original weight matrix, which significantly reduces the number of parameters that need to be updated and greatly reduces the demand for computing resources.

[0074] It should be noted that the specific mathematical expression of updating each client using the low-rank adaptation framework is as follows:

[0075] (3)

[0076] In the formula, represents the parameter matrix of the k+1th local round of the ith client in the tth communication round, Indicates the i-th client obtained through training and Low rank structure.

[0077] Due to the linear addition characteristics of low-rank structures and pre-trained weights, transmitting only the parameters of the low-rank structure has the same effect as transmitting all the parameters. The mathematical explanation is as follows:

[0078] (4)

[0079] In the formula, Represents the client parameter matrix broadcast in the tth communication round.

[0080] It should be noted that in formula (4), represents the parameter matrix of the i-th client in the new round updated and transmitted using the global update parameters. The first equation describes that the parameter matrix is ​​equal to the aggregate parameters of all clients at the end of the previous round. The second equation describes the actual use of The parameter update process of low-rank adaptive federated learning, and the third equation shows that the full amount of transmission parameters (leftmost ) is equivalent to transmitting only Low-rank adaptation, verifying that only training and transmission The feasibility of low-rank adaptation.

[0081] Step 203: Add random perturbations to the low-rank parameter matrix of each client and perform perturbation training to determine the zero-order gradient estimate.

[0082] It should be noted that the zero-order optimization framework is introduced to perform zero-order fine-tuning on the federated large language model. In the entire process of federated learning, random perturbations are added to approximate the gradient calculation, thereby reducing the memory overhead and completing the training convergence process of the model.

[0083] Further, step 203 may include the following sub-steps:

[0084] S11. Add random perturbations to the low-rank parameter matrix of each client.

[0085] In the embodiment of the present invention, each client performs forward propagation locally on two low-rank parameter matrices and Add small random perturbations to the parameters of .

[0086] S12. Perform bidirectional perturbations on the low-rank parameter matrix with added random perturbations to obtain forward perturbation loss and backward perturbation loss.

[0087] In an embodiment of the present invention, a low-rank parameter matrix with random perturbation is bidirectionally perturbed, and the bidirectional perturbation includes forward perturbation and backward perturbation, specifically:

[0088] Forward perturbation: perturb the current parameter matrix Add a small positive perturbation , calculate the new loss function for each client ;

[0089] Backward perturbation: perturb the current parameter matrix Add a small positive perturbation , calculate the new loss function for each client .

[0090] It should be noted that zero-order optimization is a method that can be optimized without directly calculating gradient information. Its core idea is based on the first approximation of Taylor expansion, which estimates the gradient direction by adding random perturbations to the model parameters and evaluating the changes in the objective function. Specifically, for a given loss function, its gradient can be approximated in the following way.

[0091] When each client performs forward propagation locally, and The parameters of the model are randomly perturbed and the change in the loss function value before and after the perturbation is calculated. By comparing these two values, an unbiased estimate of the gradient at that position can be obtained. This process generates a gradient estimate for each parameter, which is used to guide subsequent parameter updates.

[0092] S13, using the forward perturbation loss and the backward perturbation loss to perform a difference operation to obtain a target difference.

[0093] S14, performing a multiplication operation on the target difference and the preset disturbance coefficient to obtain a zero-order gradient estimate.

[0094] In a specific implementation, in order to facilitate the implementation of the method, the above S13-S14 process can be converted into a formula encapsulation form, wherein the formula of the zero-order gradient estimation value can be as follows:

[0095] (5)

[0096] In the formula, represents the zero-order gradient estimate, represents the preset disturbance coefficient, express is a Gaussian distributed random perturbation vector, Indicates the degree of disturbance.

[0097] It is worth mentioning that the zero-order gradient estimate calculated by the client in the present invention replaces the true gradient relied on in the traditional first-order optimization method to update the model parameters. Since the calculation of the true gradient often involves traversal of the entire data set and complex derivative operations, especially for large-scale data sets and complex model structures (such as deep neural networks), the amount of calculation is extremely large, and the zero-order gradient estimate can be quickly obtained through some approximate methods, which greatly reduces the consumption of computing resources and computing time. At the same time, in distributed training scenarios (such as federated learning), the communication bandwidth between the client and the server is often a bottleneck. After calculating the true gradient, transmitting it to the server or other nodes may take up a lot of communication resources. The zero-order gradient estimate usually has a more concise representation and a smaller transmission amount.

[0098] Step 204: Determine a global update parameter based on the zero-order gradient estimate and a preset learning rate.

[0099] It should be noted that the model parameters are updated according to the calculated approximate calculation gradients of each client. Specifically, an optimization algorithm (such as SGD and Adam) can be used to update the local low-rank parameter matrix of the client according to the gradient.

[0100] Further, step 204 may include the following sub-steps:

[0101] S21. Perform multiplication operation on the zero-order gradient estimate and the preset learning rate to obtain a scaled gradient estimate.

[0102] In an embodiment of the present invention, a zero-order gradient estimate is multiplied by a preset learning rate to obtain a scaled gradient estimate, and in an optimization algorithm based on gradient descent, the multiplication operation is used to update model parameters. The zero-order gradient estimate indicates the direction of change of the objective function at the current parameter point, while the learning rate controls the step size of each parameter update. By multiplying the zero-order gradient estimate by the learning rate, the distance that the model parameters should move in the opposite direction of the gradient can be determined, thereby gradually adjusting the parameters to reduce the value of the objective function.

[0103] It should be noted that the preset learning rate here is a pre-set hyperparameter used to control the update step size, and its size will affect the speed and effect of model convergence.

[0104] S22. Update low-rank parameter matrices of multiple clients according to the scaled gradient estimate.

[0105] In an embodiment of the present invention, a difference operation is performed between the current low-rank parameter matrix and the scaled gradient estimate, thereby obtaining the updated low-rank parameter matrices of multiple clients.

[0106] S23. Aggregate multiple updated low-rank parameter matrices through a central server to determine a global update parameter.

[0107] In the embodiment of the present invention, the updated and It is uploaded to the central server, which collects the updated low-rank parameter matrices uploaded by all clients, and uses the weighted average method to aggregate multiple updated low-rank parameter matrices to obtain the global updated parameters.

[0108] It should be noted that in step 205, global update parameters are used to update multiple clients, that is, each client receives a global update matrix after global aggregation from the central server. The global update parameters in the matrix are the result of joint training by all participants (clients). The client deploys the received global update parameters to the local environment and combines them with the fixed part of the pre-trained large language model to form a complete local model.

[0109] The complete local model is as follows:

[0110] (6)

[0111] in, Characterizes the global update parameter after aggregation.

[0112] Step 205: Use global update parameters to update multiple clients, and use a preset data set to iteratively train the updated federated large language model.

[0113] In an embodiment of the present invention, the central server distributes the aggregated global update parameters to each client, and the client uses these global update parameters to replace part of the parameters in the local original low-rank parameter matrix, thereby obtaining a new federated large language model.

[0114] Through this series of operations, we can complete zero-order fine-tuning of the federated large language model while reducing the amount of model parameters used and communication transmission, thereby improving the model's adaptability and performance on different client data.

[0115] Each client then uses its local mini-batch dataset right and Perform zero-order fine-tuning for iterative training to obtain training results, which include single-round training loss difference data, multi-round loss difference data, and model parameter update difference data.

[0116] It is worth mentioning that the present invention proposes a personalized federated learning strategy, which dynamically adjusts the learning rate according to the data distribution and computing power of different clients. Through this strategy, the specific characteristics of each client can be better adapted, thereby accelerating the decline of the loss function and improving the model training effect.

[0117] A personalized federated learning strategy mainly considers the degree of adaptation of the learning rate of each client. It is worth mentioning that in the hypothesis derivation of step 207, the convergence rate of non-independent and identically distributed It is shown that the heterogeneity value The higher the value, the higher the overall convergence rate. Based on this analysis, the heterogeneity measure is incorporated into the personalized federated learning strategy. It is an abstract concept. Since heterogeneity cannot be directly measured, it needs to be represented by a proxy. Instantiation, specifically, requires the use of training states and signals to base the model training process on. Based on the model training process, three methods for fitting heterogeneity are provided:

[0118] Single round training loss difference data:

[0119] The loss value of each client in the previous round of training And the global loss value The difference between:

[0120] (7)

[0121] In the embodiment of the present invention, the updated federated large language model is iteratively trained using a preset data set. Based on the training output, the loss value of each client in the previous round of training and the global loss value can be calculated using formulas (1) and (2), and then the difference between them is calculated, that is, the loss difference data of a single round of training.

[0122] Multi-round loss difference data:

[0123] The average loss deviation of each client relative to the global loss in the past multiple rounds (preferably five rounds in this embodiment) of training (preferably five rounds in this embodiment):

[0124] (8)

[0125] In the embodiment of the present invention, the updated federated large language model is iteratively trained using a preset data set. Based on the training output, the loss value of each client in the first few rounds of training and the global loss value can be calculated using formulas (1) and (2), and then the average loss deviation, that is, the multi-round loss difference data, can be calculated.

[0126] Model parameter update difference data:

[0127] The parameter update amplitude of each client in the previous round With global update amplitude The difference between:

[0128] (9)

[0129] In an embodiment of the present invention, the updated federated large language model is iteratively trained using a preset data set. The client parameter update amplitude measures the degree of change of the model parameters of each client in a round of training. Before each round of training begins, the current model parameters of each client are recorded. After the round of training is completed, the updated model parameters are obtained. For each client, the Euclidean norm method can be used to calculate the Euclidean distance of the parameter vector to represent the update amplitude, or the average absolute change method can be used to calculate the average value of the absolute value of the change of each element of the parameter vector to represent the update amplitude. The global update amplitude requires the aggregation of global parameters first (such as the above steps 204-205). After each round of training, the model parameters of each client need to be aggregated to obtain the global model parameters. The calculation of the global update amplitude is similar to the calculation of the client parameter update amplitude. The Euclidean norm method or the average absolute change method is used to calculate the global update amplitude, and then the difference between the parameter update amplitude and the global update amplitude is calculated, that is, the model parameter update difference data.

[0130] Step 206: Update the learning rate of each client according to the training result, and jump to the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate until the preset number of iterations is met.

[0131] Further, step 206 may include the following sub-steps:

[0132] S31. Normalize the single-round training loss difference data, multi-round loss difference data, and model parameter update difference data of each client.

[0133] In an embodiment of the present invention, multiple single-round training loss difference data, multi-round loss difference data, and model parameter update difference data calculated by the above formulas (7) to (9) are normalized to the range of (-1, 1).

[0134] S32. Use the normalized single-round training loss difference data, multi-round loss difference data, and model parameter update difference data to determine the heterogeneity value.

[0135] In the specific implementation, in order to facilitate the implementation of the method, the above S32 process can be converted into a formula encapsulation form, where the heterogeneity value The formula can be as follows:

[0136] (10)

[0137] In the embodiment of the present invention, the normalized single-round training loss difference data, multi-round loss difference data and model parameter update difference data are used to calculate the heterogeneity value of each client through formula (10).

[0138] S33. Using the heterogeneity value and the preset trade-off coefficient, determine and update the new learning rate of each client.

[0139] It is worth mentioning that the learning rate of clients with higher heterogeneity is dynamically increased and the preset trade-off coefficient is used Constrain the degree of personalization, using Characterizes the heterogeneity of the i-th client, which means the degree of difference in data distribution between different clients.

[0140] In a specific implementation, in order to facilitate the implementation of the method, the above S33 process can be converted into a formula encapsulation form, wherein the formula of the updated learning rate can be as follows:

[0141] (11)

[0142] In the embodiment of the present invention, the heterogeneity degree value and the preset weight coefficient are used. Perform a multiplication operation, and perform a sum operation on the value after the multiplication operation and 1 to obtain the updated learning rate .

[0143] S34. Based on the updated learning rate, jump to the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate until the preset number of iterations is met.

[0144] In the embodiment of the present invention, the updated learning rate is used to replace the previously preset learning rate, and the process jumps to step 204, and the iteration is repeated until the preset number of iterations is reached, and the iterative training is stopped, and the updated federated large language model is output to complete the zero-order fine-tuning. In this way, even in an environment with limited computing resources, high-quality model training can be effectively completed, while significantly reducing memory usage and communication costs.

[0145] It should be noted that the learning rate of each customer is updated through the above-mentioned personalized federated learning strategy to ensure that the model can meet the specific needs of different users or tasks while ensuring an efficient convergence rate, thereby achieving a more flexible and efficient federated learning mechanism.

[0146] Step 207: Determine model convergence influencing parameters from the global update parameters.

[0147] In the embodiment of the present invention, it can be known from the above formula (6) that To characterize the global update parameters after aggregation, it is necessary to analyze and determine the key parameters that affect the model optimization performance as model convergence influencing parameters, so as to improve the convergence speed of the model training process.

[0148] The mathematical derivation of the training process consists of two assumptions to form the convergence rate drop of each step, and then the overall convergence analysis is obtained. According to the convex optimization theory, the convergence rate and related properties of the present invention are obtained based on the scenario of federated learning and large models. It is necessary to adopt certain assumptions and derive step by step. The specific derivation steps are as follows:

[0149] S41. Setting multiple assumptions;

[0150] The application scenario of the present invention is a large language model and a federated large language model constructed by a federated learning framework, so certain relevant assumptions are adopted.

[0151] Specifically, this derivation is based on a non-independent and identically distributed scenario, which means that the data of each client is different, which is consistent with the actual data situation. The relevant assumptions are as follows:

[0152] Loss function lower bound assumption: global loss function The lower bound of , that is, ,have ;

[0153] L-smooth condition assumption: All loss functions satisfy the L-smooth condition, taking the loss function of the i-th client as For example, the following equation is satisfied, where is the gradient of the loss function. This assumption can be described as the gradient change rate of the loss function is limited by a constant L, thereby ensuring that the function changes smoothly:

[0154] (12)

[0155] (13)

[0156] In the formula, and Represents two different parameter matrices in the parameter space, specifically the state of the model at different times or under different parameter matrix settings.

[0157] Mini-batch learning error assumption: For any parameter matrix , the second-order moment of the stochastic gradient is restricted to the following formula, which can be described as using a small batch dataset The gradient of the loss function of zero-order optimization will not be too far away from the true gradient. and Restrictions:

[0158] (14)

[0159] In the formula, represents the expectation about the mini-batch dataset, represents a constant, represents a constant, represents a function of the parameter matrix and the mini-batch dataset, Represents the gradient vector of the function with respect to the parameter matrix.

[0160] Non-IID data assumption: For any , the global gradient Local gradients with each client There is a difference, which can be explained as the difference between the loss function gradient of each client and the global gradient and Restrictions, namely:

[0161] (15)

[0162] In the formula, represents a constant, It represents the variance term, which specifically reflects the degree of difference between different data subsets or the fluctuation of the data in the context of non-independent and identically distributed data.

[0163] The above assumptions describe the properties of the function and the scenario limitations of federated learning. Considering that large models have a large number of parameters, it is also necessary to adopt the assumption that the large model has a low effective rank, that is, to assume that fine-tuning the large model can be equivalent to occurring only in a low-dimensional subspace (dimension r), which can be formally expressed as:

[0164] There is a Hessian matrix in the loss function , which is the second-order partial derivative matrix of the loss function with respect to the model parameters, which provides information about the curvature of the function and helps to understand the local geometric characteristics in the parameter space, thus affecting the convergence and stability of the optimization algorithm. The matrix satisfies the following conditions:

[0165] For all parameter matrices , when the parameter matrix satisfy Sometimes ,in ;

[0166] in, represents the parameter matrix of the model, represents the parameter matrix of the model for the tth communication round in federated learning, It represents the maximum norm of the gradient of the loss function under a given data distribution, reflecting the sensitivity of the loss function to data changes under this parameter value. represents the second-order derivative of the loss function with respect to the parameter matrix, represents the Hessian matrix, The maximum effective rank of , characterized by ,in is the trace of the matrix, is the operator norm, Does it represent a dimension parameter related to data or model? Represents the learning rate.

[0167] S42. According to a plurality of assumptions, a preset step-by-step convergence function in a non-federated scenario is used to obtain a convergence value of the federated large language model.

[0168] It should be noted that, assuming the loss function of the current federated large language model is The L-smooth condition is satisfied, and the large model low effective rank assumption in step S41 is met. Using convex optimization theory, the gradual convergence formula of the large model in the non-federated scenario can be obtained, that is, the parameter changes of each training.

[0169] Specifically, the two-point gradient estimation method of step 203 is used. If the Hessian matrix has a value of of low effective rank, and the constant and If exists, then the expected value of each step of the loss function will be limited as follows, where is the overall learning rate, and the preset gradual convergence function is:

[0170] (16)

[0171] in, Represents the convergence value, specifically the expected value of the loss function when the parameter is, and the overall learning rate , used to limit or adjust the effect of the positive term of the gradient estimate on the decrease of the loss function , is the number of gradient estimates performed in each gradient update iteration, represents the decrease in the loss function due to updating the parameters along the gradient direction, represents a correction term associated with the positive term of the gradient estimate.

[0172] The formula shows that the rate of decrease of the loss function of the centralized zero-order optimization algorithm is affected by the low effective rank and the gradient estimate is limited, where the gradient negative term Limited by the overall learning rate , the gradient estimate positive term Limited by influencing parameters .

[0173] The descent lemma is based on the parameter assumption of the large model, which characterizes the large number of parameters of the large model as a low effective rank , and further convergence analysis is carried out on this basis.

[0174] S43. Introduce the convergence value into the federation scenario, obtain the convergence rate of each parameter in the global update parameter in the federation scenario, and use the parameter associated with the maximum convergence rate as the model convergence influencing parameter.

[0175] In an embodiment of the present invention, in order to introduce the gradual convergence formula of the centralized zero-order optimization algorithm in step S42 into the federation scenario, it is necessary to apply the assumptions in the federation scenario such as the multiple non-independent and identically distributed data assumptions mentioned in step S41.

[0176] Specifically, through all the above assumptions, the learning rate Fixed to , the degree of disturbance Fixed to When , the convergence rate can be obtained in the federation scenario:

[0177] (17)

[0178] in, It means that within the range of the communication round of federated learning, the expected value of the norm square of the gradient of the loss function with respect to the parameter matrix is ​​minimized. represents the parameter convergence rate involving the number of clients, the iteration rounds of client local training, and the communication rounds in federated learning, represents the number of clients involved and the parameter convergence rate of the iterative rounds of client local training, represents the parameter convergence rate with respect to the number of clients involved, is the number of clients, Iteration rounds of local training on the client. is the communication round in federated learning, is the dimension of low effective rank in the low effective rank hypothesis.

[0179] and The two items characterize the heterogeneity degree value and the difference value between different clients, and introduce constraints on the heterogeneity degree value.

[0180] Since the number of model parameters is a relatively large value (in the billions), so the second term of formula (17) is small and cannot make a significant contribution to the convergence rate. Similarly, in the third term It also makes little contribution to the overall convergence, so only the term By observing the first term of the convergence rate formula, we can find that the heterogeneity value The higher the value, the higher the overall convergence rate, indicating that the heterogeneity between clients is conducive to the overall convergence. Therefore, the key parameter affecting the convergence performance of the federated large language model is the number of clients. , the number of iterations of client local training and communication rounds .

[0181] Step 208: adjust the model convergence influencing parameters according to the preset ascending gradient and update the global update parameters.

[0182] It should be noted that the number of clients The increase of not only brings more data diversity, but also enables each global update to integrate more local information, thereby approaching the optimal solution more effectively.

[0183] In addition, the number of iterations of client local training It is also crucial; by appropriately increasing the number of local training times on each client, the model can be better fitted on the local data without frequent and expensive cross-client communication, thereby improving the quality of the global model.

[0184] The communication round It determines the frequency of model updates and the degree of synchronization. A higher communication round means more frequent global parameter synchronization, which helps to ensure model consistency between clients and further stabilize the entire training process.

[0185] Therefore, in the actual training process, the number of clients is kept Faster convergence. Choose a larger number of local iterations at the beginning of training, and further increase the number of local iterations as the iterations progress. . However, due to its contribution to the convergence rate, , so further expansion may not necessarily produce significant convergence help, and the current training overhead needs to be measured. Through in-depth analysis of these hyperparameters, it can not only provide theoretical support for the training process, but also provide valuable guidance for optimizing algorithm design.

[0186] In the embodiment of the present invention, the model convergence influencing parameters are adjusted according to the preset ascending gradient, and the global update parameters are updated to obtain the updated global update parameters.

[0187] Step 209: Based on the updated global update parameters, jump to the step of updating multiple clients with the global update parameters and iteratively training the updated federated large language model with the preset data set.

[0188] In the embodiment of the present invention, based on the updated global update parameters, the process jumps to step 205 .

[0189] In the present invention, by introducing memory-efficient zero-order optimization technology, efficient fine-tuning of large language models is achieved within the framework of federated learning, avoiding the additional computational burden caused by gradient calculation, thereby significantly reducing memory requirements. In addition, the model convergence influencing parameters that affect the optimization performance are determined through theoretical analysis, and the corresponding convergence properties are established. At the same time, a personalized federated learning strategy is proposed, which dynamically adjusts the learning rate according to the data distribution and computing power of different clients to better adapt to the specific characteristics of each client, accelerate the decline rate of the loss function, and improve the model training effect. The present invention effectively solves the problems of high memory requirements and lack of personalized strategies faced when fine-tuning large language models under the condition of limited client computing resources. It not only improves the training efficiency of large language models in federated learning, but also significantly reduces the usage of graphics card memory, and provides theoretical and technical support for designing more flexible and personalized federated learning solutions, which has certain reference significance for other application scenarios that require model fine-tuning in resource-constrained environments.

[0190] See also Figure 4 , Figure 4 This is a structural block diagram of a federated large language model fine-tuning system based on zero-order optimization provided in Example 3 of the present invention.

[0191] The present invention provides a federated large language model fine-tuning system based on zero-order optimization, comprising:

[0192] The framework deployment module 301 is used to deploy the large language model in the federated learning framework to build a federated large language model;

[0193] A low-rank decomposition module 302 is used to perform low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix;

[0194] The zero-order fine-tuning module 303 is used to perform zero-order fine-tuning on the federated large language model according to the low-rank parameter matrix.

[0195] Furthermore, the federated large language model includes a central server and a plurality of clients in communication connection with the central server, and the zero-order fine-tuning module 303 includes:

[0196] The zero-order gradient estimation submodule is used to add random perturbations to the low-rank parameter matrix of each client and perform perturbation training to determine the zero-order gradient estimation value;

[0197] A global update parameter submodule, used to determine the global update parameter based on the zero-order gradient estimate and the preset learning rate;

[0198] An iterative training submodule, which is used to update multiple clients using global update parameters and iteratively train the updated federated large language model using a preset data set;

[0199] The first jump submodule is used to update the learning rate of each client according to the training result, and jump to the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate until the preset number of iterations is met.

[0200] Furthermore, the zero-order fine-tuning module 303 also includes:

[0201] A model convergence influencing parameter submodule is used to determine the model convergence influencing parameter from the global update parameters;

[0202] The parameter update submodule is used to adjust the model convergence influencing parameters according to the preset ascending gradient and update the global update parameters;

[0203] The second jump submodule is used to jump to the step of updating multiple clients with the global update parameters based on the updated global update parameters, and iteratively training the updated federated large language model using a preset data set.

[0204] Furthermore, the zero-order gradient estimation submodule includes:

[0205] A disturbance adding unit, used to add random disturbance to the low-rank parameter matrix of each client;

[0206] A bidirectional perturbation unit is used to bidirectionally perturb the low-rank parameter matrix with added random perturbations to obtain forward perturbation loss and backward perturbation loss;

[0207] A difference operation unit, used for performing a difference operation using a forward disturbance loss and a backward disturbance loss to obtain a target difference;

[0208] The multiplication operation unit is used to perform a multiplication operation using the target difference and the preset disturbance coefficient to obtain a zero-order gradient estimate.

[0209] Furthermore, the global update parameter submodule includes:

[0210] A scaled gradient estimate unit, used to perform a multiplication operation by using a zero-order gradient estimate and a preset learning rate to obtain a scaled gradient estimate;

[0211] A matrix update unit for updating low-rank parameter matrices of multiple clients based on the scaled gradient estimates;

[0212] The aggregation unit is used to aggregate multiple updated low-rank parameter matrices through a central server to determine a global update parameter.

[0213] Furthermore, the training results include single-round training loss difference data, multi-round loss difference data, and model parameter update difference data. The first jump submodule includes:

[0214] A normalization unit, used to normalize the single-round training loss difference data, multi-round loss difference data, and model parameter update difference data of each client;

[0215] A heterogeneity degree value unit, used to determine a heterogeneity degree value by using normalized single-round training loss difference data, multi-round loss difference data, and model parameter update difference data;

[0216] A learning rate updating unit, used to determine and update a new learning rate of each client by using a heterogeneity degree value and a preset trade-off coefficient;

[0217] The iterative convergence unit is used to jump to the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate based on the updated learning rate until the preset number of iterations is met.

[0218] In the present invention, a large language model is deployed in a federated learning framework, a federated large language model is constructed, a low-rank decomposition is performed on the federated large language model, a low-rank parameter matrix is ​​obtained, and zero-order fine-tuning is performed on the federated large language model according to the low-rank parameter matrix; the present invention deploys a large language model in a federated learning framework, and uses low-rank decomposition to reduce the usage of model parameters during training and the amount of transmission during communication, and then zero-order fine-tuning is performed on the local model to complete the model update. The technical problem of high memory requirements in the prior art is solved.

[0219] See also Figure 5 , Figure 5 This is a structural block diagram of a computer device provided in Embodiment 4 of the present invention.

[0220] An electronic device according to an embodiment of the present invention includes: a memory 401 and a processor 402, wherein the memory 402 stores a computer program; when the computer program is executed by the processor 402, the processor 402 executes the federated large language model fine-tuning method based on zero-order optimization as described in any of the above embodiments.

[0221] The memory 401 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. The memory 401 has a storage space 403 for a program code 413 for executing any method step in the above method. For example, the storage space 403 for the program code may include individual program codes 413 for implementing the various steps in the above method, respectively. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The program code may be compressed, for example, in an appropriate form. When these codes are run by a computing and processing device, the computing and processing device performs the various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The program code may be compressed, for example, in an appropriate form. When these codes are executed by a computing processing device, they cause the computing processing device to execute the various steps in the zero-order optimization-based federated large language model fine-tuning method described above.

[0222] Embodiment 5 of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for fine-tuning a federated large language model based on zero-order optimization as described in any of the above embodiments is implemented.

[0223] Embodiment 6 of the present invention further provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the federated large language model fine-tuning method based on zero-order optimization as described in any of the above embodiments.

[0224] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0225] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0226] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0227] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0228] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0229] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for fine-tuning a federated large language model based on zero-order optimization, characterized in that: include: Deploy the large language model in the federated learning framework to build a federated large language model; Performing low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix; The federated large language model is zero-order fine-tuned according to the low-rank parameter matrix.

2. The method for fine-tuning a federated large language model based on zero-order optimization according to claim 1, characterized in that: The federated large language model includes a central server and a plurality of clients in communication with the central server, and performing zero-order differentiation on the federated large language model according to the low-rank parameter matrix includes: Adding random perturbations to the low-rank parameter matrix of each client and performing perturbation training to determine a zero-order gradient estimate; Determining a global update parameter based on the zero-order gradient estimate and a preset learning rate; The global update parameters are used to update the plurality of clients, and a preset data set is used to iteratively train the updated federated large language model; The learning rate of each client is updated according to the training result, and the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate is jumped to execute until the preset number of iterations is met.

3. The method for fine-tuning a federated large language model based on zero-order optimization according to claim 2, characterized in that: Also includes: Determining a model convergence influencing parameter from the global update parameters; Adjust the model convergence influencing parameters according to the preset ascending gradient, and update the global update parameters; Based on the updated global update parameters, the step of using the global update parameters to update the multiple clients and iteratively training the updated federated large language model using a preset data set is jumped to execution.

4. The method for fine-tuning a federated large language model based on zero-order optimization according to claim 2, characterized in that: The adding random perturbation to the low-rank parameter matrix of each client and performing perturbation training to determine the zero-order gradient estimate includes: Adding random perturbations to the low-rank parameter matrix of each of the clients; Perform bidirectional perturbations on the low-rank parameter matrix with added random perturbations to obtain forward perturbation loss and backward perturbation loss; Performing a difference operation using the forward disturbance loss and the backward disturbance loss to obtain a target difference; The target difference value is multiplied by a preset disturbance coefficient to obtain a zero-order gradient estimate.

5. The method for fine-tuning a federated large language model based on zero-order optimization according to claim 2, characterized in that: The determining of the global update parameter based on the zero-order gradient estimate and a preset learning rate includes: Performing a multiplication operation by using the zero-order gradient estimate and a preset learning rate to obtain a scaled gradient estimate; Update low-rank parameter matrices of a plurality of the clients according to the scaled gradient estimate; The central server aggregates multiple updated low-rank parameter matrices to determine global update parameters.

6. The method for fine-tuning a federated large language model based on zero-order optimization according to claim 2, characterized in that: The training results include single-round training loss difference data, multi-round loss difference data, and model parameter update difference data. The learning rate of each client is updated according to the training results, and the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate is jumped to execute until the preset number of iterations is met, including: Normalizing the single-round training loss difference data, the multi-round loss difference data, and the model parameter update difference data of each client; The normalized single-round training loss difference data, multi-round loss difference data, and model parameter update difference data are used to determine the heterogeneity value; Using the heterogeneity value and the preset trade-off coefficient, determine and update a new learning rate for each of the clients; Based on the updated learning rate, jump to the step of determining the global update parameter based on the zero-order gradient estimate and the preset learning rate until the preset number of iterations is met.

7. A federated large language model fine-tuning system based on zero-order optimization, based on the federated large language model fine-tuning method based on zero-order optimization according to any one of claims 1 to 6, characterized in that: include: The framework deployment module is used to deploy the large language model in the federated learning framework and build a federated large language model; A low-rank decomposition module, used for performing low-rank decomposition on the federated large language model to obtain a low-rank parameter matrix; A zero-order fine-tuning module is used to perform zero-order fine-tuning on the federated large language model according to the low-rank parameter matrix.

8. An electronic device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the method for fine-tuning a federated large language model based on zero-order optimization as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the method for fine-tuning a federated large language model based on zero-order optimization as described in any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to execute the federated large language model fine-tuning method based on zero-order optimization as described in any one of claims 1-6.

Citation Information

Cited By

  • Large model parameter adjustment method based on cloud edge collaboration and related equipment

    CN120851140A

  • Cloud-edge collaboration-based large model parameter adjustment method and related device

    CN120851140B

  • Large language model zero-order fine tuning method and system based on low-rank projection matrix learning

    CN122045829A

  • Low-rank projection matrix learning based large language model zeroth-order fine-tuning method and system

    CN122045829B