Training method and device of federal large language model, equipment and medium
By using distillation technology, singular value decomposition and knowledge migration methods in federal large language model training, the problem that the client's small model knowledge cannot be enriched into large models on the server side is solved, which reduces communication costs and computing resource requirements, and improves the generalization ability and privacy protection effect of the model.
Patent Information
- Application Number
- CN202510064621.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The existing federal large language model training method ignores the problem that the client's small large language model's insights on its unique domain knowledge can be enriched into the server-side large language model. At the same time, since each client has to pass on knowledge with the server-side large model, the overall communication cost of the system is high.
By obtaining the preset large language model on the server side, using distillation technology to obtain a small language model and send it to the client. The client trains the small language model based on local private data and uploads the model weight parameters. The server side filters the key singular values and the corresponding singular vectors through singular value decomposition, and updates the model weight parameters, and obtains the client summary model through weighting, and finally transfers the small language model and the preset large language model to update the server-side model.
It reduces the redundant information of model weight parameters, reduces the amount of computing, improves response speed, protects user privacy, improves the generalization ability and performance of the model, and realizes the distributed utilization of computing resources, reducing server hardware costs and energy consumption.
Smart Images

Figure CN119990367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model training technology, and in particular to a method, device, equipment and medium for training a federated large language model. Background Art
[0002] In recent years, large language models have received widespread attention and have been applied in various fields, but they still face challenges in the development of real-world scenarios. On the one hand, these challenges stem from the fact that high-quality public domain data is becoming increasingly scarce, which greatly limits the improvement of model knowledge coverage and generalization capabilities; on the other hand, in private fields such as medical, financial, and scientific research industries, data contains sensitive information, privacy protection is crucial, and traditional data aggregation training models are difficult to implement. In this context, in order to solve these problems, the federated large language model came into being. It combines the technical architecture of federated learning and large language models. Under the framework of federated learning, multiple participants such as different institutions, organizations, or devices can collaborate to train large language models without sharing local private data, which can effectively utilize the data resources of all parties while protecting data privacy.
[0003] However, current federated large language models mainly focus on enabling clients to collaboratively fine-tune their locally deployed smaller large language models, or transferring knowledge from the server-side larger large language model to the client-side small large language model. These approaches all ignore the fact that the insights of the client-side small large language model into its own unique domain knowledge can also be enriched into the server-side large language model. At the same time, if each client needs to transfer knowledge with the server-side large model, the alignment work of the small model and the large model will continue to increase with the increase of clients, resulting in a higher overall communication cost for the system. Summary of the invention
[0004] In order to solve the above technical problems, one or more embodiments of this specification provide a method, apparatus, device and medium for training a federated large language model.
[0005] One or more embodiments of this specification adopt the following technical solutions:
[0006] One or more embodiments of this specification provide a method for training a federated large language model, the method comprising:
[0007] Obtain the preset large language model on the server side of the current scenario, obtain the corresponding small language model based on the distillation technology, and send it to each client;
[0008] The small language model is trained according to the local private data of each client, so as to upload the model weight parameter matrix corresponding to the small language model trained by each client to the server;
[0009] Decomposing each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors, and updating the model weight parameters according to the key singular values and corresponding singular vectors;
[0010] Weighting the updated model weight parameters to obtain a client summary model, and weighting the client summary model and the small language model to obtain a current small language model on the server side;
[0011] The knowledge of the current small language model is migrated based on a common data set of the preset large language model and the current small language model to achieve training update for the preset large language model.
[0012] Optionally, in one or more embodiments of the present specification, the decomposing each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors specifically includes:
[0013] Performing singular value decomposition on the model weight parameter matrix uploaded by each client to obtain singular values and singular vectors corresponding to each model weight parameter matrix;
[0014] Sorting the singular values corresponding to each of the model weight parameter matrices and obtaining the sum of the singular values;
[0015] Determine the maximum value of the singular value as a starting point based on the sorting order, so as to accumulate the singular values in sequence according to the sorting order to obtain the number of accumulated singular values and the sum of accumulated singular values;
[0016] Obtaining a ratio of the accumulated singular value sum to the sum of the singular values to determine whether the ratio is greater than a preset ratio;
[0017] If so, the accumulated number of singular values is used as the reserved number, so as to sequentially screen multiple singular values corresponding to the reserved number based on the starting point as key singular values, and obtain a singular value vector corresponding to the key singular value.
[0018] Optionally, in one or more embodiments of the present specification, updating the model weight parameters according to the key singular values and the corresponding singular vectors specifically includes:
[0019] Constructing a current matrix based on the key singular value and the singular vector corresponding to the key singular value;
[0020] The current matrix is scaled according to the specification of the model weight parameter matrix to obtain a current model parameter matrix, so that the model weight parameters corresponding to the current model parameter matrix are used as updated model weight parameters.
[0021] Optionally, in one or more embodiments of the present specification, weighting the updated model weight parameter to obtain a client summary model, and weighting the client summary model and the small language model to obtain the current small language model on the server side, specifically includes:
[0022] Acquire the size of the private data set of each client, and determine the weighted value corresponding to each client based on the ratio of the total size of the private data set of each client to the size of the private data set of the client;
[0023] According to the weighted value corresponding to each of the clients, each updated model weight parameter is weighted to obtain a client summary model.
[0024] Optionally, in one or more embodiments of the present specification, obtaining a preset large language model on the server side of the current scene to obtain a corresponding small language model based on the distillation technology and sending it to each client specifically includes:
[0025] Acquire a preset large language model corresponding to the server side based on the task requirements of the current scenario, and load the selected large language model into a memory corresponding to the server side;
[0026] Performing a configuration check on the preset large language model to determine whether to perform distillation on the preset large language model based on the check result;
[0027] If yes, then obtain a public data set corresponding to the task requirement, predict the input text of the public data set based on the preset large language model to obtain a soft label, and determine a hard label based on the existing label of the public data set;
[0028] The loss function corresponding to the distillation technology is determined according to the soft label and the hard label, so as to obtain a small language model by training on the public data set through a back propagation algorithm according to the loss function, so as to send the small language model to each client.
[0029] Optionally, in one or more embodiments of the present specification, obtaining a preset large language model corresponding to the server based on the task requirements of the current scenario specifically includes:
[0030] Matching the functional keywords of each existing large language model based on the task requirement to obtain one or more initial large language models corresponding to the task requirement, and determining a first matching weight of each initial large language model based on the matching number of the keywords;
[0031] Determining key performance indicators corresponding to each of the task requirements according to the task requirements and historical tasks in which each of the initial large language models matches the task requirements;
[0032] Determining a second matching weight of each of the initial large language models according to a value of a key performance indicator corresponding to each of the initial large language models;
[0033] According to the first matching weight and the second matching weight, a matching value of each of the initial large language models is determined, so as to filter and obtain a preset large language model corresponding to the task requirement of the current scene based on the matching value.
[0034] Optionally, in one or more embodiments of the present specification, after performing knowledge migration on the current small language model based on the common data set of the preset large language model and the current small language model to implement training update for the preset large language model, the method further includes:
[0035] Perform performance evaluation on the updated preset large language model based on the preset test dataset;
[0036] According to the performance evaluation result, adjusting parameters of the updated preset large language model;
[0037] The deployment environment of the adjusted preset large language model is determined based on the task requirement, so as to call a corresponding script to complete the deployment of the adjusted preset large language model based on the deployment environment.
[0038] One or more embodiments of this specification provide a training device for a federated large language model, the device comprising:
[0039] The distillation unit is used to obtain the preset large language model of the server side of the current scene, obtain the corresponding small language model based on the distillation technology, and send it to each client;
[0040] A training unit, used for training the small language model according to the local private data of each client, so as to upload the model weight parameter matrix corresponding to the small language model trained by each client to the server;
[0041] An updating unit, configured to decompose each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors, and update the model weight parameters according to the key singular values and the corresponding singular vectors;
[0042] an acquisition unit, configured to weight the updated model weight parameter to obtain a client summary model, and to weight the client summary model and the small language model to obtain a current small language model on the server side;
[0043] A migration unit is used to perform knowledge migration on the current small language model based on a common data set of the preset large language model and the current small language model, so as to implement training update for the preset large language model.
[0044] One or more embodiments of this specification provide a training device for a federated large language model, the device comprising:
[0045] at least one processor; and,
[0046] a memory communicatively connected to the at least one processor; wherein,
[0047] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform any of the above methods.
[0048] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute any of the above-described methods.
[0049] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:
[0050] Based on singular value decomposition, the key singular values and corresponding singular vectors are screened to update the model weight parameters, which reduces the redundant information of the model weight parameters and the subsequent calculation amount, thereby improving the overall response speed. In addition, each client uses local private data to train the small language model, which can customize the model according to the specific needs and usage scenarios of different clients, and the client only uploads the trained model weight parameter matrix to the server, rather than the original data, which effectively protects the user's privacy information. By weighting the updated model weight parameters to obtain the client summary model, and then weighting it with the small language model, the current small language model on the server is obtained, so that the final model can integrate the training results of different clients and improve the generalization ability and performance of the model. In addition, allowing each client to participate in model training utilizes the idle computing resources of the client and realizes the distributed use of computing resources. This avoids the excessive dependence of centralized training on powerful computing resources on the server side and reduces the hardware cost and energy consumption of the server. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:
[0052] Figure 1 A flowchart of a method for training a federated large language model provided in an embodiment of this specification;
[0053] Figure 2 A schematic diagram of the structure of a training device for a federated large language model provided in an embodiment of this specification;
[0054] Figure 3 A schematic diagram of the structure of a training device for a federated large language model provided in an embodiment of this specification;
[0055] Figure 4 A schematic diagram of the structure of a non-volatile storage medium provided in an embodiment of this specification. DETAILED DESCRIPTION
[0056] The embodiments of this specification provide a method, apparatus, device and medium for training a federated large language model.
[0057] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0058] like Figure 1 As shown, the present specification embodiment provides a flowchart of a method for training a federated large language model. Figure 1 It can be seen that in one or more embodiments of this specification, a method for training a federated large language model includes the following steps:
[0059] S101: Obtain a preset large language model on the server side of the current scene, obtain a corresponding small language model based on distillation technology, and send it to each client.
[0060] Large language models usually have a huge parameter scale, which will take up a lot of storage space. For the client, in order to generate results faster and thus improve the operating efficiency of the client, a model with smaller parameters is required to execute. Therefore, under the framework of federated learning, the preset large language model of the server side of the current scene will be obtained in the embodiment of this specification, and the preset large language model will be processed according to the distillation technology to obtain the corresponding small language model, which will be sent to each client. At this time, the small language model sent can be used as the basic model for local training of the client. Since they all come from the distillation of the same large model, they have certain commonalities, which enables better collaboration when the client uploads the trained model parameters back to the server side for aggregation, reducing problems such as aggregation difficulties caused by large model differences.
[0061] Specifically, in one or more embodiments of the present specification, a preset large language model of the server side of the current scene is obtained to obtain a corresponding small language model based on the distillation technology, and the small language model is sent to each client, which specifically includes the following process:
[0062] It can ensure that the selected model has a high correlation with the task in terms of architecture, pre-training knowledge, etc. For example, if it is a sentiment analysis task, you can choose a preset model that performs well in text sentiment understanding, so that the subsequent distilled small model is more in line with the task requirements from the source, and the practicality of the model in a specific scenario is improved. Therefore, in the embodiment of this specification, the preset large language model corresponding to the server side will be obtained according to the task requirements of the current scenario, and the selected large language model will be loaded into the memory corresponding to the server side. Then the preset large language model is configured to determine whether to distill the preset large language model based on the inspection results. If distillation is required, then the public data set corresponding to the task requirement is obtained, so as to obtain soft labels based on the input text of the public data set based on the preset large language model, and determine the hard labels based on the existing labels of the public data set. So as to determine the loss function corresponding to the distillation technology according to the soft labels and hard labels. The soft labels are obtained by presetting the large language model, and the loss function is determined in combination with the existing hard labels of the public data set. This method comprehensively considers the output probability distribution and the true label of the model. The small model can better learn the knowledge of the large model and the real requirements of the task during the training process, thereby improving the performance of the small model. Then, according to the loss function, a small language model is obtained by training on the public dataset through a back propagation algorithm, so that the small language model can be sent to each client to provide a unified basis for subsequent federated learning or other collaborative training.
[0063] Further, in one or more embodiments of the present specification, obtaining a preset large language model corresponding to the server based on the task requirements of the current scenario specifically includes the following process:
[0064] First, according to the current task requirements and the functional keywords of each existing large language model, one or more initial large language models corresponding to the task requirements are obtained, and the first matching weight of each initial large language model is determined according to the number of keyword matches. According to the historical tasks that match the task requirements and each initial large language model with the task requirements, the key performance indicators corresponding to each task requirement are determined. At the same time, according to the values of the key performance indicators corresponding to each initial large language model, the second matching weight of each initial large language model is determined. And according to the first matching weight and the second matching weight, the matching value of each initial large language model is determined to obtain the preset large language model corresponding to the task requirements of the current scene based on the matching value screening. In this process, the preset large language model is determined by combining the functional keyword matching and the key performance indicator matching, so that the model that best suits the current task requirements can be found more comprehensively and accurately. This comprehensive consideration method avoids the model selection errors that may be caused by relying on a single factor, and improves the adaptability of the model to the task.
[0065] S102: The small language model is trained according to the local private data of each client, so as to upload the model weight parameter matrix corresponding to the small language model trained by each client to the server.
[0066] Each client has its own local private data, which may be text data, such as medical records owned by medical institutions, customer transaction information and comments held by financial institutions, etc. The client uses local data to train the initial model against the small language model obtained from the server. During the local model training process, the model parameters are updated according to the characteristics of the local data. These parameters contain the language knowledge learned from the local data, such as vocabulary usage habits, semantic understanding patterns, etc. Then, the model weight parameter matrix corresponding to the small language model trained by each client can be obtained and uploaded to the server. In this process, the client uses local private data to train the model locally, avoiding uploading sensitive data such as medical records of medical institutions and customer transaction information of financial institutions to the server, thereby fundamentally protecting data privacy. This enables institutions involving sensitive information to participate in model training without leaking data, ensuring that sensitive information will not be obtained externally. In addition, after distributing the model training tasks to each client and uploading the model weight parameter matrix corresponding to the small language model trained by the client to the server, the local computing resources of the client are utilized, which reduces the computing pressure on the server. The server does not need to process all the training tasks for the data, so it can more efficiently coordinate other tasks such as aggregating model parameters.
[0067] S103: Decomposing each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors, and updating the model weight parameters according to the key singular values and the corresponding singular vectors.
[0068] After the model weight parameter matrix is obtained based on the above steps, the model weight parameter matrix may contain a large amount of information, some of which is redundant, so in order to remove these relatively unimportant parts and reduce the redundancy of the data. In the embodiment of this specification, the model weight parameters corresponding to each client are decomposed according to the singular value decomposition, so as to screen out the key singular values and the singular vectors corresponding to the key singular values, so as to update the model weight parameters according to the key singular values and their corresponding singular value vectors. In this process, the unfavorable factors are removed by updating, so that the model can better adapt to the task requirements.
[0069] Specifically, in one or more embodiments of the present specification, each model weight parameter matrix is decomposed based on singular value decomposition to screen key singular values and corresponding singular vectors, specifically including the following process:
[0070] In a certain application scenario, the singular values corresponding to each client can be taken from large to small. The first k values taken by each client are different. The value of k for each client depends on the ratio of the sum of the k singular values to the total sum of the singular values of the client to be greater than or equal to a set value. Then obtain the singular vectors corresponding to these k singular values respectively. That is to say, firstly, the singular value decomposition is performed on the model weight parameter matrix uploaded by each client to obtain the singular values and singular vectors corresponding to each model weight parameter matrix. Then the singular values corresponding to each model weight parameter matrix are sorted, and the sum of the singular values is obtained. Thus, the maximum value of the singular value is determined as the starting point according to the sorting order, so as to accumulate each singular value in turn according to the sorting order to obtain the number of accumulated singular values and the sum of accumulated singular values. Then obtain the ratio of the sum of the accumulated singular values to the sum of the singular values to determine whether the ratio is greater than the preset ratio. If yes, then the accumulated number of singular values is used as the reserved number, so as to sequentially screen multiple singular values corresponding to the reserved number based on the starting point as key singular values, and obtain a singular value vector corresponding to the key singular value.
[0071] In this process, by sorting by singular value size and calculating the ratio of the sum of accumulated singular values to the total, the number of key singular values can be adaptively determined, and the distribution of singular values of each model weight parameter matrix can be screened, ensuring that the most critical information can be effectively extracted in different situations. The key singular values are determined by accumulating singular values in sequence starting from the maximum value, taking into account the order of importance of the data information represented by the singular values. The size of the singular value reflects the importance of the information contained in the corresponding eigenvector. This method prioritizes the most important information part, ensuring that the selected key singular values and singular vectors can represent the core knowledge in the model weight parameter matrix to the greatest extent. After the key singular values and corresponding singular vectors are screened, the unnecessary amount of calculation caused by redundant data is reduced in the subsequent use of this information to update the model weight parameters and other calculation processes.
[0072] Specifically, in one or more embodiments of the present specification, updating the model weight parameters according to the key singular values and the corresponding singular vectors specifically includes:
[0073] Based on the key singular values and the singular vectors corresponding to the key singular values, a current matrix is constructed. Then, the current matrix is scaled according to the specifications of the model weight parameter matrix to obtain the current model parameter matrix, so that the model weight parameters corresponding to the current model parameter matrix are used as updated model weight parameters.
[0074] S104: weighting the updated model weight parameters to obtain a client summary model, and weighting the client summary model and the small language model to obtain a current small language model on the server side.
[0075] After the model weight parameters of the small language model of the client are obtained after local training and screening based on the above step S103, the updated model weight parameters will be weighted to obtain the client summary model, that is, the new model weight parameters of each client are weighted and summed according to the scale of the client private data set to obtain the client summary model. Since the data scales of different clients are different, the weighting of this step can balance their contributions to the summary model. It avoids the situation that clients with small data volumes and clients with large data volumes have equal influence in the model aggregation process, so that the summary model can more accurately reflect the comprehensive characteristics of all client data. Then the client summary model and the small language model are weighted to obtain the current small language model of the server side, that is, the client summary model is obtained by weighting, and then weighted with the small language model, which can effectively integrate the knowledge of each client and the knowledge of the original small language model of the server side. This fusion enables the current small language model of the server side to absorb personalized knowledge from different clients, and improves the accuracy and generalization ability of the model when processing various tasks and data types.
[0076] Specifically, in one or more embodiments of the present specification, weighting the updated model weight parameters to obtain the client summary model, and weighting the client summary model and the small language model to obtain the current small language model on the server side, specifically includes the following process:
[0077] First, the size of each client's private data set is obtained, and then the weighted value corresponding to each client is determined based on the ratio of the total size of each client's private data set to the size of the client's private data set. Then, according to the weighted value corresponding to each client, the updated model weight parameters are weighted to obtain the client summary model. By calculating the ratio of the size of each client's private data set to the total size to determine the weighted value, the proportion of each client's data in the whole can be accurately measured, and the important features in different client data can be more effectively captured, so that the summary model integrates the data characteristics of clients of different sizes.
[0078] S105: Performing knowledge migration on the current small language model based on a common data set of the preset large language model and the current small language model, so as to implement training update for the preset large language model.
[0079] The current small language model obtained in the above process is transferred according to the public data set of the preset large language model and the current small language model to realize the training update of the preset large language model, that is, the original output of the two models on the public data set is used as the medium for transferring knowledge. The original outputs of the two models are aligned through the vocabulary mapping tables of the two models to ensure efficient knowledge transfer between the two models. During the knowledge transfer, the original outputs of the two models on the public data set are compared to select the original output with better performance for transfer. Specifically, it can be implemented based on the following process: first, after the server side obtains the updated current small language model, a knowledge set is calculated on the same public data set, each element in the set is a loss-original output pair, and the total number of elements in the set is the number of data in the data set. Then, the loss value and original output of the preset large language model reasoning are calculated for each data in the public data set, which also constitutes a knowledge set of loss-original output pairs. According to the vocabulary mapping table between the preset large language model and the current small language model, the original output in the knowledge set of the current small language model is converted to the original output of the preset large language model, and the original output in the knowledge set of the preset large language model is converted to the original output of the current small language model. Then compare the loss of the preset large language model and the current small language model when inferring on each public data. If the loss of the preset large language model is smaller, the original output of the preset large language model after conversion on the data is included in the current small language model knowledge base. Otherwise, the original output of the current small language model after conversion on the data is included in the preset large language model knowledge base, and use these knowledge bases to fine-tune each model to achieve updated training.
[0080] Further, in one or more embodiments of the present specification, after performing knowledge migration on the current small language model based on the public data set of the preset large language model and the current small language model to implement training update for the preset large language model, the method further includes the following process:
[0081] Based on the preset test data set, the performance of the updated preset large language model is evaluated, and then the parameters of the updated preset large language model are adjusted according to the performance evaluation results. At the same time, the deployment environment of the adjusted preset large language model is determined based on the task requirements, and the corresponding script is called to complete the deployment of the adjusted preset large language model based on the deployment environment. Adjusting the parameters of the model based on the performance evaluation results helps to continuously optimize the model and gradually improve its performance. Selecting a suitable deployment environment according to task requirements and calling the corresponding script to complete the deployment helps to improve the applicability of the model in different scenarios.
[0082] like Figure 2 As shown, the present specification embodiment provides a structural diagram of a training device for a federated large language model. Figure 2It can be seen that in one or more embodiments of this specification, a training device for a federated large language model includes:
[0083] The distillation unit 201 is used to obtain the preset large language model of the server side of the current scene, obtain the corresponding small language model based on the distillation technology, and send it to each client;
[0084] A training unit 202 is used to train the small language model according to the local private data of each client, so as to upload the model weight parameter matrix corresponding to the small language model trained by each client to the server;
[0085] An updating unit 203 is used to decompose each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors, and update the model weight parameters according to the key singular values and the corresponding singular vectors;
[0086] An acquisition unit 204 is configured to weight the updated model weight parameters to obtain a client summary model, and to weight the client summary model and the small language model to obtain a current small language model on the server side;
[0087] The migration unit 205 is used to perform knowledge migration on the current small language model based on the common data set of the preset large language model and the current small language model, so as to implement training update for the preset large language model.
[0088] like Figure 3 As shown, the present specification provides a structural diagram of a training device for a federated large language model. Figure 3 It can be seen that in one or more embodiments of this specification, a training device for a federated large language model includes:
[0089] at least one processor; and,
[0090] a memory communicatively connected to the at least one processor; wherein,
[0091] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform any of the above methods.
[0092] like Figure 4 As shown, a schematic diagram of the structure of a non-volatile storage medium is provided in an embodiment of the present specification. Figure 4 It can be seen that in one or more embodiments of the present specification, a non-volatile storage medium stores computer-executable instructions, and the computer-executable instructions can execute any of the methods described above.
[0093] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0094] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0095] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.
Claims
1. A method for training a federated large language model, characterized in that: The method comprises: Obtain the preset large language model on the server side of the current scenario, obtain the corresponding small language model based on the distillation technology, and send it to each client; The small language model is trained according to the local private data of each client, so as to upload the model weight parameter matrix corresponding to the small language model trained by each client to the server; Decomposing each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors, and updating the model weight parameters according to the key singular values and corresponding singular vectors; Weighting the updated model weight parameters to obtain a client summary model, and weighting the client summary model and the small language model to obtain a current small language model on the server side; The knowledge of the current small language model is migrated based on a common data set of the preset large language model and the current small language model to achieve training update for the preset large language model.
2. The method for training a federated large language model according to claim 1, characterized in that: Decomposing each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors specifically includes: Performing singular value decomposition on the model weight parameter matrix uploaded by each client to obtain singular values and singular vectors corresponding to each model weight parameter matrix; Sorting the singular values corresponding to each of the model weight parameter matrices and obtaining the sum of the singular values; Determine the maximum value of the singular value as a starting point based on the sorting order, so as to accumulate the singular values in sequence according to the sorting order to obtain the number of accumulated singular values and the sum of accumulated singular values; Obtaining a ratio of the accumulated singular value sum to the sum of the singular values to determine whether the ratio is greater than a preset ratio; If so, the accumulated number of singular values is used as the reserved number, so as to sequentially screen multiple singular values corresponding to the reserved number based on the starting point as key singular values, and obtain a singular value vector corresponding to the key singular value.
3. The method for training a federated large language model according to claim 1, characterized in that: The updating of the model weight parameters according to the key singular values and the corresponding singular vectors specifically includes: Constructing a current matrix based on the key singular value and the singular vector corresponding to the key singular value; The current matrix is scaled according to the specification of the model weight parameter matrix to obtain a current model parameter matrix, so that the model weight parameters corresponding to the current model parameter matrix are used as updated model weight parameters.
4. The method for training a federated large language model according to claim 1, characterized in that: The step of weighting the updated model weight parameters to obtain a client summary model, and weighting the client summary model and the small language model to obtain a current small language model on the server side, specifically includes: Acquire the size of the private data set of each client, and determine the weighted value corresponding to each client based on the ratio of the total size of the private data set of each client to the size of the private data set of the client; According to the weighted value corresponding to each of the clients, each updated model weight parameter is weighted to obtain a client summary model.
5. The method for training a federated large language model according to claim 1, characterized in that: The method of obtaining a preset large language model on the server side of the current scene, obtaining a corresponding small language model based on the distillation technology, and sending the small language model to each client specifically includes: Acquire a preset large language model corresponding to the server side based on the task requirements of the current scenario, and load the selected large language model into a memory corresponding to the server side; Performing a configuration check on the preset large language model to determine whether to perform distillation on the preset large language model based on the check result; If yes, then obtain a public data set corresponding to the task requirement, predict the input text of the public data set based on the preset large language model to obtain a soft label, and determine a hard label based on the existing label of the public data set; The loss function corresponding to the distillation technology is determined according to the soft label and the hard label, so as to obtain a small language model by training on the public data set through a back propagation algorithm according to the loss function, so as to send the small language model to each client.
6. The method for training a federated large language model according to claim 5, characterized in that: Acquiring a preset large language model corresponding to the server based on the task requirements of the current scenario specifically includes: Matching the functional keywords of each existing large language model based on the task requirement to obtain one or more initial large language models corresponding to the task requirement, and determining a first matching weight of each initial large language model based on the matching number of the keywords; Determining key performance indicators corresponding to each of the task requirements according to the task requirements and historical tasks in which each of the initial large language models matches the task requirements; Determining a second matching weight of each of the initial large language models according to a value of a key performance indicator corresponding to each of the initial large language models; According to the first matching weight and the second matching weight, a matching value of each of the initial large language models is determined, so as to filter and obtain a preset large language model corresponding to the task requirement of the current scene based on the matching value.
7. The method for training a federated large language model according to claim 1, characterized in that: After performing knowledge migration on the current small language model based on the common data set of the preset large language model and the current small language model to implement training update for the preset large language model, the method further includes: Perform performance evaluation on the updated preset large language model based on the preset test dataset; According to the performance evaluation result, adjusting parameters of the updated preset large language model; The deployment environment of the adjusted preset large language model is determined based on the task requirement, so as to call a corresponding script to complete the deployment of the adjusted preset large language model based on the deployment environment.
8. A training device for a federated large language model, characterized in that: The device comprises: The distillation unit is used to obtain the preset large language model of the server side of the current scene, obtain the corresponding small language model based on the distillation technology, and send it to each client; A training unit, used for training the small language model according to the local private data of each client, so as to upload the model weight parameter matrix corresponding to the small language model trained by each client to the server; An updating unit, configured to decompose each of the model weight parameter matrices based on singular value decomposition to screen key singular values and corresponding singular vectors, and update the model weight parameters according to the key singular values and the corresponding singular vectors; an acquisition unit, configured to weight the updated model weight parameter to obtain a client summary model, and to weight the client summary model and the small language model to obtain a current small language model on the server side; A migration unit is used to perform knowledge migration on the current small language model based on a common data set of the preset large language model and the current small language model, so as to implement training update for the preset large language model.
9. A training device for a federated large language model, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the methods described in claims 1 to 7.
10. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 7.