A Fine-Tuning Method, Device, and Medium for Federal Large Models Based on Mutual Learning
By introducing a fine-tuning method based on mutual learning in the federal large language model, using the client small model to calculate the public data set, obtain the loss value and knowledge set, and fine-tune the large language model on the server side, the problem of ignoring the unique domain knowledge of the client small model in the existing technology is solved, and the performance of both language models is improved.
Patent Information
- Application Number
- CN202510038816.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The existing federal large language model mainly focuses on the transfer of knowledge between the client and the server side, ignoring the enrichment effect of the unique domain knowledge of the client small model on the server side large model.
The first language model is fine-tuned through the client and the local private data set to obtain the second language model, and the second language model is used to calculate the public data set, obtain feature vectors, attention distribution and gradient information, calculate the loss value, form a knowledge set, calculate the public data set through the third language model, and compare the loss value to fine-tune the third language model.
It realizes that without leaking the client's privacy data set, the knowledge of the server-side large language model is adaptively transferred to the client's small language model, and at the same time, the unique insights of the client's small language model are enriched into the server-side large language model, achieving the common improvement of the performance of both language models.
Smart Images

Figure CN119443209B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large language models, and in particular, to a method, device, and medium for fine-tuning a federated large model based on mutual learning. Background Art
[0002] In recent years, large language models have received extensive attention and have been applied in various fields. However, they still face challenges in the development of real-world scenarios. These challenges stem from the scarcity of public domain data and the need to maintain privacy in private domain data. To address these issues, federated large language models have emerged.
[0003] Research on federated large language models mainly focuses on enabling clients to collaboratively fine-tune their respective smaller-scale large language models deployed locally, or transferring the knowledge of a larger-scale large language model on the server side to the smaller-scale large language models on the client side. However, these methods ignore the fact that the insights of the smaller-scale large language models on the client side regarding their unique domain knowledge can also enrich the larger-scale large language model on the server side.
[0004] Through the above analysis, the problems and deficiencies of the existing technology are as follows:
[0005] The existing federated large language models in the prior art mainly focus on knowledge transfer between the client side and the server side, but ignore the enrichment effect of the unique domain knowledge of the smaller-scale models on the client side on the larger-scale model on the server side. Summary of the Invention
[0006] The embodiments of this application provide a method, device, and medium for fine-tuning a federated large model based on mutual learning, which can solve the problem that the existing federated large language models mainly focus on knowledge transfer between the client side and the server side, but ignore the enrichment effect of the unique domain knowledge of the smaller-scale models on the client side on the larger-scale model on the server side.
[0007] In a first aspect, the embodiments of this application provide a method for fine-tuning a federated large model based on mutual learning. The method includes: fine-tuning a first language model through a client and a local private dataset to obtain a second language model; calculating a second result for a public dataset through the second language model, where the public dataset already has a first result, and the second result includes feature vectors, attention distributions, and gradient information; calculating a first loss value between the second result and the first result, and forming a first knowledge set with the second result and the first loss value; obtaining the knowledge set, and calculating a third result for the public dataset through a third language model; calculating a second loss value of the third result, and comparing the first loss value and the second loss value to fine-tune the third language model.
[0008] In an implementation of the present application, calculate the second loss value of the third result, and compare the first loss value and the second loss value to fine-tune the third language model, specifically including: obtaining the smallest first loss value, and comparing the smallest first loss value and the second loss value according to the vocabulary mapping table with the client; if it is less than the second loss value, obtain the first knowledge set of the client and fine-tune the third language model.
[0009] In an implementation of the present application, after calculating the second loss value of the third result, and comparing the first loss value and the second loss value to fine-tune the third language model, the method further includes: forming a second knowledge set with the third result and the second loss value; sending the third result to the client according to the vocabulary mapping table with the client.
[0010] In an implementation of the present application, before fine-tuning the first language model through the client and the local private data set to obtain the second language model, the method further includes: collecting the local private data set and the public data set, extracting the vocabulary, and counting the occurrence frequency and context information of the vocabulary; determining the mapping relationship between the vocabulary based on the similarity and context information of the vocabulary to obtain the vocabulary mapping table.
[0011] In an implementation of the present application, the method further includes: removing duplicate vocabulary, merging vocabulary with a similarity greater than a preset threshold to periodically update the vocabulary mapping table; establishing a user feedback mechanism and a sharing mechanism to allow different clients to share the first knowledge set.
[0012] In an implementation of the present application, fine-tuning the first language model through the client and the local private data set to obtain the second language model, specifically including: introducing two low-rank matrices in the attention mechanism layer and the feed-forward neural network layer of the first language model; initializing the low-rank matrices, and training the low-rank matrices using the local private data set to obtain the fine-tuned model.
[0013] In an implementation of the present application, the method further includes: evaluating the accuracy and efficiency of the fine-tuned model through synonym replacement; when the accuracy and efficiency reach the preset threshold, save the current low-rank matrices to obtain the second language model.
[0014] In an implementation of the present application, obtaining the smallest first loss value, and comparing the smallest first loss value and the second loss value according to the vocabulary mapping table with the client, specifically including: based on the first loss value of the first result including the calculated feature vector, attention distribution, and gradient information, respectively comparing the first loss values of the feature vector, attention distribution, and gradient information; respectively selecting the output forms with the smallest first loss values of the feature vector, attention distribution, and gradient information, and combining them to obtain the smallest first loss value.
[0015] In a second aspect, an embodiment of the present application further provides a device for fine-tuning a federated large model based on mutual learning. The device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: fine-tune a first language model through a client and a local private dataset to obtain a second language model; calculate a second result by using the second language model for a public dataset, wherein the public dataset already has a first result, and the second result includes a feature vector, an attention distribution, and gradient information; calculate a first loss value between the second result and the first result, and form a first knowledge set with the second result and the first loss value; obtain the knowledge set, and calculate a third result by using a third language model for the public dataset; calculate a second loss value of the third result, and compare the first loss value and the second loss value to fine-tune the third language model.
[0016] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for fine-tuning a federated large model based on mutual learning, storing computer-executable instructions, and the computer-executable instructions are set to: fine-tune a first language model through a client and a local private dataset to obtain a second language model; calculate a second result by using the second language model for a public dataset, wherein the public dataset already has a first result, and the second result includes a feature vector, an attention distribution, and gradient information; calculate a first loss value between the second result and the first result, and form a first knowledge set with the second result and the first loss value; obtain the knowledge set, and calculate a third result by using a third language model for the public dataset; calculate a second loss value of the third result, and compare the first loss value and the second loss value to fine-tune the third language model.
[0017] A fine-tuning method, device, and medium for a federated large model based on mutual learning provided by an embodiment of the present application make full use of the knowledge of the client's small language model to enrich the server's large language model, and at the same time use the knowledge of the server's large model to improve the capabilities of the client's small language model; borrow the original output of the fine-tuned model on the public dataset as a medium for transferring knowledge, and align the original outputs of the two models by establishing a vocabulary mapping table for the server and client language models to ensure efficient knowledge transfer between the two models. When transferring knowledge, by comparing the original outputs of the two models on the public dataset, select the original output with better performance for transfer. This method can adaptively transfer the knowledge of the server's large language model to the client's small language model without disclosing the client's private dataset, and at the same time enrich the unique insights of the client's small language model into the server's large language model, achieving a common improvement in the performance of the server and client language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a flowchart of a fine-tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0020] Figure 2 is a schematic diagram of the test set results of a fine-tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of the first loss value of a fine-tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0022] Figure 4 is a schematic diagram of the component relationship of a fine-tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0023] Figure 5 is a schematic diagram comparing the first loss value and the second loss value of a fine-tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0024] Figure 6 is a schematic diagram of the first loss value of the feature vector, attention distribution, and gradient information of a fine-tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0025] Figure 7Schematic diagram for comparing the losses between the server - side and the local model of a fine - tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0026] Figure 8 Overall logic schematic diagram of a fine - tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0027] Figure 9 Schematic diagrams before and after model fine - tuning of a fine - tuning method for a federated large model based on mutual learning provided by an embodiment of the present application;
[0028] Figure 10 Internal structure schematic diagram of a fine - tuning device for a federated large model based on mutual learning provided by an embodiment of the present application. Detailed implementation manners
[0029] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0030] An embodiment of the present application provides a fine - tuning method, device, and medium for a federated large model based on mutual learning, which solves the problem in the prior art that the federated large - language model mainly focuses on knowledge transfer between the client - side and the server - side, but ignores the enrichment effect of the unique domain knowledge of the client - side small model on the server - side large model.
[0031] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the drawings.
[0032] Figure 1 Flowchart of a fine - tuning method for a federated large model based on mutual learning provided by an embodiment of the present application. As Figure 1 shown, a fine - tuning method for a federated large model based on mutual learning provided by an embodiment of the present application specifically includes the following steps:
[0033] Step 10: Fine - tune the first language model through the client and the local private dataset to obtain the second language model;
[0034] In this step, several clients use the local private dataset to fine - tune their respective small - language models. For example, the client private dataset contains 10,245 maintenance dialogue records in the electronic field, involving questions and answers between maintenance personnel and equipment holders, covering various electronic problems, including electronic device card ejections, specific usage instructions, and the private dataset has been cleaned and pre - processed to remove irrelevant information and noise.
[0035] As an alternative embodiment, the first language model is fine-tuned through a client and a local private dataset to obtain a second language model, which may specifically include: Step 101: Introduce two low-rank matrices in the attention mechanism layer and the feed-forward neural network layer of the first language model.
[0036] In this step, two low-rank matrices A and B are introduced in the decisive levels of the model, such as the attention mechanism layer or the feed-forward neural network layer of Transformer. The ranks of these matrices are usually much lower than the rank of the original matrix, so they have fewer parameters.
[0037] Step 102: Initialize the low-rank matrices and use the local private dataset to train the low-rank matrices to obtain a fine-tuned model.
[0038] In this step, for the private dataset of each client, it is input into the first language model, and the loss of the model on the task is calculated. During the backpropagation process, only the low-rank matrices A and B are updated, while other parameters of the original model remain unchanged. Through multiple iterations of training, the low-rank matrices will gradually learn the specific features of the electronic device conversation, thereby achieving fine-tuning of the model. For example, the original model parameter W can be decomposed into the product of two low-rank matrices A and B plus a very small residual matrix ΔW, that is, W = A * B^T + ΔW. First, it can be understood that ΔW represents the residual matrix, which is a very small matrix used to capture the part that cannot be fully represented by A * B^T in W; the letter T represents the transpose operation, indicating the transpose matrix of matrix B, in order to enable matrix A and B to be multiplied. Then, during the fine-tuning process, only A, B, and ΔW are updated, while most of the parameters of the original model remain unchanged. Assume that the size of the original weight matrix W is d×d, and two low-rank matrices A (size d×r) and B (size r×d) are defined. The original weight matrix W is a 4×4 square matrix (i.e., d = 4), and a smaller rank r = 2 is selected. The low-rank matrix A (size 4×2) and B (size 2×4), and the residual matrix ΔW is also a 4×4 square matrix, but its element values are very small and close to zero initially. The learning rate is set to 0.0001. Further, the predetermined number of iterations is 1000 times. During the training process, a batch of conversation records is used in each iteration to calculate the gradient and update the parameters of the low-rank matrices A, B, and the residual matrix ΔW. Fine-tuning can also include updating the weights of the model, adding additional layers, or adjusting the architecture of the model. In the embodiment of this application, it is to update the parameters, which can be determined according to the actual situation. As the iteration progresses, the loss value will gradually decrease, indicating that the model's ability to understand the electronic device conversation is gradually improving.
[0039] As an alternative embodiment, the method may further include: Step 103: Evaluate the accuracy and efficiency of the fine-tuned model through synonym replacement;
[0040] In this step, for example, replace "mobile phone" with "cellular phone", "screen" with "display screen", etc., and calculate the accuracy between the model prediction result and the true label. For example, the test set contains 100 samples, 50 of which belong to the electronic product category. After synonym replacement, the model correctly predicts 48 samples in the electronic product category. Then the accuracy is: 48 / 50 = 96%; for the generation task, indicators such as the quality, fluency, and matching degree of the generated text with the input text can be evaluated. For example, use the BLEU (Bilingual Evaluation Understudy) score to evaluate the quality of the generated text. Assume the BLEU score ranges from 0 to 1, where 1 represents a perfect match. After evaluation, the BLEU score of the generated text is 0.85, indicating a high similarity between the generated text and the reference text; at the same time, record the time required for the model to process the test set to evaluate its efficiency. For example, the total time for the model to process 100 samples is 10 seconds. Then the throughput of the model is: 100 / 10 = 10 samples / second, or the average time to process each sample is: 10 / 100 = 0.1 second / sample, as Figure 2 shown.
[0041] Step 104: When the accuracy and efficiency reach the preset threshold, save the current low-rank matrix to obtain a second language model.
[0042] In this step, for example, for the prediction task (predicting repair steps), the accuracy may be required to be not less than 90%; for the generation task (generating repair steps), the quality score of the generated text may be required to be not less than 85 points, and the processing time does not exceed 1 second.
[0043] Step 20: Calculate the second result for the public dataset through the second language model, where the public dataset already has the first result, and the second result includes feature vectors, attention distributions, and gradient information;
[0044] That is to say, the purpose of the embodiments of the present application is to input maintenance requirements and obtain maintenance steps. Each client has its own dedicated field. In this step, the public dataset is the decrypted maintenance process conversations of all items, and each conversation is attached with a manually revised reference answer. The first result is the reference answer, and its feature vectors, etc. are all known. All conversations are input into the second language model, and through natural language processing, the following are obtained for the current conversation respectively: 1. Feature vectors, the feature representations generated inside the model, which can be used for subsequent classification, clustering, or similarity calculation tasks; 2. The attention weight distribution of each part when the model processes the input text, which helps to understand the key points of the model's attention to the input data; 3. The gradient values calculated during the model training process, which can be used to analyze the training dynamics and optimization direction of the model.
[0045] Step 30: Calculate the first loss value between the second result and the first result, and form the first knowledge set with the second result and the first loss value;
[0046] In this step, the mean square error (MSE) can be used as the loss function to calculate the loss value of the above-mentioned feature vectors. MSE =1 / n , n is the dimension of the feature vector. yi is the i th eigenvalue in the first result. y ^ i is the i th eigenvalue in the second result. Assuming the first result is [1, 2, 3, 4, 5, 6, 7, 8, 9, 10] and the second result is [1.1, 1.9, 3.2, 3.8, 5.1, 6.3, 6.8, 8.2, 8.9, 10.1], the first loss value MSE =101[(1 - 1.1) 2 +(2 - 1.9) 2 +…+(10 - 10.1) 2 =0.05. As Figure 3 shown, after the client fine-tuning is completed, each client calculates a knowledge set respectively on the same public dataset. As Figure 4 shown, each element in the set is a loss - original output pair, and the total number of elements in the set is the number of data in the dataset.
[0047] Step 40: Obtain the knowledge set and calculate the public dataset through the third language model to obtain the third result;
[0048] In this step, for example, the third language model can be Llama2-13B, which is a pre-trained global model that has been pre-trained on a large amount of text data and then outputs feature vectors; each client uploads the knowledge set calculated by itself to the server; the server converts the original output in the knowledge set uploaded by the client into the original output of the server model according to the vocabulary mapping table between the server and each client model.
[0049] Step 50: Calculate the second loss value of the third result, and compare the first loss value and the second loss value to fine-tune the third language model.
[0050] In this step, as Figure 5 shown, for example, the third result feature vector is [0.9, 2.1, 2.9, 4.2, 4.9, 6.1, 7.2, 7.8, 9.1, 9.8], MSE = 101[(1 - 0.9) 2 +(2 - 2.1) 2 +…+(10 - 9.8) 2 , assuming the calculated MSE is 0.08, if the second loss value (0.08) is greater than the first loss value (0.05), then the knowledge of the second language model (through knowledge distillation) can be considered to fine-tune the third language model.
[0051] As an alternative embodiment, calculating the second loss value of the third result and comparing the first loss value and the second loss value to fine-tune the third language model may specifically include: Step 501: Obtain the minimum first loss value, and compare the minimum first loss value and the second loss value according to the vocabulary mapping table with the client;
[0052] As an alternative embodiment, as Figure 6 shown, obtaining the minimum first loss value and comparing the minimum first loss value and the second loss value according to the vocabulary mapping table with the client may specifically include:
[0053] Step 5011: Based on the first loss values of the first result including the calculated feature vector, attention distribution, and gradient information, compare the first loss values of the feature vector, attention distribution, and gradient information respectively;
[0054] In this step, for example, if the first loss value of the feature vector of client 1 is the smallest and the first loss value of the attention distribution of client 2 is the smallest, then the feature vector of client 1 and the attention distribution of client 2 are respectively selected.
[0055] Step 5012: Select the output forms with the smallest first loss values of the feature vector, attention distribution, and gradient information respectively, and combine them to obtain the minimum first loss value.
[0056] In this step, the feature vector of Client 1, the attention distribution of Client 2, etc. are combined to obtain the smallest first loss value.
[0057] Step 502: If it is less than the second loss value, obtain the first knowledge set of this client and fine-tune the third language model.
[0058] In this step, the server compares the losses of each client's inference on each piece of data, takes the loss value of the client with the smallest loss, and compares it with the loss calculated by the server-side model. If it is less than the loss value calculated by the server-side, then the original output after conversion of this client on this data is incorporated into the server-side knowledge base, and the server-side model is fine-tuned using these knowledge bases.
[0059] As an optional embodiment, after calculating the second loss value of the third result and comparing the first loss value and the second loss value to fine-tune the third language model, the method may further include:
[0060] Step 60: Combine the third result and the second loss value to form a second knowledge set;
[0061] In this step, the server calculates the loss value and the original output of the inference of the local large language model for each piece of data in the public dataset, and also forms a knowledge set of loss-original output pairs.
[0062] Step 70: Send the third result to the client according to the vocabulary mapping table with the client.
[0063] In this step, the server sends the knowledge set on the public dataset to each client; each client converts the original output in the knowledge set sent by the server into the original output of its own model according to the vocabulary mapping table between itself and the server; each client compares the loss of the server on each piece of data with the loss calculated by its local model. If the server loss is smaller, then the original output after conversion of the server on this data is incorporated into the knowledge base of this client, and the respective models are fine-tuned using these knowledge bases; as Figure 7 shown, for example, on data item A of Client 1, the loss of the server L_server_A = 0.025, while the loss of the local model L_local_A = 0.03. Because L_server_A < L_local_A, Client 1 incorporates the original output after conversion of the server on data item A into its knowledge base and fine-tunes the local model using this knowledge. Refer to Figure 8 , in this way, finally select the knowledge set, parameters, and knowledge with smaller loss values (through knowledge distillation) to fine-tune the large language model on the local or server side. As Figure 9 shown, the accuracy rate has increased.
[0064] As an alternative embodiment, before fine-tuning the first language model through the client and the local private dataset to obtain the second language model, the method may further include: Step 01: Collect the local private dataset and the public dataset, extract the vocabulary, and count the occurrence frequency and context information of the vocabulary; Step 02: Determine the mapping relationship between the vocabulary based on the similarity of the vocabulary and the context information to obtain a vocabulary mapping table.
[0065] As an alternative embodiment, the method may further include: Step 03: Remove duplicate vocabulary and merge vocabulary with a similarity greater than a preset threshold to periodically update the vocabulary mapping table; Step 04: Establish a user feedback mechanism and a sharing mechanism to allow different clients to share the first knowledge set.
[0066] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiments of this application also provide a federated large model fine-tuning device based on mutual learning, and its structure is as Figure 10 shown.
[0067] Figure 10 FIG. is a schematic internal structure diagram of a federated large model fine-tuning device based on mutual learning provided by an embodiment of this application. As shown in Figure 10 shown, the device includes:
[0068] At least one processor 1001;
[0069] And a memory 1002 communicatively connected to the at least one processor;
[0070] Wherein, the memory 1002 stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor 1001 so that the at least one processor 1001 can: fine-tune the first language model through the client and the local private dataset to obtain the second language model; calculate the public dataset through the second language model to obtain a second result, wherein the public dataset already has a first result, and the second result includes feature vectors, attention distributions, and gradient information; calculate a first loss value between the second result and the first result, and form a first knowledge set with the second result and the first loss value; obtain the knowledge set, and calculate the public dataset through a third language model to obtain a third result; calculate a second loss value of the third result, and compare the first loss value and the second loss value to fine-tune the third language model.
[0071] Some embodiments of this application provide corresponding to Figure 1A non-volatile computer storage medium for fine-tuning a federated large model based on mutual learning stores computer-executable instructions, which are set to: fine-tune a first language model through a client and a local private dataset to obtain a second language model; calculate a second result for a public dataset through the second language model, where the public dataset already has a first result, and the second result includes a feature vector, an attention distribution, and gradient information; calculate a first loss value between the second result and the first result, and form a first knowledge set with the second result and the first loss value; obtain the knowledge set, and calculate a third result for the public dataset through a third language model; calculate a second loss value of the third result, and compare the first loss value and the second loss value to fine-tune the third language model.
[0072] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0073] The systems and media provided by the embodiments of this application correspond one-to-one with the methods. Therefore, the systems and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be elaborated here.
[0074] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0075] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or blocks Figure 1means for the functions specified in one or more blocks.
[0076] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.
[0077] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.
[0078] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0079] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0080] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0081] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0082] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for fine-tuning a federated large model based on mutual learning, characterized in that: The method comprises: The first language model is fine-tuned through the client and the local private dataset to obtain the second language model, which includes: In the attention mechanism layer and feedforward neural network layer of the first language model, two low-rank matrices are introduced; Initializing the low-rank matrix, and using the local private data set to train the low-rank matrix to obtain a fine-tuned model; the first language model is a client language model; Calculate the public data set by using the second language model to obtain a second result, wherein the public data set already has the first result, and the second result includes a feature vector, attention distribution, and gradient information; Calculate the first loss value of the second result and the first result, and form the second result and the first loss value into a first knowledge set; Acquire the knowledge set, and calculate the public data set through a third language model to obtain a third result, where the third language model is a server-side language model; Calculating a second loss value of the third result and the first result, and comparing the first loss value with the second loss value to fine-tune the third language model, specifically comprising: Obtaining a minimum first loss value, and comparing the minimum first loss value with a second loss value according to a vocabulary mapping table with the client; If the first loss value is less than the second loss value, obtaining a first knowledge set of the client, and fine-tuning the third language model; Combining the third result and the second loss value into a second knowledge set; The third result is sent to the client according to the vocabulary mapping table with the client.
2. According to the mutual learning-based federated large model fine-tuning method of claim 1, it is characterized in that: Before fine-tuning the first language model through the client and the local private data set to obtain the second language model, the method further includes: Collect local private data sets and public data sets, extract vocabulary, and count the occurrence frequency and context information of the vocabulary; Based on the similarity of the words and context information, a mapping relationship between the words is determined to obtain the word mapping table.
3. According to claim 2, a method for fine-tuning a federated large model based on mutual learning is characterized in that: The method further comprises: Deduplication processing is performed on repeated words, and words with similarity greater than a preset threshold are merged to regularly update the word mapping table; A user feedback mechanism and a sharing mechanism are established to allow the first knowledge set to be shared between different clients.
4. According to the mutual learning-based federated large model fine-tuning method of claim 1, it is characterized in that: The method further comprises: By synonym replacement, evaluating the accuracy and efficiency of the fine-tuned first language model; When the accuracy and efficiency reach a preset threshold, the current low-rank matrix is saved to obtain the second language model.
5. According to the mutual learning-based federated large model fine-tuning method of claim 1, it is characterized in that: Obtaining a minimum first loss value, and comparing the minimum first loss value with a second loss value according to a vocabulary mapping table with the client, specifically includes: Based on the first result, the first loss values of the feature vector, the attention distribution and the gradient information are calculated, and the first loss values of the feature vector, the attention distribution and the gradient information are compared respectively; The output forms with the smallest first loss values of the feature vector, attention distribution, and gradient information are respectively selected, and combined to obtain the minimum first loss value.
6. A federated large model fine-tuning device based on mutual learning, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Execute the steps of a method for fine-tuning a federated large model based on mutual learning as described in any one of claims 1-5.
7. A non-volatile computer storage medium for fine-tuning a federated large model based on mutual learning, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Execute the steps of a method for fine-tuning a federated large model based on mutual learning as described in any one of claims 1-5.
Citation Information
Patent Citations
Cross-language social media event detection method based on federal knowledge distillation
CN118586431A
Reverse knowledge distillation method and system based on federal large model
CN119129708A