Federal learning training method and device for code task-oriented large language model
Through the combination of federated learning and adapter models, the problem of scarcity of computing resources and data in the application of large language models in specific fields is solved, efficient model training and performance improvement are achieved, and the ability to adapt to different programming languages is achieved.
Patent Information
- Application Number
- CN202510409307.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
AI Technical Summary
Large language models have degraded performance when applied in specific fields, limited computing resources and scarce high-quality code samples. The existing technology solutions lead to excessive training cycles of model and high data processing requirements.
Combined with federated learning technology, by screening organization nodes participating in training and pre-processing their data, fine-tuning the large language model using a lightweight adapter model, reducing unnecessary data screening stages, and inserting code language tags during the training process to distinguish different programming languages.
It improves the training efficiency and performance of large language models in specific fields, reduces computing resource consumption and model training cycles, and enhances the model's adaptability to different programming languages.
Smart Images

Figure CN120278232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and particularly to a federated learning training method and device for large language models for code tasks. Background Art
[0002] The application of large language models (LLMs) in the field of machine learning has shown significant potential, especially in tasks such as code generation and code translation. LLMs represented by GPT, CodeBERT, Codex, etc., through pre-training on a large amount of code and text, have demonstrated remarkable capabilities in tasks such as code generation, code completion, and cross-language code translation.
[0003] However, when LLMs are applied to specific fields, such as financial system development, embedded software design, or privatized software development tasks, their performance often drops significantly. This problem can be attributed to two core reasons: on the one hand, the pre-training data of LLMs mainly comes from public open-source code repositories (such as GitHub), making it difficult to cover specific domain or enterprise private code logic, specific design patterns, or internal API call specifications; on the other hand, there are significant deviations in the statistical characteristics of the data distribution of specific tasks and public code data (such as the real-time constraints of medical device control code), resulting in limited model generalization ability. To address the above problems, domain adaptation fine-tuning of LLMs has become a necessary means, but this process faces two challenges: firstly, the problem of limited computing resources, as the number of parameters of mainstream LLMs reaches tens of billions or even hundreds of billions, and the cluster resources required for full-parameter fine-tuning far exceed the capabilities of most organizations; secondly, the scarcity of high-quality code samples for specific tasks, and enterprises are restricted by code asset confidentiality and compliance requirements, making it difficult to build a large-scale training set through data sharing.
[0004] To address the above problems, the prior art has proposed combining federated learning technology to fine-tune the adapter part of LLMs to adjust model parameters by obtaining a large number of high-quality models while ensuring the privacy of training data; and only fine-tuning the adapter part of LLMs to reduce the burden of computing resources. For example, the industry large model training method and system based on efficient fine-tuning and federated learning proposed in CN118982074A discloses fine-tuning large models by combining federated learning; the large model training and task processing method based on federated learning proposed in CN117350411A discloses combining task processing to select clients consistent with the task type and then fine-tuning large models by combining federated learning.
[0005] However, although the above technical solution solves the problems of computing resources and data privacy in large model training, during the large model training process, a large amount of data needs to be analyzed, interpreted, and screened. The model training end needs to be equipped with an efficient data analysis and screening module, which still sets relatively high data processing requirements for the model training end. If the data analysis or interpretation process is too long, it is likely to lead to an overly long model training cycle, which is not conducive to the further application of large models in specific fields. Additionally, there are multiple coding languages in existing code tasks, and each coding language has specific built-in logics or data structures. When each client conducts model training for different code tasks, it needs to traverse the possibilities of multiple coding languages, which also easily leads to an overly long model training cycle. Summary of the Invention
[0006] Based on this, the present invention provides a federated learning training method and device for large language models for code tasks, which combines code tasks to screen organizational nodes participating in large language model training, and structurally preprocesses the training data of organizational nodes, reducing the screening stage for unnecessary data during the large language model training process. Furthermore, by performing tagging processing on the specific code languages of each organizational node, parameters can be quickly adjusted for the training data of specific code languages during the large language model training process, without the need for overly complicated code language traversal or recognition processes, improving the efficiency of large language model training.
[0007] In the first aspect, the present invention provides a federated learning training method for large language models for code tasks. The federated learning training method is applied to a federated learning system, which includes a server node and several organizational nodes. The federated learning training method is executed by the organizational nodes and includes the following steps:
[0008] Step S101, preprocess the local data according to the code task to obtain structured training data;
[0009] Step S102, negotiate training conditions, a base model, and an adapter model with the server node according to the code task;
[0010] Step S103, receive the base model sent by the server node;
[0011] Step S104, receive the initial adapter model sent by the server node, and fuse the initial adapter model and the base model to obtain an initial global model;
[0012] Step S105, train the initial global model according to the structured training data, and adjust the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model;
[0013] Step S106: Send the number of the fine-tuning adapter models and the structured training data to the server node, so that the server node aggregates all the fine-tuning adapter models to update the initial adapter model;
[0014] Step S107: Repeat Steps S104 - S106 until the training condition is satisfied, so that the server node aggregates all the fine-tuning adapter models and the base model to obtain the trained global model.
[0015] Further, before executing Step S103, it further includes:
[0016] Correct the structured training data by inserting code language tags.
[0017] Further, the structured training data includes code content and annotation content.
[0018] Further, when training the initial global model according to the structured training data and adjusting the parameters of the initial adapter model in the initial global model to obtain the specific expression of the optimization objective of the initial adapter model in the fine-tuning adapter model is:
[0019] ,
[0020] Wherein, is the fine-tuning adapter model, is the number of batches of structured training data of the organization node, is the size of the structured training data batch, is the loss function, is the base model, is the fine-tuning adapter model of the previous training round, is the regularization coefficient, is the th structured training data in the
[0021] In a second aspect, the present invention further provides a federated learning training method for a large language model for code tasks. The federated learning training method is applied to a federated learning system. The federated learning system includes a server node and a plurality of organization nodes, and includes the following steps:
[0022] Step S201: The server node randomly selects a plurality of organization nodes as training organization nodes according to the code task;
[0023] Step S202: Each organization node preprocesses the local data according to the code task to obtain structured training data;
[0024] Step S203: The training organization node and the server node negotiate to determine the base model, adapter model, and training conditions for joint training.
[0025] Step S204: The server node sends the base model and the initial adapter model to each training organization node.
[0026] Step S205: Each training organization node fine-tunes the adapter model based on local data to obtain a fine-tuned adapter model, and sends the fine-tuned adapter model and the quantity of structured training data to the server node.
[0027] Step S206: The server node aggregates the fine-tuned adapter models and the quantities of structured training data feedback by each training organization node to obtain a trained adapter model.
[0028] Step S207: Send the trained adapter model as the initial adapter back to each training organization node, and repeat steps S205 - S206 until the training conditions are met. The server node fuses the trained adapter model and the base model to obtain a trained global model.
[0029] Further, the server node aggregates the fine-tuned adapter models and the quantities of structured training data feedback by each training organization node to obtain a trained adapter model. The specific expression is:
[0030] ,
[0031] where, is the trained adapter model in the th training round, is the number of training organization nodes, is the quantity of structured training data of the th training organization node, is the total quantity of structured training data in the th training round, is the fine-tuned adapter model feedback by the th training organization node in the th training round.
[0032] Further, the structured training data includes code content and annotation content.
[0033] Further, before executing step S205, it also includes:
[0034] Correct the structured training data by inserting code language tags.
[0035] Third aspect, a federated learning training device for large language models for code tasks, the federated learning training device is applied to a federated learning system, the federated learning system includes a server node and several organizational nodes, and the federated learning training device is executed by the organizational nodes, including:
[0036] A data preprocessing module, configured to preprocess local data according to the code task to obtain structured training data;
[0037] A model selection module, configured to negotiate training conditions, a base model, and an adapter model with the server node according to the code task;
[0038] A base model receiving module, configured to receive the base model sent by the server node;
[0039] A model fusion module, configured to receive the initial adapter model sent by the server node, and fuse the initial adapter model and the base model to obtain an initial global model;
[0040] A parameter adjustment module, configured to train the initial global model according to the structured training data, and adjust the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model;
[0041] A model update module, configured to send the fine-tuned adapter model and the quantity of the structured training data to the server node, so that the server node aggregates all the fine-tuned adapter models to update the initial adapter model;
[0042] A global model output module, configured to repeatedly execute the model fusion module, the parameter adjustment module, and the model update module until the training conditions are met, so that the server node aggregates all the fine-tuned adapter models and the base model to obtain a trained global model.
[0043] Fourth aspect, a federated learning training device for large language models for code tasks, the federated learning training device is applied to a federated learning system, the federated learning system includes a server node and several organizational nodes, including:
[0044] An organizational node extraction module, configured to randomly extract several organizational nodes as training organizational nodes by the server node according to the code task;
[0045] A data preprocessing module, configured to preprocess local data by each organizational node according to the code task to obtain structured training data;
[0046] A model selection module, configured to negotiate and determine a jointly trained base model, adapter model, and training conditions by the training organizational nodes and the server node according to the code task;
[0047] A model sending module, configured to enable the server node to send a base model and an initial adapter model to each training organization node;
[0048] A parameter adjustment module, configured to enable each of the training organization nodes to fine-tune the adapter model according to local data to obtain a fine-tuned adapter model, and send the fine-tuned adapter model and the quantity of structured training data to the server node;
[0049] An adapter aggregation module, configured to enable the server node to aggregate according to the fine-tuned adapter models and the quantities of structured training data fed back by each training organization node to obtain a trained adapter model;
[0050] A global model output module, configured to re-send the trained adapter model to each training organization node as an initial adapter, and repeatedly execute the parameter adjustment module and the adapter aggregation module until the training condition is met, and the server node fuses the trained adapter model and the base model to obtain a trained global model.
[0051] In a fifth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the federated learning training method for a large language model for code tasks in any one of the first aspect and the second aspect are implemented.
[0052] In a sixth aspect, the present invention further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it executes the federated learning training method for a large language model for code tasks in any one of the first aspect and the second aspect.
[0053] The beneficial effects of adopting the above technical solutions are as follows: In this embodiment, the organization nodes participating in the training of the large language model are screened in combination with code tasks, and the training data of the organization nodes is subjected to structured preprocessing, reducing the screening stage of unnecessary data in the training process of the large language model. Furthermore, considering that different organization nodes use different code languages, direct fine-tuning may cause interference between different code languages. By adding tags to the code languages, a fine-tuning method under heterogeneous code data is designed to improve the performance of the model after federated learning fine-tuning. Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art.
[0055] Figure 1 It is a schematic diagram of a federated learning system architecture in an embodiment of the present application;
[0056] Figure 2 Schematic diagram of the training process of the organization node model in an embodiment of the present application;
[0057] Figure 3 Schematic diagram of the training process of the federated learning system model in an embodiment of the present application. Detailed implementation manners
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. In order to describe the present invention in more detail, the federated learning training method and device for a large language model for code tasks provided by the present invention will be specifically described below in conjunction with the accompanying drawings.
[0059] Unless otherwise defined, the technical terms or scientific terms used in the disclosure of the present application should have the ordinary meaning understood by those of ordinary skill in the art in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an" or "the" do not indicate a quantity limitation, but mean that there is at least one. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "up", "down", "left", "right" are only used to represent relative positional relationships, and when the absolute position of the object to be described changes, the relative positional relationship may also change accordingly.
[0060] The present invention provides a federated learning training method for a large language model for code tasks, which screens organization nodes participating in the training of the large language model in combination with code tasks, and performs structured preprocessing on the training data of the organization nodes to reduce the screening stage of unnecessary data in the training process of the large language model. Further, the specific code languages of each organization node are labeled, and the parameters can be quickly adjusted for the training data of the specific code language during the training process of the large language model, without the need for a too complicated code language traversal or recognition process, improving the training efficiency of the large language model. Taking the application of this method to a federated learning system as an example for illustration.
[0061] Combined with the attached Figure 1The shown federated learning system includes a server node and several organizational nodes. Suppose in a specific downstream task of software engineering, there are a total of N organizational nodes participating in fine-tuning a large language model. They have high-quality code-related data available for model fine-tuning locally, but due to competition between organizational nodes or certain privacy policies, this data is prohibited from being used outside the organizational nodes. By jointly executing the federated learning training method of the large language model for code tasks in this embodiment by the organizational nodes, it is possible to efficiently jointly fine-tune the large language model while maintaining the privacy of local data. The organizational nodes are composed of various data owners and are mainly responsible for data collation and model fine-tuning. The server node is a trusted server that undertakes the tasks of model aggregation and distribution.
[0062] It should be noted that the above server node and organizational nodes are terminal devices with a storage structure. The terminal device includes, but is not limited to, a smart phone and a computer device, where the computer device can be at least one of a desktop computer, a portable computer, a laptop computer, a mainframe computer, a tablet computer, and other devices. The user operates the terminal device, and through the cooperation of each organizational node and the server node, a trained large language model is obtained. For the specific process, please refer to the federated learning training method of the large language model for code tasks.
[0063] Combined with the Figure 2 training process of the organizational nodes shown, the federated learning training method of the large language model for code tasks in this embodiment will be described.
[0064] Step S101, preprocess the local data according to the code task to obtain structured training data.
[0065] Specifically, preprocessing the local data according to the code task includes, but is not limited to, data cleaning and data extraction steps. After preprocessing, the local data obtains structured training data related to the code task. Among them, data cleaning of the local data includes eliminating the noise of the local data and improving the consistency of the local data. For example, performing missing value processing, deduplication processing, outlier processing, format standardization, data type correction, or logical correction on the local data. Through the above data cleaning, the organizational nodes provide high-quality local data for large language model training; data extraction of the local data is to extract the necessary data for large language model training based on the specific requirements or task purposes of the code task on the basis of data cleaning. For example, performing field screening and condition filtering on the local data after data cleaning, reducing the processing of invalid data during the training process of the large language model, and improving the training efficiency of the large language model.
[0066] Taking the software development task related to the database as an example, this code task requires the assistance of a large language model to generate code comments. Therefore, the organizational node filters out the code content unrelated to the software development task after data cleaning, extracts the code areas related to the local data after cleaning and the corresponding comment areas, and obtains structured training data. Further, the structured training data can be recorded in the form of fragments of "[code content][comment content]". It should be noted that this application does not stipulate the positions of each module in the structured training data, but within the same organizational node, the structured training data can only be recorded in a unified format.
[0067] Step S102, negotiate the training conditions, base model, and adapter model with the server node according to the code task.
[0068] Specifically, a large language model is a deep learning model that pre-trains on a large amount of text data to learn the complex patterns and structures of language. Large language models usually have a large number of parameters, which can be billions or even hundreds of billions, enabling them to understand and generate natural language text. Large language models perform well in a variety of natural language processing tasks, such as text classification, question answering systems, machine translation, and text generation. Currently, the vast majority of large language models are implemented based on the Transformer architecture. The Transformer architecture allows the model to consider the information at all positions in the sequence while processing sequence data, and can more effectively capture long-distance dependencies.
[0069] The base model of a large language model refers to the main part of the pre-trained model, which contains the main architecture and parameters of the model. It is usually obtained by pre-training on a large-scale dataset and has rich language knowledge and understanding ability. In this embodiment, the base model uses a general open-source model, which is a large neural network obtained by self-supervised training on a large amount of unlabeled corpus, such as Llama, GPT-3, PaLM, or a specially designed model. Negotiating the base model with the server node according to the code task means negotiating and selecting a base model that performs well in similar code tasks and has a moderate parameter scale according to the specific code task, the organizational nodes participating in the training, and the server node.
[0070] The adapter model of this embodiment is a lightweight network structure that can be inserted into specific layers of a pre-trained model for fine-tuning, enabling the large language model to adapt to new tasks without significantly increasing the number of parameters and saving computing resources. The adapter model is usually composed of a multi-layer perceptron (MLP) and is inserted into a specific position of the Transformer model. Considering that the core of the Transformer model is the multi-head self-attention mechanism and the feed-forward neural network, and the multi-head self-attention mechanism contains three types of matrices (query matrix, key matrix, and value matrix), the adapter can be inserted into all query matrices, key matrices, and value matrices. For example, for a query matrix , the adapter generates two small matrices and , namely and ( is much smaller than ), so that and are much smaller in size than . When fine-tuning the adapter, the query matrix is frozen, and the matrices and are fine-tuned. This process is the process of inserting the adapter model in this embodiment. All the small matrices and together constitute the adapter model. After fine-tuning, if the preset training effect is achieved, the adapter can be fused into the base model using 's method.
[0071] The training conditions are the end conditions for each organizational node to train the model, such as reaching the preset number of training rounds or the loss function of the model reaching the preset value.
[0072] Step S103: Receive the base model sent by the server node.
[0073] Step S104: Receive the initial adapter model sent by the server node, and fuse the initial adapter model and the base model to obtain an initial global model.
[0074] Step S105: Train the initial global model according to the structured training data, and adjust the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model.
[0075] Specifically, during the training of the initial global model, the parameters of the base model are frozen, and the parameters of the initial adapter model are adjusted. During the parameter adjustment of the adapter model, the adapter model is divided into several low-rank matrices for parameter adjustment.
[0076] For the base model, the th key matrix in the th Transformer module is , where represents a real number field space with a matrix size of . The adapter model injects a low-rank matrix into this matrix to replace for training. consists of two sub-matrices and , where is much smaller than and , and . The adapter is composed of all the matrices injected into the model base. The base model and the adapter model will perform a multiplication operation on the same input , and the calculated vectors will be summed up one by one according to the coordinate correspondence. Assuming that only the base model exists, there is , then the output after using the adapter can be expressed as: .
[0077] For specific modules in the Transformer model, such as the adapter for the query matrix , the value matrix and the output matrix , the way of introducing the rank decomposition matrix is similar to the above key matrix, which are , , , and , respectively. The combination represents the adapter model.
[0078] During the training phase, the parameters of the adapter are optimized through the backpropagation algorithm to minimize the loss function of a specific task. At the same time, the parameters of the base model remain unchanged and do not need to participate in the training, saving a large amount of computing resources. The introduction of this adapter structure allows the model to be fine-tuned for specific tasks while keeping the parameters of the basic Transformer module unchanged. This means that once the basic model is trained, the model fine-tuned through the adapter can quickly adapt to new tasks while retaining the general knowledge learned from a large amount of data.
[0079] Specifically, when training a model in the field of deep learning, the Mini-batch Gradient Descent (MBGD) is a commonly used method. MBGD only samples a part of the training samples in the dataset for training and updating parameters each time, taking into account the advantages and disadvantages of both the Stochastic Gradient Descent and the Batch Gradient Descent. However, the MBGD method is not very suitable for the present invention because each update of the global model in a distributed environment requires a large amount of communication resources. For the purpose of saving communication resources, in this embodiment, the local data is divided into different batches according to the batch size B before training. When using these batches for training, the global model parameters are not updated, but only the adapter model is updated. Only after all batches of training are completed, the global model aggregation is uniformly executed. At the same time, due to the unstable gradients during large model training, the divergence of the model gradients may be exacerbated when different organizations use different data for fine-tuning. In this embodiment, the L2 norm is used to constrain the divergence of the model during the parameter adjustment process of the initial adapter model. The specific expression of the optimization objective of the initial adapter model is:
[0080] ,
[0081] where, is the fine-tuned adapter model, is the number of batches of structured training data of the organizational node, is the size of the structured training data batch, is the loss function, is the base model, is the fine-tuned adapter model of the previous training round, is the regularization coefficient, is the th structured training data in the th training batch.
[0082] For example, taking the code task as intelligent question answering as an example, the above loss function selects the cross-entropy loss function. For each time step (i.e., each word in the sequence), the model outputs a probability distribution indicating the probability that the next word is any word in the vocabulary. The cross-entropy loss function calculates the difference between this predicted probability distribution and the true distribution of the actual word. The true distribution here is a one-hot encoded vector representing the actual next word.
[0083] Step S106, send the number of the fine-tuned adapter models and the structured training data to the server node, so that the server node aggregates all the fine-tuned adapter models to update the initial adapter model.
[0084] Step S107, repeatedly execute Step S104 - Step S106 until the training condition is met, so that the server node aggregates all the fine-tuned adapter models and the base model to obtain the trained global model.
[0085] Based on the above embodiments, considering that the organizational nodes have a multilingual environment, before executing Step S103, it further includes: correcting the structured training data by inserting code language tags.
[0086] Specifically, insert code language tags into the structured training data. For example, if the local data of an organizational node uses the Python language, then this organizational node needs to add the code language tag "[Code Language: Python]" to the structured training data. By this method, the large language model can identify and distinguish codes in different programming languages, and learn their common features and patterns during the training process. The insertion of tags can be achieved through automated detection tools or manual correction to ensure that each piece of structured data has an accurate language identifier. In the experiment, four different language organizations were used for joint fine-tuning. The results show that compared with the same language environment, not adding code tags will cause a performance drop of about 10% in the model, and different languages will interfere with each other; while adding code language tags enables the model to distinguish codes in different programming languages and learn the common patterns and features of different programming languages, and the model performance is improved by 43%.
[0087] Furthermore, in this embodiment, before inputting the structured training data into the initial global model, text auxiliary processing can also be set, which is implemented through a prompter and a tokenizer. The role of the prompter is to provide guiding prompts or examples for the initial global model to help the initial global model better understand the task or problem. The prompter can be some text fragments, problem descriptions, or example input-output, etc., which helps to guide the initial global model to generate more expected results. For example, when inputting the data "[Code Content][Comment Content]", add a prompt word and modify it to: "[This is a code about a database. Please generate comment content based on the code][Code Content][Comment Content]". The prompter adds guiding information to the data fragment, provides clear task goals and examples for the model, and thus guides the large language model to better understand the context information and task requirements contained in the data.
[0088] The role of the tokenizer is to convert the text data into a digital vector representation that the model can understand, that is, to split the text into words or sub-words and convert them into corresponding digital vectors. The tokenizer preprocesses the input text data to make it meet the input requirements of the initial global model, and at the same time converts the output of the initial global model into a readable text result.
[0089] Combined with the appendixFigure 3 Schematic diagram of the federated learning training method for large language models oriented to code tasks. Taking the federated learning system executing a specific federated learning training method as an example for illustration.
[0090] Step S201: The server node randomly selects a number of organizational nodes as training organizational nodes according to the code task.
[0091] Specifically, in this embodiment, there are multiple ways to randomly select a number of organizational nodes as training organizational nodes according to the code task, including but not limited to selecting organizational nodes according to the historical processing of similar code tasks, or selecting organizational nodes according to the similar structure of local data, or the server node announces a mechanism for task organizational nodes to actively sign up, or a metadata-driven dynamic screening mechanism.
[0092] Among them, the mechanism for the server node to announce the task organizational node to actively sign up is: the server node makes the public task code annotation generation, and each organizational node signs up according to its own needs.
[0093] The metadata-driven dynamic screening mechanism is: each organizational node maintains a set of metadata to describe the local data characteristics, such as programming language type, framework version, project scale, annotation coverage rate, etc. After the server node receives the code task, it parses the task requirements, extracts the key metadata fields (such as target language, required code library scale, specific technology stack, etc.), and then the server node matches the metadata tags of the organizational nodes according to the extracted task metadata fields, calculates the similarity through multi-dimensional feature vectors, and screens out the most matching organizational nodes. For example: map the task requirements and node metadata to vectors , , calculate the cosine similarity . A specific example can be: the task requirement metadata is Python(1), Machine Learning(1), Data Volume(0.8); the metadata of organizational node A is Python(1), Machine Learning(0.9), Data Volume(0.6); calculate the cosine similarity between the two as , through this method, the matching degree between the task and the organizational node can be quantified, and the organizational nodes with higher similarity are preferentially selected to participate in the training, such as selecting organizational nodes with a cosine similarity higher than 0.4 to participate in the training.
[0094] Step S202: Each organizational node preprocesses the local data according to the code task to obtain structured training data.
[0095] Specifically, the preprocessing of local data according to the code task includes, but is not limited to, data cleaning and data extraction steps. After preprocessing, the local data obtains structured training data related to the code task. Among them, data cleaning of local data includes eliminating the noise of local data and improving the consistency of local data. For example, missing value processing, duplicate removal, outlier processing, format standardization, data type correction, or logical correction are performed on local data. Through the above data cleaning, the organizational node provides high-quality local data for large language model training. Data extraction from local data is to extract the necessary data for large language model training based on the specific requirements or task objectives of the code task on the basis of data cleaning. For example, field screening, conditional filtering, etc. are performed on the local data after data cleaning, reducing the processing of invalid data during the training process of the large language model and improving the training efficiency of the large language model.
[0096] Taking the software development task related to the database as an example, this code task requires the large language model to assist in generating code comments. Therefore, the organizational node screens out the code content unrelated to the software development task after data cleaning, and extracts the code area and the corresponding comment area related to the cleaned local data to obtain structured training data. Further, this structured training data can be recorded in the form of fragments of "[code content][comment content]". It should be noted that this application does not stipulate the positions of each module in the structured training data, but within the same organizational node, the structured training data can only be recorded in a unified format.
[0097] Step S203, the training organizational node and the server node negotiate and determine the base model, adapter model, and training conditions for joint training according to the code task.
[0098] Specifically, a large language model is a deep learning model that pre-trains on a large amount of text data to learn the complex patterns and structures of language. Large language models usually have a large number of parameters, which can be billions or even hundreds of billions, enabling them to understand and generate natural language text. Large language models perform well in a variety of natural language processing tasks, such as text classification, question answering systems, machine translation, and text generation. Currently, the vast majority of large language models are implemented based on the Transformer architecture. The Transformer architecture allows the model to consider the information at all positions in the sequence while processing sequence data, and can more effectively capture long-distance dependencies.
[0099] The base model of the large language model refers to the main part of the pre-trained model, which contains the main architecture and parameters of the model. It is usually pre-trained on a large-scale dataset and has rich language knowledge and understanding ability. In this embodiment, the base model uses a general open-source model, which is a large neural network obtained through self-supervised training on a large amount of unlabeled corpus, such as Llama, GPT-3, PaLM or a specially designed model. Negotiating the base model according to the code task and the server node means selecting a base model that performs well in similar code tasks and has a moderate parameter scale through negotiation between the specific code task, the participating training organization nodes and the server node.
[0100] The adapter model of this embodiment is a lightweight network structure that can be inserted into specific layers of the pre-trained model for fine-tuning, enabling the large language model to adapt to new tasks without significantly increasing the number of parameters and saving computing resources. The adapter model is usually composed of a multi-layer perceptron (MLP) and is inserted into a specific position of the Transformer model. Considering that the core of the Transformer model is the multi-head self-attention mechanism and the feed-forward neural network, and the multi-head self-attention mechanism contains three types of matrices (query matrix, key matrix, and value matrix), the adapter can be inserted into all query matrices, key matrices, and value matrices. For example, for a certain query matrix , the adapter generates two small matrices and , which are and ( much smaller than ), so that and are much smaller than . When fine-tuning the adapter, the query matrix is frozen, and the matrices and are fine-tuned. This process is the process of inserting the adapter model in this embodiment. All the small matrices and together form the adapter model. After fine-tuning, if the preset training effect is achieved, the adapter can be fused into the base model using 's method.
[0101] The training conditions refer to the end conditions for model training by each organization node, such as reaching a preset number of training rounds or the loss function of the model reaching a preset value.
[0102] Step S204, the server node sends the base model and the initial adapter model to each training organization node.
[0103] Step S205: Each of the training organization nodes fine-tunes the adapter model based on local data to obtain a fine-tuned adapter model, and sends the fine-tuned adapter model and the quantity of structured training data to the server node.
[0104] Specifically, during the initial global model training process, the parameters of the base model are frozen, and the parameters of the initial adapter model are adjusted. During the parameter adjustment process of the adapter model, the adapter model is divided into several low-rank matrices for parameter adjustment.
[0105] For the base model, the th key matrix in the th Transformer module is , where represents the real number field space with matrix size . The adapter model injects a low-rank matrix to replace the training of . consists of two sub-matrices and , where is much smaller than and , and . The adapter is composed of all the matrices injected into the model base. The base model and the adapter model will perform a multiplication operation on the same input , and the calculated vectors are summed up one by one according to the coordinate correspondence. Assuming that only the base model exists, there is , then the output after using the adapter can be expressed as: .
[0106] For specific modules in the Transformer model, such as the query matrix , the value matrix and the output matrix adapters, the way of introducing rank decomposition matrices is similar to the above key matrix, which are , , , and , 's combination represents the adapter model.
[0107] During the training phase, the parameters of the adapter are optimized through the backpropagation algorithm to minimize the loss function for a specific task. Meanwhile, the parameters of the base model remain unchanged and do not need to participate in the training, saving a large amount of computing resources. The introduction of this adapter structure allows the model to be fine-tuned for specific tasks while keeping the parameters of the basic Transformer module unchanged. This means that once the base model is trained, the model fine-tuned through the adapter can quickly adapt to new tasks while retaining the general knowledge obtained from training on a large amount of data.
[0108] Specifically, when training a model in the field of deep learning, the Mini-batch Gradient Descent (MBGD) is a commonly used method. MBGD only samples a part of the training samples in the dataset for training and updating parameters each time, taking into account the advantages and disadvantages of both the Stochastic Gradient Descent and the Batch Gradient Descent. However, the method of MBGD is not very suitable for the present invention because each update of the global model in a distributed environment requires a large amount of communication resources. For the purpose of saving communication resources, in this embodiment, the local data is first divided into different batches according to the batch size B before training. When using these batches for training, the global model parameters are not updated, but only the adapter model is updated. Only after all batches are trained, the global model aggregation is performed uniformly. At the same time, due to the instability of the gradient during large model training, the divergence of the model gradient may be exacerbated when different organizations use different data for fine-tuning. In this embodiment, the L2 norm is used to constrain the divergence of the model during the parameter adjustment process of the initial adapter model. The specific expression of the optimization target of the initial adapter model is:
[0109] ,
[0110] where, is the fine-tuned adapter model, is the number of batches of structured training data of the organization node, is the size of the structured training data batch, is the loss function, is the base model, is the fine-tuned adapter model of the previous training round, is the regularization coefficient, is the th training batch, and is the
[0111] Step S206, the server node aggregates according to the fine-tuned adapter models and the quantities of structured training data fed back by each training organization node to obtain the trained adapter model.
[0112] Among them, the specific expression of step S206 is:
[0113] ,
[0114] wherein, is the trained adapter model after the -th training round, is the number of training organization nodes, is the number of structured training data of the -th training organization node, is the total amount of structured training data in the -th training round, is the fine-tuned adapter model feedback by the -th training organization node in the -th training round.
[0115] Step S207: Resend the trained adapter model to each training organization node as the initial adapter, and repeat steps S205 - S206 until the training conditions are met. The server node fuses the trained adapter model and the base model to obtain the trained global model.
[0116] Further, before performing step S205, it further includes:
[0117] Correct the structured training data by inserting code language tags.
[0118] Before the training starts, the organization nodes need to reach a consensus on the training parameters and processes, including the model base for training, the parameters related to model fine-tuning, the global communication rounds, and the input format of the training data. Subsequently, each organization node needs to download the same model base and perform fine-tuning of the large language model locally. According to the different local computing resources of each organization node, they can choose a suitable mini-batch size of the training data for fine-tuning. During the fine-tuning process, the base model is frozen, and only the adapter model is trained. After the fine-tuning of the current round is completed, the organization node needs to upload the updated adapter model to the server node, and at the same time upload the amount of local data used for fine-tuning. After receiving the model, the server node starts the weighted aggregation of the global adapter. After the aggregation is completed, the global model is sent back to the client, and then it is judged whether the communication rounds at this time reach the prior agreement. If so, the training ends; otherwise, the fine-tuning continues.
[0119] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence according to the arrow indication, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the accompanying drawings may include multiple sub-steps or sub-phases. These sub-steps or phases are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or phases is not necessarily sequential, but can be executed alternately or in rotation with at least a part of other steps or sub-steps or phases of other steps.
[0120] In the above-described embodiments of the present invention, the federated learning training method for large language models for code tasks is described in detail. The above methods disclosed in the present invention can be implemented by various forms of devices. Therefore, the present invention also discloses a federated learning training device for large language models for code tasks. Specific embodiments are given below for detailed description.
[0121] First, for the federated learning training device for large language models for code tasks executed by each organizational node, the federated learning training device is applied to a federated learning system. The federated learning system includes a server node and several organizational nodes, and includes the following modules:
[0122] A data preprocessing module for preprocessing local data according to the code task to obtain structured training data;
[0123] A model selection module for negotiating training conditions, a base model, and an adapter model with the server node according to the code task;
[0124] A base model receiving module for receiving the base model sent by the server node;
[0125] A model fusion module for receiving the initial adapter model sent by the server node and fusing the initial adapter model and the base model to obtain an initial global model;
[0126] A parameter adjustment module for training the initial global model according to the structured training data and adjusting the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model;
[0127] A model update module for sending the fine-tuned adapter model and the quantity of the structured training data to the server node so that the server node aggregates all the fine-tuned adapter models to update the initial adapter model;
[0128] A global model output module, which is used to repeatedly execute a model fusion module, a parameter adjustment module, and a model update module until the training condition is met, so that the server node aggregates all the fine-tuned adapter models and the base model to obtain a trained global model.
[0129] Secondly, for a federated learning training device of a large language model for code tasks executed by an overall federated learning system, the federated learning training device is applied to the federated learning system, and the federated learning system includes a server node and several organization nodes, including:
[0130] An organization node extraction module, which is used for the server node to randomly extract several organization nodes as training organization nodes according to the code task;
[0131] A data preprocessing module, which is used for each organization node to preprocess local data according to the code task to obtain structured training data;
[0132] A model selection module, which is used for the training organization nodes and the server node to negotiate and determine the base model, the adapter model, and the training conditions for joint training according to the code task;
[0133] A model sending module, which is used for the server node to send the base model and the initial adapter model to each training organization node;
[0134] A parameter adjustment module, which is used for each of the training organization nodes to fine-tune the adapter model according to local data to obtain a fine-tuned adapter model, and send the fine-tuned adapter model and the quantity of the structured training data to the server node;
[0135] An adapter aggregation module, which is used for the server node to aggregate the fine-tuned adapter models and the quantities of the structured training data fed back by each training organization node to obtain a trained adapter model;
[0136] A global model output module, which is used to resend the trained adapter model to each training organization node as an initial adapter, and repeatedly execute the parameter adjustment module and the adapter aggregation module until the training condition is met, and the server node fuses the trained adapter model and the base model to obtain a trained global model.
[0137] For the federated learning training device based on the large language model for code tasks, all can refer to the above limitations on the method, and will not be elaborated here. Each module in the above device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the terminal device in hardware form or be independent of it, or be stored in the memory of the terminal device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0138] In one embodiment, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned federated learning training method for the large language model for code tasks are implemented.
[0139] The computer-readable storage medium may be an electronic memory such as flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM (erasable programmable read-only memory), hard disk, or ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for program code for executing any of the method steps in the above-mentioned method. These program codes can be read from or written into one or more computer program products, and the program codes can be compressed in a suitable form.
[0140] In one embodiment, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned federated learning training method for the large language model for code tasks is executed.
[0141] The computer device includes a memory, a processor, and one or more computer programs, where one or more computer programs can be stored in the memory and configured to be executed by one or more processors, and one or more application programs are configured to execute the above-mentioned federated learning training method for the large language model for code tasks.
[0142] The processor may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various parts within the entire computer device. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by invoking the data stored in the memory, it performs various functions of the computer device and processes data. Optionally, the processor may be implemented in at least one of the hardware forms of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may integrate one or a combination of several of the central processing unit (CPU), the reporting validator for buried point data (graphics processing unit, GPU), and the modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for the rendering and drawing of the displayed content; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and can be implemented separately through a communication chip.
[0143] The memory may include random access memory (RAM) and may also include read-only memory. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created during the use of the terminal device.
[0144] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A federated learning training method for large language models for code tasks, the federated learning training method is applied to a federated learning system, the federated learning system includes a server node and a number of organizational nodes, characterized in that, The above-mentioned federated learning training method is executed by the organization node and includes the following steps: Step S101: Preprocess the local data according to the code task to obtain structured training data; Step S102: Negotiate the training conditions, the base model, and the adapter model with the server node according to the code task; Step S103: Receive the base model sent by the server node; Step S104: Receive the initial adapter model sent by the server node, and fuse the initial adapter model and the base model to obtain an initial global model; Step S105: Train the initial global model according to the structured training data, and adjust the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model; Step S106: Send the fine-tuned adapter model and the quantity of the structured training data to the server node, so that the server node aggregates all the fine-tuned adapter models to update the initial adapter model; Step S107: Repeat steps S104 - S106 until the training conditions are met, so that the server node aggregates all the fine-tuned adapter models and the base model to obtain a trained global model.
2. The federated learning training method for large language models for code tasks according to claim 1, wherein Before executing step S103, it further includes: Correct the structured training data by inserting code language tags.
3. The federated learning training method for large language models for code tasks according to claim 2, characterized in that, The structured training data includes code content and annotation content.
4. The federated learning training method for large language models for code tasks according to claim 3, characterized in that, In the step of training the initial global model according to the structured training data and adjusting the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model, the specific expression of the optimization objective of the initial adapter model is: , Among them, is the fine-tuning adapter model, is the number of batches of structured training data for the organization nodes, is the size of the structured training data batch, is the loss function, is the base model, is the fine-tuning adapter model of the previous training round, is the regularization coefficient, is the th th structured training data in the th training batch.
5. Federated learning training method for large language models for code tasks. The federated learning training method is applied to a federated learning system, and the federated learning system includes a server node and several organizational nodes. It is characterized in that It includes the following steps: Step S201: The server node randomly selects several organization nodes as training organization nodes according to the code task; Step S202: Each organization node preprocesses the local data according to the code task to obtain structured training data; Step S203: The training organization nodes and the server node negotiate and determine the common base model, adapter model, and training conditions according to the code task; Step S204: The server node sends the base model and the initial adapter model to each training organization node; Step S205: Each training organization node fine-tunes the adapter model according to the local data to obtain a fine-tuned adapter model, and sends the fine-tuned adapter model and the quantity of the structured training data to the server node; Step S206: The server node aggregates according to the fine-tuned adapter models and the quantities of the structured training data fed back by each training organization node to obtain a trained adapter model; Step S207: Send the trained adapter model as the initial adapter back to each training organization node, and repeat steps S205 - S206 until the training conditions are met. The server node fuses the trained adapter model and the base model to obtain a trained global model.
6. The federated learning training method for large language models for code tasks according to claim 5, characterized in that, The server node aggregates according to the fine-tuned adapter models and the quantities of structured training data fed back by each training organization node to obtain a trained adapter model. The specific expression is as follows: , Among them, is the trained adapter model for the th training round, is the number of training organization nodes, is the number of structured training data of the th training organization node, is the total amount of structured training data in the th training round, is the fine-tuned adapter model feedback by the th training organization node in the th training round.
7. The federated learning training method for large language models for code tasks according to claim 5, characterized in that, The structured training data includes code content and annotation content.
8. The federated learning training method for large language models for code tasks according to claim 7, characterized in that, Before performing step S205, it further includes: Inserting code language tags into the structured training data for correction.
9. A federated learning training device for code tasks, the federated learning training device is applied to a federated learning system, the federated learning system includes a server node and a number of organization nodes, characterized in that, The federated learning training device is executed by the organization nodes and includes: A data preprocessing module for preprocessing local data according to the code task to obtain structured training data; A model selection module for negotiating training conditions, a base model, and an adapter model with the server node according to the code task; A base model receiving module for receiving the base model sent by the server node; A model fusion module for receiving the initial adapter model sent by the server node and fusing the initial adapter model and the base model to obtain an initial global model; A parameter adjustment module for training the initial global model according to the structured training data and adjusting the parameters of the initial adapter model in the initial global model to obtain a fine-tuned adapter model; A model update module for sending the fine-tuned adapter model and the quantity of structured training data to the server node so that the server node aggregates all the fine-tuned adapter models to update the initial adapter model; A global model output module for repeatedly executing the model fusion module, the parameter adjustment module, and the model update module until the training conditions are met, so that the server node aggregates all the fine-tuned adapter models and the base model to obtain a trained global model.
10. A federated learning training device for large language models for code tasks, the federated learning training device being applied to a federated learning system, the federated learning system including a server node and a plurality of organization nodes, characterized in that, It includes: An organization node extraction module for the server node to randomly extract several organization nodes as training organization nodes according to the code task; A data preprocessing module for each organization node to preprocess local data according to the code task to obtain structured training data; A model selection module for the training organization nodes and the server node to negotiate and determine the jointly trained base model, adapter model, and training conditions according to the code task; A model sending module for the server node to send the base model and the initial adapter model to each training organization node; A parameter adjustment module for each of the training organization nodes to fine-tune the adapter model according to local data to obtain a fine-tuned adapter model and send the fine-tuned adapter model and the quantity of structured training data to the server node; An adapter aggregation module for the server node to aggregate the fine-tuned adapter models and the quantities of structured training data fed back by each training organization node to obtain a trained adapter model; A global model output module for resending the trained adapter model as the initial adapter to each training organization node and repeatedly executing the parameter adjustment module and the adapter aggregation module until the training conditions are met, and the server node fuses the trained adapter model and the base model to obtain a trained global model.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the federated learning training method of the large language model for code tasks according to any one of claims 1-8 are implemented.
12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the federated learning training method of the large language model for code tasks according to any one of claims 1-8 is executed.
Citation Information
Patent Citations
Big model training and task processing method and device based on federated learning
CN117350411A
Industrial large model training method and system based on efficient fine tuning and federated learning
CN118982074A
Cited By
Large model parameter adjustment method based on cloud edge collaboration and related equipment
CN120851140A
Cloud-edge collaboration-based large model parameter adjustment method and related device
CN120851140B
Distributed intelligent question answering method based on novel federated knowledge learning model
CN122021664A
A distributed intelligent question-answering method based on a novel federated knowledge learning model
CN122021664B