Training Method, System, Device and Storage Medium of Large Language Model

By collecting task samples in large language models, performing data augmentation and internal and external loop training, and updating model parameters in adaptive optimizers, the difficulties of large language models in adaptive to new tasks and domains are solved, and better adaptability and training effects are achieved.

CN119416853BActive Publication Date: 2025-06-24SHANDONG INSPUR SCI RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510026862.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-24
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing large-scale language models have difficulties in dealing with new tasks and domain adaptability. Traditional transfer learning methods require a large amount of annotation data, and meta-learning methods such as MAML have limited applications on large-scale language models.

Method used

By collecting task samples based on task difficulty and task diversity, obtain support sets and query sets, and perform data enhancement processing on the support sets. The meta-model is trained internally by using the support set, and the query set is used for externally by using the meta-model to obtain the meta-gradient, and the global model is updated with the adaptive optimizer.

Benefits of technology

The adaptability of large language models to new tasks and fields is improved, the comprehensiveness of samples and training effects of the data set are improved, and momentum state sharing and adaptive parameter adjustment of internal and external cycle training is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416853B_ABST
    Figure CN119416853B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and specifically provides a training method, system, device and storage medium for a large language model, including: collecting task samples based on task difficulty and task diversity requirements; obtaining a support set and a query set of the task samples, and performing data augmentation processing on the support set; using the support set to perform in-loop training on the meta-model, and using the query set to perform out-loop training on the meta-model after in-loop training to obtain meta-gradients; aggregating the meta-gradients, and using an adaptive optimizer to update the global model based on the meta-gradient aggregation data. The present invention improves the adaptability of the large language model to task types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a training method, system, device and storage medium for a large language model. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have made significant progress in the field of natural language processing. They are able to generate coherent text, understand complex instructions, and show strong context adaptation capabilities. However, despite the powerful capabilities of these models, existing technologies still face many challenges in handling new tasks and domain adaptability.

[0003] The current state of the art is not optimistic. On the one hand, traditional transfer learning methods usually require a large amount of labeled data for fine-tuning, which is particularly difficult in new tasks or data-scarce situations, and it is difficult to quickly adapt to the needs of new tasks. On the other hand, although meta-learning methods such as Model-Agnostic Meta-Learning (MAML) have shown excellent performance in small sample learning, their application is still subject to many limitations when dealing with large language models, making it difficult to fully realize their potential.

[0004] Therefore, how to maintain the powerful performance of large language models while improving their ability to adapt to new tasks and fields is an important issue that needs to be urgently addressed in the field of natural language processing. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a large-scale language model training method, system, device and storage medium to solve the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a method for training a large language model, comprising:

[0007] Collect task samples based on task difficulty and task diversity requirements;

[0008] Obtain the support set and query set of the task sample, and perform data augmentation on the support set;

[0009] The meta-model is trained in an inner loop using the support set, and the meta-model trained in an inner loop is trained in an outer loop using the query set to obtain the meta-gradient.

[0010] Aggregate meta-gradients and use the adaptive optimizer to update the global model based on the meta-gradient summary data.

[0011] In an optional implementation, collecting task samples based on task difficulty and task diversity requirements includes:

[0012] Obtain the task difficulty, where the task difficulty is the loss value of the model on the task during historical training;

[0013] Confirm the task diversity, where the task diversity is the average distance between the task and the sampled task set;

[0014] Determine the sampling probability of the task based on the task difficulty and task diversity. The calculation method of the sampling probability is:

[0015]

[0016] where, is the trade-off difficulty, is the hyperparameter of diversity, is the task difficulty, is the task diversity.

[0017] In an optional implementation manner, obtaining the task difficulty includes:

[0018] The calculation formula of the task difficulty is:

[0019]

[0020] where, is the accuracy of the model on the validation set of the task ; is the highest accuracy among all tasks;

[0021] Perform normalization processing on the task difficulty to obtain the task difficulty quantization value of the task.

[0022] In an optional implementation manner, confirming the task diversity includes:

[0023] Represent the task as a feature vector using the TF-IDF method:

[0024]

[0025] where, represents the word frequency of the word in the task ; the inverse document frequency , represents the total number of candidate tasks;

[0026] Calculate the cosine similarity between the task and each task in the sampled task set :

[0027]

[0028] where, is the TF-IDF vector of the task , is the TF-IDF vector of the task ;

[0029] Calculate the task diversity based on the cosine similarity:

[0030]

[0031] The task diversity is the average distance between the task and the sampled task set.

[0032] In an alternative embodiment, perform data augmentation on the support set, including:

[0033] Replace the words in the sample data with their synonyms with a probability of 10%;

[0034] Randomly insert context-related words into the sentences of the sample data with a probability of 10%;

[0035] Randomly delete words in the sentences of the sample data with a probability of 10%;

[0036] Translate the text data into Chinese and then back into English with a probability of 20%.

[0037] In an alternative embodiment, perform inner-loop training on the meta-model using the support set, and perform outer-loop training on the meta-model trained in the inner loop using the query set to obtain the meta-gradient, including:

[0038] Initialize the meta-model parameters, let ;

[0039] Set the optimization function: ;

[0040] Based on the optimization function, use the data-augmented support set to perform inner-loop training on the meta-model:

[0041] For each task in , perform K steps of gradient descent to adapt to the task:

[0042] ;

[0043] Calculate the gradient: ;

[0044] Update the meta-model parameters: ;

[0045] Perform outer-loop training on the meta-model using the query set:

[0046]

[0047] Among them, is the outer loop learning rate, is the task batch size, is the adaptive optimizer, represents the loss function of is after the task-specific parameters adapted within steps of the inner loop, is the i-th sample of the query set, represents the gradient of represents the distribution of the task;

[0048]

[0049] Among them represents the -th sample and its label in the query set, represents the number of query set samples.

[0050] In an optional implementation, summarize the meta-gradient, and use the adaptive optimizer to update the global model based on the meta-gradient summary data, including:

[0051] Define the update rule for the gradient at time step t, the update rule includes:

[0052]

[0053] Among them is the shared first-order and second-order momentum, represents the current gradient, represents a small constant to prevent division by zero errors, is the first-order and second-order momentum estimates after bias correction, are different values of the momentum decay rate, is the learning rate, is the weight decay coefficient.

[0054] In a second aspect, the present invention provides a training system for a large language model, including:

[0055] A sampling module for collecting task samples based on task difficulty and task diversity requirements;

[0056] A processing module, configured to obtain a support set and a query set of task samples, and perform data augmentation processing on the support set;

[0057] A training module, configured to perform inner-loop training on a meta-model using the support set, and perform outer-loop training on the meta-model after inner-loop training using the query set to obtain meta-gradients;

[0058] An update module, configured to aggregate the meta-gradients, and update the global model based on the aggregated data of the meta-gradients using an adaptive optimizer.

[0059] In a third aspect, a device is provided, including:

[0060] A memory, configured to store a training program for a large language model;

[0061] A processor, configured to implement the steps of the training method of the large language model provided in the first aspect when executing the training program of the large language model.

[0062] In a fourth aspect, a computer-readable storage medium is provided, on which a training program for a large language model is stored. When the training program of the large language model is executed by a processor, the steps of the training method of the large language model provided in the first aspect are implemented.

[0063] The beneficial effects of the present invention are as follows. The training method, system, device and storage medium of the large language model provided by the present invention perform task sampling based on considering task difficulty and task diversity, improving the comprehensiveness of the samples in the data set, and enhancing the sample size of the data set required for inner-loop training through data augmentation processing, thus improving the training effect. Further, by combining inner and outer loop training and an adaptive optimizer, the momentum state sharing and adaptive parameter adjustment of inner and outer loop training are realized, improving the adaptability of the large language model to task types.

[0064] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0066] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0067] Figure 2 It is a schematic block diagram of the system according to an embodiment of the present invention.

[0068] Figure 3 A schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners

[0069] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0071] The following explains the key terms that appear in the present invention.

[0072] Large Language Models (LLMs) are natural language processing models based on deep learning that can understand and generate natural language text. They are trained on large-scale text data to learn the grammar, semantics, and various language features of the language, so that they can perform various language tasks such as text generation, translation, summarization, and question answering. LLMs use neural network-based models and often use natural language processing (NLP) techniques to process and calculate their outputs.

[0073] Traditional transfer learning methods usually require a large amount of labeled data for fine-tuning and are difficult to quickly adapt to new tasks.

[0074] Meta-learning methods such as Model-Agnostic Meta-Learning (MAML) perform well in few-shot learning, but their application in large language models is still limited.

[0075] Existing LLM meta-training methods such as MetaICL and MetaICT mainly adopt the multi-task learning paradigm and do not fully utilize the advantages of meta-learning.

[0076] Data augmentation techniques are effective in improving the robustness of models, but their combination with meta-learning is not deep enough.

[0077] Optimization algorithms for LLMs usually adopt Adam or AdamW, but how to effectively share the optimizer state in the meta-learning scenario is still a problem.

[0078] In view of the technical problems existing in large language models, the present invention provides a training method for large language models.

[0079] The training method for large language models provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the training system for large language models runs in the computer device.

[0080] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject can be a training system for large language models. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0081] As Figure 1 shown, the method includes:

[0082] S1. Collect task samples based on task difficulty and task diversity requirements.

[0083] Construct a comprehensive and challenging task sample library. This requires that when collecting task samples, we should not only consider the difficulty level of the tasks to ensure that tasks from simple to complex are covered, but also pay attention to the diversity of the tasks, including but not limited to text types, subject fields, language styles, etc. By deeply analyzing the requirements of the target application scenario, we can select or design task samples targeted to ensure that they can fully reflect the challenges and complexities in actual applications.

[0084] S2. Obtain the support set and query set of the task samples, and perform data augmentation processing on the support set.

[0085] After collecting sufficient task samples, we divide them into two independent data sets: the support set and the query set. The purpose of doing this is to be able to use the two data sets for inner-loop and outer-loop training respectively in the subsequent meta-learning process, thereby effectively improving the generalization ability of the model. For the support set, we perform data augmentation processing, such as by synonym replacement, sentence restructuring, adding noise, etc., to increase the diversity and richness of the data, so as to improve the robustness of the model to noise and changes.

[0086] S3. Use the support set to perform inner-loop training on the meta-model, and use the query set to perform outer-loop training on the meta-model after inner-loop training to obtain the meta-gradient.

[0087] Adopt a meta - learning strategy to train the model. First, use the support set to perform inner - loop training on the meta - model, that is, train the model on one or more tasks and collect its loss and gradient information. Then, we use this information to update the parameters of the meta - model so that it can better adapt to new tasks. Next, we use the query set to perform outer - loop training on the meta - model that has undergone inner - loop training, that is, evaluate the performance of the meta - model on new tasks and calculate the meta - gradient. The meta - gradient reflects the impact of the meta - model parameters on the performance of new tasks and is the key for us to update the global model subsequently.

[0088] S4. Aggregate the meta - gradients and use an adaptive optimizer to update the global model based on the aggregated meta - gradient data.

[0089] After obtaining the meta - gradients, we aggregate them and use an adaptive optimizer (such as Adam, RMSprop, etc.) to update the parameters of the global model. The adaptive optimizer can dynamically adjust the learning rate according to the direction and magnitude of the meta - gradients, thereby more effectively optimizing the performance of the global model. By continuously performing inner - loop and outer - loop training, and updating the global model based on the meta - gradients, we can gradually improve the generalization ability of the model and its adaptability to new tasks.

[0090] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.

[0091] S101. Obtain the task difficulty, where the task difficulty is the loss value of the model on the task during historical training.

[0092] In a specific example, the calculation formula for task difficulty is:

[0093]

[0094] where is the accuracy of the model on the validation set of task ; is the highest accuracy among all tasks;

[0095] Normalize the task difficulty to obtain the task difficulty quantization value of the task. Normalize the task difficulty to the interval [0, 1]. The higher the difficulty, the closer the value is to 1.

[0096] S102. Confirm the task diversity, where the task diversity is the average distance between the task and the sampled task set.

[0097] Use the TF - IDF method to represent the task as a feature vector:

[0098]

[0099] Among them, represents the word frequency of a word in the task ; the inverse document frequency , represents the total number of candidate tasks;

[0100] Calculate the cosine similarity between the task and each task in the sampled task set :

[0101]

[0102] Among them, is the TF-IDF vector of the task , is the TF-IDF vector of the task ;

[0103] Calculate the task diversity based on the cosine similarity:

[0104]

[0105] The task diversity is the average distance between the task and the sampled task set.

[0106] S103. Determine the sampling probability of the task based on the task difficulty and task diversity. The calculation method of the sampling probability is:

[0107]

[0108] Among them, is the trade-off for difficulty, is the hyperparameter for diversity, is the task difficulty, is the task diversity.

[0109] It can be understood that the dynamic task sampling strategy provides a diverse and challenging task batch for the two-layer optimization.

[0110] The performance feedback of the model after the two-layer optimization is used to update the task difficulty , affecting the subsequent sampling probability.

[0111] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.

[0112] S201. Sample to obtain a training task set , where each task Contain a set of support set samples , where represents the th sample in the support set and its label, represents the number of support set samples, and query set samples , where represents the th sample in the query set and its label, represents the number of query set samples.

[0113] S202. Perform data augmentation on the support set.

[0114] Original task sample: Sentence: The quick brown fox jumps over the lazy dog.

[0115] Examples of data augmentation processing:

[0116] Replace words in the sample data with their synonyms with a 10% probability:

[0117] Assume this process is triggered (10% probability), we can replace "quick" with its synonym "swift":

[0118] Augmented sentence: The swift brown fox jumps over the lazy dog.

[0119] Randomly insert context-related words into the sentence of the sample data with a 10% probability: Assume this process is triggered (10% probability), we can insert a context-related word, such as "agile" (related to "quick"), into the sentence:

[0120] Augmented sentence: The agile quick brown fox jumps over the lazy dog.

[0121] Randomly delete words in the sentence of the sample data with a 10% probability:

[0122] Assume this process is triggered (10% probability), randomly delete a word, such as "brown":

[0123] Augmented sentence: The quick fox jumps over the lazy dog.

[0124] Translate the text data into Chinese and then back into English with a 20% probability:

[0125] Assume this process is triggered (with a 20% probability):

[0126] Chinese translation: The quick brown fox jumps over the lazy dog.

[0127] Translate it back to English: The agile brown fox jumps over the lazy dog.

[0128] Due to the uncertainty of translation, "agile" here may not be the direct equivalent of the original English word, but it conveys a similar meaning. Additionally, since "quick" does not directly appear in the original English sentence, this enhancement introduces a new expression.

[0129] Comprehensive example: Assume all the above enhancement methods are applied to the original sentence simultaneously (of course, in actual applications, only one or several methods will be randomly selected each time), and each method is triggered (just for example, in reality, the triggering of each method is random):

[0130] Original sentence: The quick brown fox jumps over the lazy dog.

[0131] Apply synonym replacement: "quick" -> "swift";

[0132] Insert relevant words: "agile" (assume inserted before "swift");

[0133] Randomly delete words: "brown" (assume deleted);

[0134] Translate back to English (changes that may be introduced during the Chinese translation process);

[0135] Possible enhanced sentence: The agile swift fox jumps over the lazy dog.

[0136] In an embodiment of the present invention, based on step S3, the following will give a possible embodiment to non - restrictively elaborate on its specific implementation scheme.

[0137] Adopt a two - layer optimization framework of MAML, including inner - loop adaptation and outer - loop meta - update.

[0138] S301. Initialize the meta - model parameters, let ; Before the start of training, we first initialize the parameters of the meta - model. These parameters are the objects to be optimized in the subsequent training process, and they determine the initial state and performance of the meta - model.

[0139] S302. Set the optimization function: ;

[0140] The optimization function is used to guide the update of the meta-model parameters.

[0141] S303. Based on the optimization function, use the data-augmented support set to perform in-loop training on the meta-model:

[0142] For each task in, perform K steps of gradient descent to adapt to the task:

[0143] ;

[0144] Calculate the gradient: ;

[0145] Based on the current task-specific parameters and the support set data, calculate the gradient of the loss function with respect to the task-specific parameters.

[0146] Update the meta-model parameters: ;

[0147] Using the calculated gradient, we update the task-specific parameters according to the in-loop learning rate.

[0148] For each sampled task, we perform K steps of gradient descent to adapt to the task. At each step, we calculate the gradient of the loss function on the support set and use these gradients to update the task-specific parameters.

[0149] S304. Use the query set to perform out-of-loop training on the meta-model.

[0150] For each sample in the query set, use the task-specific parameters adapted through in-loop training to calculate the value of the loss function.

[0151] Based on the calculated loss and the out-of-loop learning rate, use an adaptive optimizer to update the meta-model parameters. This update process aims to enable the meta-model to better adapt to different types of tasks, thereby improving its generalization ability on unseen tasks.

[0152]

[0153] where is the out-of-loop learning rate, is the task batch size, is the adaptive optimizer, represents 's loss function, are the task-specific parameters after steps of in-loop adaptation, is the i-th sample of the query set, represents the inner loop learning rate, represents the gradient of, represents the distribution of tasks;

[0154]

[0155] wherein represents the th sample and its label in the query set, represents the number of query set samples.

[0156] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.

[0157] It can be seen from step S4 that both the inner loop training and the outer loop training adopt an adaptive optimizer. The same AdamW optimizer is used in the inner and outer loops, sharing its momentum state.

[0158] defines the update rule for the gradient at time step t , and the update rule includes:

[0159]

[0160] wherein is the shared first - order and second - order momentum, represents the current gradient, represents a small constant to prevent division - by - zero errors, is the first - order and second - order momentum estimates after bias correction, are different values of the momentum decay rate, is the learning rate, is the weight decay coefficient. Usually .

[0161] The adaptive optimizer is shared and used in the inner and outer loops to maintain consistent optimization behavior. The sharing of momentum helps to transfer useful optimization information between tasks.

[0162] The performance of the optimizer may affect the evaluation of task difficulty, thus indirectly affecting the sampling strategy.

[0163] It can be understood that the training steps of a complete large - language model mainly include:

[0164] Initialization: Randomly initialize the meta - model parameters , and initialize the state of the adaptive optimizer.

[0165] Outer loop (meta-training): while not reaching the convergence condition do;

[0166] Dynamic task sampling: Calculate the difficulty of each task and diversity ; According to Sample a task batch;

[0167] For each sampled task Execute the inner loop:

[0168] Partition the support set and query set ;

[0169] Apply data augmentation: ;

[0170] Initialization: ;

[0171] Inner loop adaptation:

[0172] for k = 1 to K do;

[0173] Calculate the gradient: ;

[0174] Update the parameters: ;

[0175] end for;

[0176] Calculate the meta-gradient: ;

[0177] Meta-update: Calculate the average meta-gradient: , Update the meta-model: ;

[0178] Update the task statistics: Update the task difficulty and task diversity

[0179] end while.

[0180] Based on the above embodiments, in order to further strengthen the understanding of the training method of the large language model of the present invention, the following is a specific implementation equation:

[0181] 1. Data preparation

[0182] First, prepare two sets of data:

[0183] a) Meta-training dataset: It contains 50,000 historical customer question-and-answer pairs from multiple existing product categories. Each piece of data includes the customer's question and the corresponding standard answer.

[0184] b) Meta-test dataset: It contains 1,000 customer question-and-answer pairs for newly launched smart home product categories. This is used to evaluate the model's adaptability to new domains.

[0185] Example of data structure: Each record contains two fields, "question" and "answer", corresponding to the customer's query and the customer service's reply respectively.

[0186] 2. Model selection

[0187] Select the pre-trained Qwen2 13B model as the base language model.

[0188] 3. MAMA-LLM implementation

[0189] 3.1 Two-layer optimization based on MAML

[0190] Implement the Model-Agnostic Meta-Learning (MAML) framework, including inner-loop adaptation and outer-loop meta-update.

[0191] Inner-loop adaptation: For each task, perform 5 steps of gradient descent with a learning rate of 0.01. Adapt using the task-specific support set data.

[0192] Outer-loop meta-update: Calculate the meta-gradient and update the global model based on the performance of the query sets of multiple tasks.

[0193] 3.2 Dynamic task sampling strategy

[0194] Implement a sampling strategy based on task difficulty and diversity:

[0195] Calculation of task difficulty: Use the loss of the model on this task as the difficulty metric.

[0196] Calculation of task diversity:

[0197] Convert the description of each task into a TF-IDF (Term Frequency-Inverse Document Frequency) vector.

[0198] Calculate the cosine similarity between tasks.

[0199] The diversity of a task is defined as 1 minus the average similarity to other tasks.

[0200] Calculation of sampling probability: Combine difficulty and diversity and use an exponential function to calculate the sampling probability of each task.

[0201] Dynamic update: After each training cycle, update the task difficulty according to the model's performance.

[0202] 3.3 Data Augmentation Module

[0203] Implement four NLP-specific data augmentation techniques:

[0204] Synonym replacement: Replace words with their synonyms with a 10% probability.

[0205] Random insertion: Insert context-related words at random positions in the sentence with a 10% probability.

[0206] Random deletion: Randomly delete words in the sentence with a 10% probability.

[0207] Back translation: Translate the text into Chinese and then back into English with a 20% probability.

[0208] For each support set sample, randomly apply one or more of the above augmentation methods to generate augmented data.

[0209] 3.4 Adaptive Optimizer

[0210] Implement the AdamW optimizer with momentum sharing:

[0211] Initialization: Create momentum and adaptive learning rate states for each parameter.

[0212] Update rule:

[0213] Calculate the exponential moving average of the gradient (momentum) and the exponential moving average of the squared gradient.

[0214] Use these statistics to adjust the learning rate of each parameter.

[0215] Apply weight decay to prevent overfitting.

[0216] Momentum sharing: Share the momentum state between the inner loop and the outer loop to maintain the continuity of optimization.

[0217] The learning rate is set to 1e-5 and the weight decay coefficient is 0.01.

[0218] 4. Training Process

[0219] The training process is carried out for 100 epochs, and the steps for each epoch are as follows:

[0220] Select 5 training tasks using the dynamic task sampling strategy.

[0221] For each task: a. Sample 10 samples as the support set and apply data augmentation. b. Perform inner loop adaptation (5-step gradient descent). c. Calculate the loss and meta-gradient on the query set.

[0222] Aggregate the meta-gradient and update the global model using an adaptive optimizer.

[0223] Update the task difficulty and sampling probability.

[0224] During training, record and monitor the meta-training loss to evaluate the learning progress.

[0225] 5. Application to a new task (smart home product customer service)

[0226] Prepare data: Select 20 samples from the new smart home product customer service data as the support set.

[0227] Quick adaptation: a. Load the trained MAMA-LLM model. b. Perform 10 steps of inner-loop adaptation using the support set data.

[0228] Evaluate: Calculate the loss on the remaining query set to evaluate the model performance.

[0229] Practical application: Use the adapted model to answer users' new questions about smart home products.

[0230] 6. Effect demonstration

[0231] The following is an example of a practical application:

[0232] User question: "How do I connect my smart bulb to my phone App?"

[0233] Model answer: "The steps to connect the smart bulb to the phone App are as follows: 1. Ensure that your phone has the latest version of our smart home App installed. 2. Open the App and click 'Add Device'. 3. Select the 'Lighting Devices' category and then choose your smart bulb model. 4. Plug the bulb into the power supply and ensure it is in pairing mode (usually a fast blinking state). 5. Follow the instructions in the App to connect to the temporary Wi-Fi network created by the bulb. 6. Enter the password of your home Wi-Fi, and the App will automatically complete the pairing process. 7. Name the bulb and assign it to the corresponding room. Now, you can control the smart bulb through the App. If you encounter any problems, don't hesitate to ask for further help."

[0234] This answer demonstrates that the model can generate detailed and targeted responses, covering specific installation steps for smart home devices, indicating that the large language model trained using the training method provided by the present invention has successfully adapted to the new product field.

[0235] In some embodiments, the training system of the large language model may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the training system of the large language model can be stored in the memory of the computer device and executed by at least one processor to perform the functions of training the large language model (see details in Figure 1 description).

[0236] In this embodiment, according to the functions it performs, the training system of the large language model can be divided into multiple functional modules, as shown in Figure 2 . The functional modules of system 200 may include: a sampling module 210, a processing module 220, a training module 230, and an updating module 240. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0237] The sampling module is used to collect task samples based on task difficulty and task diversity requirements;

[0238] The processing module is used to obtain the support set and query set of the task samples and perform data augmentation processing on the support set;

[0239] The training module is used to perform inner-loop training on the meta-model using the support set and perform outer-loop training on the meta-model after inner-loop training using the query set to obtain meta-gradients;

[0240] The updating module is used to aggregate the meta-gradients and update the global model using an adaptive optimizer based on the meta-gradient aggregation data.

[0241] Optionally, as an embodiment of the present invention, collecting task samples based on task difficulty and task diversity requirements includes:

[0242] Obtaining the task difficulty, where the task difficulty is the loss value of the model on the task during historical training;

[0243] Confirming the task diversity, where the task diversity is the average distance between the task and the sampled task set;

[0244] Determining the sampling probability of the task based on the task difficulty and task diversity, and the calculation method of the sampling probability is:

[0245]

[0246] where, is the trade-off difficulty, is the hyperparameter of diversity, is the task difficulty, is the task diversity.

[0247] Optionally, as an embodiment of the present invention, obtaining the task difficulty includes:

[0248] The calculation formula of the task difficulty is:

[0249]

[0250] where, is the accuracy rate of the model on the validation set of the task ; is the highest accuracy rate among all tasks;

[0251] Perform normalization processing on the task difficulty to obtain the task difficulty quantization value of the task.

[0252] Optionally, as an embodiment of the present invention, confirming the task diversity includes:

[0253] Use the TF-IDF method to represent the task as a feature vector:

[0254]

[0255] where, represents the word frequency of the word in the task ; the inverse document frequency , represents the total number of candidate tasks;

[0256] Calculate the cosine similarity between the task and each task in the sampled task set :

[0257]

[0258] where, is the TF-IDF vector of the task , is the TF-IDF vector of the task ;

[0259] Calculate the task diversity based on the cosine similarity:

[0260]

[0261] The task diversity is the average distance between the task and the sampled task set.

[0262] Optionally, as an embodiment of the present invention, performing data augmentation processing on the support set includes:

[0263] Replace the words in the sample data with their synonyms with a probability of 10%.

[0264] Randomly insert context-related words into the sentences of the sample data with a probability of 10%.

[0265] Randomly delete words in the sentences of the sample data with a probability of 10%.

[0266] Translate the text data into Chinese and then back into English with a probability of 20%.

[0267] Optionally, as an embodiment of the present invention, use the support set to perform inner-loop training on the meta-model, and use the query set to perform outer-loop training on the meta-model that has undergone inner-loop training to obtain the meta-gradient, including:

[0268] Initialize the meta-model parameters, let ;

[0269] Set the optimization function: ;

[0270] Based on the optimization function, use the data-augmented support set to perform inner-loop training on the meta-model:

[0271] For each task in , perform K steps of gradient descent to adapt to the task:

[0272] ;

[0273] Calculate the gradient: ;

[0274] Update the meta-model parameters: ;

[0275] Use the query set to perform outer-loop training on the meta-model:

[0276]

[0277] where, is the outer-loop learning rate, is the task batch size, is the adaptive optimizer, represents 's loss function, is the task-specific parameter after steps of inner-loop adaptation, is the i-th sample of the query set, represents the inner-loop learning rate, represents 's gradient, represents the distribution of the task;

[0278]

[0279] wherein represents the th sample and its label in the query set, represents the number of samples in the query set.

[0280] Optionally, as an embodiment of the present invention, aggregating the meta-gradient and updating the global model based on the aggregated data of the meta-gradient using an adaptive optimizer includes:

[0281] Define an update rule for the gradient at time step t , the update rule includes:

[0282]

[0283] where is the shared first and second order momenta, represents the current gradient, represents a small constant to prevent division by zero error, is the first and second order momentum estimates after bias correction, are different values of the momentum decay rate, is the learning rate, is the weight decay coefficient.

[0284] Figure 3 The training method for the large language model provided by the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.

[0285] Among them, the device 300 may include: a processor 310, a memory 320, and a communication unit 330. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0286] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 can execute some or all of the steps in the above method embodiments.

[0287] The processor 310 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 320, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.

[0288] The communication unit 330 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.

[0289] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0290] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., various media that can store program codes, including several instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0291] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0292] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or modules can be in electrical, mechanical, or other forms.

[0293] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0294] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0295] Although the present invention has been described in detail by reference to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and all such modifications or substitutions should be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of changes or substitutions, and all of them should be covered within the protection scope of the present invention.

Claims

1. A method for training a large language model, characterized in that: include: Collect task samples based on task difficulty and task diversity requirements; Obtain the support set and query set of the task sample, and perform data augmentation on the support set; The meta-model is trained in an inner loop using the support set, and the meta-model trained in an inner loop is trained in an outer loop using the query set to obtain the meta-gradient. Aggregate meta-gradients and use the adaptive optimizer to update the global model based on the meta-gradient summary data; Identify task diversity, including: The TF-IDF method is used to represent the task as a feature vector: Among them, TF(w,T i ) represents the word w in task T i term frequency in ; inverse document frequency The total number of tasks representing candidates; Computational task T i With the sampled task set The cosine similarity of each task in: in, It is task T i The TF-IDF vector of It is task T j TF-IDF vector of Calculate task diversity based on the cosine similarity: The task diversity is task T i The average distance to the sampled task set; Aggregate meta-gradients and use the adaptive optimizer to update the global model based on the meta-gradient summary data, including: AdaptiveOptimizer() defines the update rule, for the gradient g at time step t t , the update rules include: m t =β1m t-1 +(1-β1)g t Where m t ,v t is the shared first and second order momentum, g t represents the current gradient, ∈ represents a small constant to prevent division by zero errors, are the bias-corrected first-order and second-order momentum estimates, β1, β2 are different values ​​of the momentum decay rate, η is the learning rate, and λ is the weight decay coefficient; Get the difficulty of the task, including: The calculation formula for task difficulty is: in, is the model f θ In Task T i The validation set The accuracy of It is the highest accuracy among all tasks; The task difficulty is normalized to obtain the quantitative value of the task difficulty.

2. The method according to claim 1, characterized in that Task samples are collected based on task difficulty and task diversity requirements, including: Obtaining a task difficulty, where the task difficulty is the loss value of the model on the task during historical training; Determine task diversity, where the task diversity is the average distance between the task and a sampled task set; The sampling probability of the task is determined based on the task difficulty and task diversity. The sampling probability is calculated as follows: Among them, λ1 is the trade-off difficulty, λ2 is the hyperparameter of diversity, and d(T i ) is the task difficulty, For task diversity.

3. The method according to claim 1, characterized in that Perform data augmentation on the support set, including: Replace the words in the sample data with their synonyms with a probability of 10%; Randomly insert context-related words into sentences in the sample data with a probability of 10%; Randomly delete words from sentences in sample data with a probability of 10%; The text data is translated into Chinese and then translated back into English with a probability of 20%.

4. The method according to claim 1, characterized in that The meta-model is trained in an inner loop using the support set, and the meta-model trained in an inner loop is trained in an outer loop using the query set to obtain the meta-gradient, including: Initialize the metamodel parameters, let Set the optimization function: Based on the optimization function, the support set of data enhancement is used Perform inner loop training on the meta-model: for For each task in , perform K steps of gradient descent to adapt to the task: Compute the gradient: Update metamodel parameters: Use the query set to train the metamodel in an outer loop: Among them, β is the outer loop learning rate, B is the task batch size, AdaptiveOptimizer() is the adaptive optimizer, Represents T i The loss function is are the task-specific parameters after K steps of inner loop adaptation, is the i-th sample in the query set, α represents the inner loop learning rate, represent The gradient of Represents the distribution of tasks; in represents the kth sample and its label in the query set, n q Represents the number of query set samples.

5. A large language model training system, characterized in that: include: Sampling module, used to collect task samples based on task difficulty and task diversity requirements; The processing module is used to obtain the support set and query set of the task sample and perform data enhancement processing on the support set; A training module, used for performing inner loop training on the meta-model using the support set, and performing outer loop training on the meta-model trained by the inner loop using the query set to obtain a meta-gradient; An update module is used to summarize meta-gradients and use an adaptive optimizer to update the global model based on the meta-gradient summary data; Identify task diversity, including: The TF-IDF method is used to represent the task as a feature vector: Among them, TF(w,T i ) represents the word w in task T i term frequency in ; inverse document frequency The total number of tasks representing candidates; Computational task T i With the sampled task set The cosine similarity of each task in: in, It is task T i The TF-IDF vector of It is task T j TF-IDF vector of Calculate task diversity based on the cosine similarity: The task diversity is task T i The average distance to the sampled task set; Aggregate meta-gradients and use the adaptive optimizer to update the global model based on the meta-gradient summary data, including: AdaptiveOptimizer() defines the update rule, for the gradient g at time step t t , the update rules include: m t =β1m t-1 +(1-β1)g t Where m t ,v t is the shared first and second order momentum, g t represents the current gradient, ∈ represents a small constant to prevent division by zero errors, are the bias-corrected first-order and second-order momentum estimates, β1, β2 are different values ​​of the momentum decay rate, η is the learning rate, and λ is the weight decay coefficient; Get the difficulty of the task, including: The calculation formula for task difficulty is: in, is the model f θ In Task T i The validation set The accuracy of It is the highest accuracy among all tasks; The task difficulty is normalized to obtain the quantitative value of the task difficulty.

6. A large language model training device, characterized in that: include: Memory, used to store training programs for large language models; A processor, used to implement the steps of the large language model training method as described in any one of claims 1-4 when executing the training program of the large language model.

7. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a training program for a large language model, and when the training program for the large language model is executed by a processor, the steps of the training method for a large language model as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Risk management method and system for cross-border e-commerce transaction behavior

    CN118469715A