Natural language processing, model training method, device and storage medium
Patent Information
- Application Number
- CN202310652747.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-06-02
AI Technical Summary
在对第一语言模型进行更新迭代时,若直接对整个模型进行重训练,则将导致模型发生灾难性遗忘,从而降低模型的任务处理性能
[0014] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the method provided in this application.
Smart Images

Figure CN116775807B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a natural language processing, model training method, device and storage medium. Background Technology
[0002] With the development of artificial intelligence, deep learning is widely used in various application scenarios, such as image processing and natural language processing. In these scenarios, the task processing logic and data distribution of the first language model change continuously over time, requiring constant updates and iterations after deployment and operation. However, directly retraining the entire model during these updates would lead to catastrophic forgetting, thus reducing its task processing performance. Therefore, a new solution is needed. Summary of the Invention
[0003] This application provides a natural language processing method, a device, and a storage medium for learning new features and reducing the probability of catastrophic forgetting during incremental training of a first language model.
[0004] This application provides a natural language processing method, comprising: acquiring corpus data to be processed for a target task; inputting the corpus data to be processed into a first language model to obtain a processing result; wherein the first language model is obtained by adding a task adaptation network to the feature transformation network of a second language model; the second language model is trained using historical corpus samples; the first language model includes parameters of the task adaptation network and main model parameters inherited from the second language model; the parameters of the task adaptation network are obtained by incrementally learning the target task using incremental corpus samples.
[0005] Optionally, inputting the corpus data to be processed into a first language model to obtain a processing result includes: in the first language model, determining a target task adaptation network corresponding to the target task based on the type of the target task; and processing the corpus data to be processed based on the main model parameters and the model parameters of the target task adaptation network to obtain the processing result.
[0006] This application embodiment also provides a model training method, including: acquiring incremental corpus samples required for incremental learning of a target task; inputting the incremental corpus samples into a first language model to obtain the processing result of the incremental corpus samples; the first language model is obtained by adding a task adaptation network to the feature transformation network of a second language model; the second language model is trained using historical corpus samples; based on the processing result, determining the training loss of the first language model for the target task; and optimizing the parameters of the task adaptation network with the goal of satisfying a set convergence condition for the training loss.
[0007] Optionally, optimizing the parameters of the task adaptation network based on the training loss includes: determining the task type of the target task; determining a target task adaptation network corresponding to the task type from at least one task adaptation network of the first language model; and optimizing the parameters of the target task adaptation network based on the training loss.
[0008] Optionally, before inputting the incremental corpus samples into the first language model, the method further includes: in a pre-training sub-stage, obtaining corpus samples related to the target task from the historical corpus samples of the second language model as initialization corpus samples; inputting the initialization corpus samples into the first language model to obtain the initialization processing result of the initialization corpus samples; determining the initialization training loss of the first language model based on the initialization processing result; determining a first type of parameter in the task adaptation network; the first type of parameter being used to learn knowledge of the target task; and updating the first type of parameter based on the initialization training loss.
[0009] Optionally, acquiring incremental corpus samples required for incremental learning of the target task includes: in the incremental training sub-stage, obtaining real-time application data associated with the target task from the application data during the operation of the first language model, as the incremental corpus samples; optimizing the parameters of the task adaptation network according to the training loss, including: determining a second type of parameters in the task adaptation network; the second type of parameters being used to learn multi-task fusion knowledge; and updating the second type of parameters according to the training loss.
[0010] Optionally, based on the processing result, determining the training loss of the first language model for the target task includes: determining the task learning loss of the target task based on the difference in feature representation of positive and negative sample pairs in the incremental corpus samples by the first language model; and determining the distillation loss of the first language model based on the difference in feature representation of the incremental corpus samples by the first language model and the second language model; and determining the training loss of the target task based on the task learning loss and the distillation loss.
[0011] Optionally, determining the task learning loss for the target task based on the feature representation differences between positive and negative sample pairs in the incremental corpus samples by the first language model includes: determining the positive and negative samples of any sample from the incremental corpus samples; the sample and its positive sample forming a positive sample pair, and the sample and its negative sample forming a negative sample pair; determining a first similarity of the positive sample pair based on the features extracted from the sample and the positive sample by the first language model respectively; and determining a second similarity of the negative sample pair based on the features extracted from the sample and the negative sample by the first language model respectively; and constructing a contrastive learning loss based on the first similarity and the second similarity as the task learning loss for the target task.
[0012] Optionally, determining the distillation loss of the first language model based on the differences in feature representation of the incremental corpus samples by the first language model and the second language model includes: determining a third similarity of the positive sample pair based on the features extracted by the second language model from the sample and the positive sample respectively; and determining a fourth similarity of the negative sample pair based on the features extracted by the second language model from the sample and the negative sample respectively; calculating the entropy of the similarity of the positive sample based on the first similarity and the third similarity, and calculating the entropy of the similarity of the negative sample based on the second similarity and the fourth similarity; and determining the distillation loss based on the cumulative value of the entropy of the similarity of the positive sample and the entropy of the similarity of the negative sample.
[0013] This application also provides a server, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to perform the steps in the method provided in this application.
[0014] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the method provided in this application.
[0015] In this embodiment, the first language model used for incremental learning of the target task is obtained by adding a task adaptation network to the feature transformation network of the second language model. Incremental corpus samples required for incremental learning are input into the first language model to obtain the processing results of the incremental corpus samples. Based on the processing results and the training ground truth of the incremental corpus samples, the training loss of the first language model is determined, and the parameters of the task adaptation network can be optimized according to the training loss. In this incremental learning process, without updating the main model parameters of the base model, the knowledge of new data is learned by updating the parameters of the task adaptation network, thereby enabling the first language model to adaptively learn new features and reducing the probability of catastrophic forgetting. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0017] Figure 1 A schematic flowchart of a model training method provided for an exemplary embodiment of this application;
[0018] Figure 2 A schematic diagram illustrating incremental learning of a second language model to obtain a first language model, provided as an exemplary embodiment of this application;
[0019] Figure 3a A partial structural diagram of a second language model provided for an exemplary embodiment of this application;
[0020] Figure 3b A schematic diagram of the structure of a task adaptation network provided in another exemplary embodiment of this application;
[0021] Figure 4 A flowchart illustrating a natural language processing method provided in an exemplary embodiment of this application;
[0022] Figure 5 This is a schematic diagram of the structure of a server provided for an exemplary embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” used in the embodiments of this invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one.
[0025] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0026] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.
[0027] In view of the technical problem that in the prior art, when updating and iterating the first language model, directly retraining the entire model will lead to catastrophic forgetting of the model, thereby reducing the model's task processing performance, a solution is provided in some embodiments of this application. The technical solutions provided by the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0028] Figure 1 This is a flowchart illustrating a natural language processing and model training method provided in an exemplary embodiment of this application. The method may include, for example: Figure 1 The steps shown are as follows:
[0029] Step 101: Obtain the incremental corpus samples required for incremental learning of the target task.
[0030] Step 102: Input the incremental corpus sample into the first language model to obtain the processing result of the incremental corpus sample; the first language model is obtained by adding a task adaptation network to the feature transformation network of the second language model; the second language model is trained by historical corpus samples.
[0031] Step 103: Determine the training loss of the first language model based on the processing result.
[0032] Step 104: Optimize the parameters of the task adaptation network with the goal of satisfying the set convergence conditions for the training loss.
[0033] In this embodiment, for ease of description, the model training phase is broadly divided into a basic training phase and a self-learning phase. In the basic training phase, the collected corpus samples are used as training data to perform offline training on basic processing, enabling the second language model to learn knowledge from the existing training data. The trained second language model can then be deployed as a basic model in application scenarios. For example, in the field of natural language processing, the basic model can be deployed in scenarios such as chatbot customer service Q&A, user sentiment analysis, and machine translation. In the self-learning phase, the robot can acquire new training data to incrementally train the basic model, thereby improving its performance in handling tasks in different scenarios.
[0034] During the self-learning phase, to reduce the probability of forgetting previously learned knowledge while learning new knowledge, the structure of the base model was improved by adding a "pluggable" knowledge integration network module (hereinafter referred to as the task adapter network). In this embodiment, the improved base model after adding the knowledge integration network module is denoted as the first language model. That is, the first language model is obtained by adding the task adapter network to the second language model. The task adapter network is a knowledge integration network module containing a certain number of parameters. The task adapter network can be added to the second language model as needed and can be removed from the second language model when necessary, achieving a "pluggable" knowledge integration. It should be noted that the use of "first" and "second" to define the model here is only for ease of description and distinction and does not impose any restrictions on the training order or number of models.
[0035] like Figure 2 As shown, a second language model can be trained offline based on historical corpus samples. After model evaluation, the second language model can be deployed online to perform corresponding tasks based on the learned knowledge. During the online operation of the second language model, application data generated during the process can be fed back through data tracking. Based on the collected feedback data, new samples, i.e., incremental corpus samples, can be constructed. Sample construction includes labeling the samples and identifying and annotating positive and negative samples, which will not be elaborated here. After obtaining the incremental corpus samples, a task adaptation network can be added to the second language model to obtain a first language model. Incremental learning can then be performed based on the task adaptation network to obtain a converged first language model. After the first language model passes model evaluation, the online deployed second language model can be updated.
[0036] Optionally, a task adaptation network may be added to one or more feature transformation networks of the second language model, or a task adaptation network may be added to each computing layer of the second language model, so as to improve and obtain the first language model, which is not limited in this embodiment.
[0037] Wherein, the second language model can be created based on any model framework comprising a feature transformation network. Wherein, optional model architectures may include, but are not limited to: at least one of a BERT (Bidirectional Encoder Representation from Transformers, bidirectional encoder representation from feature transformation networks) model framework, T5 (Transfer Text-to-Text Transformer, feature transformation network for converting text to text) and GPT (Generative Pre-Trained Transformer, generative pre-trained feature transformation network model) framework.
[0038] Taking a model framework comprising a deep self-attention transformation network (transformer) as an example, such as Figure 3a shown in, in some embodiments, a task adaptation network may be added in the deep self-attention transformation network of the second language model, and the task adaptation network may be inserted after the feed-forward layer of the deep self-attention transformation network. Figure 3b An exemplary structure of the task adaptation network is schematically illustrated, such as Figure 3b shown in, each task adaptation network may comprise two feed-forward sub-layers, the first feed-forward sub-layer (namely Figure 3b the upstream feed-forward sub-layer illustrated) is configured to take the output of the deep self-attention network as an input, project the dimension d of the original input to a dimension m, and limit the number of parameters of the task adaptation network by controlling the size of m, wherein m<<d. In the output stage of the task adaptation network, the dimension m is re-projected to the dimension d through the second feed-forward sub-layer (namely Figure 3b the downstream feed-forward sub-layer illustrated) to input the dimension and output a calculation result. Based on this implementation, adding a task adaptation network can produce an easily scalable downstream model. When a new downstream task occurs, adding a task adaptation network can avoid the problems of full-model fine-tuning and catastrophic forgetting. Meanwhile, incremental learning is performed based on the task adaptation network, which does not require fine-tuning the main model parameters of the base model, and can store knowledge about the target task by introducing a small number of parameters for the target task, reducing the computing power requirement for model fine-tuning.
[0039] The following example uses any task to illustrate the incremental training method for the first language model during the self-learning phase. For ease of description and distinction, this task will be described as the target task. The target task can be any task learned by the first language model during the basic training phase, or it can be a new task learned through self-learning; this embodiment does not impose any limitations.
[0040] The incremental corpus samples required for incremental training on the target task are related to the task type. For example, when the target task is a question-answering task, the incremental corpus samples may include question-answer data pairs. For example, when the target task is a translation task, the incremental corpus samples include training data pairs consisting of the source text and the translation. And for example, when the target task is a sentiment analysis task, the incremental corpus samples include training data pairs consisting of the text and the sentiment analysis results, and so on.
[0041] After determining the incremental corpus samples, they can be input into a first language model. The first language model can extract features from these samples based on its parameters and output processing results according to the extracted features. The processing results differ depending on the task type. For example, if the target task is translation, the first language model's output may include translation results. Similarly, if the target task is sentiment analysis, the first language model's output may include sentiment analysis results. And if the target task is question-answering, the first language model's output may include responses to questions.
[0042] After obtaining the processing results of the first language model for the training data, the training loss of the first language model can be determined based on the training ground truth (GT) of the training data and the processing results, and the first language model can be optimized based on the training loss. In this embodiment and subsequent embodiments, for ease of description, the parameters of the first language model are divided into two parts: the main model parameters and the task adaptation network parameters. The main model parameters refer to the model parameters in the first language model other than the task adaptation network parameters, i.e., the inherent model parameters of the base model. During the incremental training phase, the main model parameters of the first language model can be kept unchanged, and the parameters of the task adaptation network are optimized with the goal of converging the training loss of the first language model. Incremental training can be executed iteratively until the training loss of the first language model meets the preset convergence condition, at which point the trained first language model is output. The convergence condition may include: the training loss is less than a set threshold, or the training loss fluctuates within a set range; this embodiment does not impose any limitations on this.
[0043] In this implementation, the first language model used for incremental learning of the target task is obtained by adding a task adaptation network to the feature transformation network of the second language model. Incremental corpus samples required for incremental learning are input into the first language model to obtain the processing results of the incremental corpus samples. Based on the processing results and the training ground truth of the incremental corpus samples, the training loss of the first language model is determined, and the parameters of the task adaptation network can be optimized according to the training loss. In this incremental learning process, without updating the main model parameters of the base model, the knowledge of new data is learned by updating the parameters of the task adaptation network, thereby enabling the first language model to adaptively learn new features and reducing the probability of catastrophic forgetting.
[0044] In some optional embodiments, the base model may include a multi-task model. This multi-task model is used to perform prediction or classification of multiple tasks. Typically, the multiple tasks may share lower network layers and each has its own branch network. Based on the model training method provided in this application, incremental self-learning of multiple tasks can be achieved on top of the base model. The following will provide an exemplary description.
[0045] Optionally, multiple task adaptation networks can be added to the base model to obtain the first language model, with each set of task adaptation networks corresponding to a specific task type. For example, when using a multi-task model to handle question-answering, translation, and sentiment analysis tasks, a set of task adaptation networks corresponding to question-answering, a set of task adaptation networks corresponding to translation, and a set of adaptation networks corresponding to sentiment analysis can be added to the base model. When performing incremental self-learning on a particular task type, the parameters of the task adaptation network corresponding to that task type can be optimized.
[0046] Continuing with the example of the target task, when optimizing the parameters of the task adaptation network based on the training loss of the first language model, the task type of the target task can be determined. Then, from at least one task adaptation network in the first language model, a target task adaptation network corresponding to the task type of the target task can be selected. This target adaptation network is used to learn the knowledge of the target task. Therefore, based on the training loss, the parameters of this target task adaptation network can be optimized to improve the performance of the first language model in the target task. For example, when the target task is a question-answering task, after obtaining the training loss of the first language model for question-answering tasks, a set of task adaptation networks corresponding to the question-answering task can be determined in the first language model, and the parameters of this set of task adaptation networks corresponding to the question-answering task can be optimized based on the training loss.
[0047] In this implementation, by setting up task adaptation networks for different types of tasks, differentiated self-learning for different types of tasks can be achieved, thereby improving the processing performance of the first language model on different types of tasks.
[0048] In some optional embodiments, the training of the first language model may include a pre-training sub-stage and an incremental training sub-stage (i.e., a self-learning stage). The pre-training sub-stage is offline training, while the incremental training sub-stage is online training after the first language model is deployed. The pre-training sub-stage is primarily used to pre-train the task adaptation network for any given task based on corpus samples, thereby initializing the model parameters of the task adaptation network for that task. In the self-learning stage, the first language model can learn specific knowledge for each task based on incrementally collected samples and encapsulate this learned knowledge in the parameters of the task adaptation network for each task. After the first language model is deployed, the incremental training sub-stage can be repeated at a set period, or it can be started after a certain number of incremental samples have been collected; this embodiment does not impose any restrictions.
[0049] The incremental training sub-stage is mainly used to incrementally train the task adaptation network for different tasks based on the application data generated after the basic model is launched and running online, so as to further optimize the parameters of the task adaptation network corresponding to the task.
[0050] In some optional embodiments, if the first language model is a multi-task model, in the incremental training sub-stage, the learned multi-task knowledge can be further fused and transferred.
[0051] Optionally, continuing with the target task as an example, the task adaptation network for the target task includes a first type of parameters and a second type of parameters. The first type of parameters is used to learn knowledge about the target task during the pre-training sub-stage, while the second type of parameters is used to learn knowledge about other tasks during the incremental training sub-stage, i.e., learning multi-task fusion knowledge. During the pre-training sub-stage for learning the target task, after obtaining the training loss of the first language model for the target task, the first type of parameters can be optimized to encapsulate the knowledge of the target task within the first type of parameters.
[0052] In the incremental training sub-stage, the first type of parameters can be kept unchanged, while the second type of parameters are updated. It's worth noting that in the pre-training sub-stage, the first type of parameters in the task adaptation networks for each task have already learned the knowledge corresponding to their respective tasks. In the incremental training sub-stage, the second type of parameters in the task adaptation networks can be used to combine and learn the knowledge already learned by the first type of parameters in the various task adaptation networks, thereby achieving knowledge fusion and transfer across multiple tasks.
[0053] Let's continue with the example of the target task.
[0054] In the pre-training sub-stage, corpus samples related to the target task can be obtained from the historical corpus samples of the second language model as initialization corpus samples. After inputting the initialization corpus samples into the first language model, the initialization processing result of the initialization corpus samples can be obtained. Based on the initialization processing result, the initialization training loss of the first language model can be determined. Based on this initialization training loss, the first type of parameters in the task adaptation network can be updated. During this update process, the parameters of the base model and the second type of parameters remain unchanged, while the first type of parameters are updated. Thus, the first type of parameters can be initialized based on the historical corpus samples to improve the performance of the target task on the historical corpus samples.
[0055] After initializing the first type of parameters in the first language model based on the pre-training sub-stage, the first language model can be put into operation online and enter the incremental training sub-stage (i.e., the self-learning stage).
[0056] In the incremental training sub-stage, the incremental corpus samples required for incremental learning of the target task may include real-time application data corresponding to the target task. This real-time application data can be obtained from the data generated after the first language model is launched and running online. After inputting the incremental corpus samples into the first language model, the processing result of the incremental corpus samples is obtained, and the training loss is determined based on the processing result. After obtaining the training loss based on the incremental corpus samples, the parameters of the task adaptation network can be optimized based on the training loss. Optionally, during the optimization of the parameters of the task adaptation network based on the training loss, the second type of parameters in the task adaptation network are determined, and the second type of parameters are updated based on the training loss. In this update process, the parameters of the base model and the first type of parameters remain unchanged, while the second type of parameters are updated.
[0057] Based on this implementation method, initializing the parameters of the task adaptation network added to the first language model using historical data can improve the performance of the first language model in different tasks to a certain extent. Further optimization of the task adaptation network parameters using real-time application data generated after the model goes live allows for real-time adjustments to the task adaptation network based on the model's application requirements, enabling the first language model to adapt to changes in the scenario and improving its robustness in different application scenarios. Simultaneously, updating different parameters at different training stages effectively solves the problems of inter-task interference and training instability, while also effectively mitigating the catastrophic forgetting problem.
[0058] It is worth noting that increasing the number of task-adaptive networks also leads to an increase in the number of parameters, which reduces the inference speed of the model and is not conducive to running on some resource-limited terminals. In some embodiments, after the language model is trained, some task-adaptive networks can be dynamically removed based on the "pluggable" characteristic of the task-adaptive networks to reduce the number of parameters and computational cost of the language model, thereby improving the prediction efficiency of the model in the inference stage and making the model more lightweight.
[0059] In some alternative embodiments, the training loss of the first language model in each round of training may include at least two parts: task learning loss and distillation loss.
[0060] Among them, the task learning loss is used to express the loss of the first language model (new model) based on the incremental corpus sample processing task, so as to continuously learn from the incremental new data.
[0061] Distillation loss is used to express the performance difference between the first language model (new model) and the base model (old model), to measure the gap between the first language model and the base model, thereby reducing the forgetting of representations learned in the base training phase.
[0062] In some optional embodiments, a contrastive learning approach can be used to determine the task learning loss and distillation loss. In this contrastive learning approach, the task learning loss can be determined based on the performance difference of the first language model on positive and negative sample data. The smaller the performance difference between the first language model and the negative sample data, the stronger the first language model's ability to distinguish between positive and negative sample data, and the higher the task processing performance. The distillation loss can be determined based on the representation difference between the first language model and the base model on positive and negative sample data. The smaller the representation difference between the first language model and the base model on positive and negative sample data, the less historical representation the first language model forgets. This will be explained in detail below.
[0063] Optionally, the task learning loss of the target task can be determined based on the processing results of the first language model on the incremental corpus samples and the training ground value of the incremental corpus samples; and the distillation loss of the first language model can be determined based on the difference in feature representation of the incremental corpus samples by the first language model and the second language model; and the training loss of the target task can be determined based on the task learning loss and the distillation loss.
[0064] Optionally, positive and negative samples of any sample can be determined from the incremental corpus samples; the sample and its positive sample form a positive sample pair, and the sample and its negative sample form a negative sample pair. Based on the features extracted from the sample and the positive sample by the first language model, a first similarity of the positive sample pair is determined, and a second similarity of the negative sample pair can be determined based on the features extracted from the sample and the negative sample by the first language model. A contrastive learning loss is constructed based on the first and second similarities, serving as the task learning loss for the target task.
[0065] The following will provide further illustrative examples of how to calculate task learning loss.
[0066] For any sample in the incremental corpus, the first language model can obtain the feature q of that sample through encoding, and can encode other samples in the incremental corpus to obtain a feature library: [k0, k1, k2, ... k n-1 Here, n represents the total number of incremental corpus samples. During the contrastive learning process, positive sample features that match feature k can be queried from the feature library. + Feature q and feature k + These are positive sample pairs. In the feature library, except for feature k... + Samples other than those in the original text can be used as negative sample features of feature q, denoted as feature k. - Feature q and any feature k - These are negative sample pairs. The learning objective of the first language model is to learn a set of encoding parameters such that the encoded feature q corresponds to feature k. + As similar as possible, and feature q and feature k - The goal is to minimize dissimilarity. In some embodiments, the task loss function can be constructed based on the negative logarithm of the ratio of the similarity between positive and negative sample pairs. Therefore, the higher the similarity of positive samples, the smaller the task loss function; conversely, the higher the similarity of negative samples, the larger the task loss function, thereby optimizing the encoding parameters. Task Loss Function L q One way to express it is as shown in the following formula:
[0067]
[0068] Where τ is a hyperparameter, which can be set according to empirical values, m is the number of negative sample features of feature q, and k i- Let represent the i-th negative sample. Based on this task loss function, the first language model can be guided to learn to distinguish between positive and negative samples, thereby improving the first language model's ability to extract features from incremental corpus samples and thus improving its performance on the target task.
[0069] The following will provide an exemplary explanation of how to calculate distillation loss.
[0070] Optionally, the incremental corpus samples include positive samples and negative samples. For any sample, its positive and negative samples can be determined to form a triple. In the triple, the sample and its positive sample can be considered as a positive sample pair, and the sample and its negative sample can be considered as a negative sample pair.
[0071] Optionally, the third similarity of a positive sample pair can be determined based on features extracted from the sample and positive samples respectively by the first language model, and the fourth similarity of a negative sample pair can be determined based on features extracted from the sample and negative samples respectively by the first language model. Furthermore, the third similarity of a positive sample pair can be determined based on features extracted from the sample and positive samples respectively by the second language model, and the fourth similarity of a negative sample pair can be determined based on features extracted from the sample and negative samples respectively by the second language model.
[0072] Based on the aforementioned similarity, the cross-entropy loss between the first and second language models for positive and negative samples can be calculated as the distillation loss. Optionally, the entropy of the similarity of positive samples can be calculated based on the first and third similarity scores, and the entropy of the similarity of negative samples can be calculated based on the second and fourth similarity scores; the distillation loss can be determined based on the cumulative value of the entropy of the similarity of positive and negative samples. Optionally, the distillation loss function L... d One way to express it is as shown in the following formula:
[0073]
[0074] in, This indicates that the base model (old model) is derived from negative sample pairs x. i The similarity between the extracted features, i.e., the fourth similarity; The first language model (new model) is represented by x. i negative sample pairs x i The similarity between the extracted features, i.e., the second similarity; This indicates that the base model (old model) is derived from positive sample pairs x. i The similarity between the extracted features, i.e., the third similarity; This indicates that the first language model (new model) is derived from positive sample pairs x. i The similarity between the extracted features is the first similarity; N represents the total number of sample pairs.
[0075] in, Entropy, which can express the similarity of negative samples; The entropy represents the similarity of positive samples. By calculating the cumulative value of the entropy of the similarity of negative samples and the entropy of the similarity of positive samples, the representational differences between the first language model and the base model on positive samples and the representational differences on negative samples can be determined, resulting in distillation loss. Based on distillation loss, the parameters of the task adaptation network in the first language model can be optimized to reduce this distillation loss, thereby reducing the probability of historical knowledge forgetting caused by incremental training.
[0076] After determining the task learning loss and distillation loss based on the aforementioned embodiments, the task learning loss and distillation loss can be weighted and calculated to obtain the total training loss of the first language model on the target task. Optimizing the parameters in the task adaptation network of the first language model based on this total training loss enables the first language model to learn new knowledge and store old knowledge, thereby enhancing its performance on the target task.
[0077] In addition to the foregoing embodiments, this application also provides a natural language processing method, such as... Figure 4 As shown, this natural language processing method mainly includes the following steps:
[0078] Step 401: Obtain the corpus data to be processed for the target task.
[0079] Step 401: Input the corpus data to be processed into the first language model to obtain the processing result; wherein, the first language model is obtained by adding a task adaptation network to the feature transformation network of the second language model; the second language model is trained by historical corpus samples; the first language model includes the parameters of the task adaptation network and the main model parameters inherited from the second language model; the parameters of the task adaptation network are obtained by incrementally learning the target task using incremental corpus samples.
[0080] The incremental training method for the first language model can be found in the description of the aforementioned embodiments, and will not be repeated here.
[0081] Optionally, when the first language processing model is a multi-task model, it can be obtained by adding task adaptation networks for different tasks to the second language processing model. In the first language processing model, different types of tasks correspond to different task adaptation networks, and these networks are used to store knowledge learned from different tasks. Accordingly, after inputting the corpus data to be processed into the first language model, the model can determine the target task adaptation network corresponding to the target task based on the type of the target task. Then, based on the parameters of the main model and the model parameters of the target task adaptation network, the corpus data to be processed is processed to obtain the processing result.
[0082] Furthermore, in this implementation, on the one hand, the first language model can be ensured to have high prediction accuracy for corpus data that matches historical corpus samples based on the main model parameters inherited from the second processing model; on the other hand, the first language model can be ensured to have high prediction accuracy for corpus data that matches incremental corpus sample types based on the parameters learned by the task adaptation network.
[0083] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 104 can be device A; or the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
[0084] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0085] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0086] Figure 5 This illustration shows a schematic diagram of the server structure provided in an exemplary embodiment of this application, such as... Figure 5 As shown, the server includes: a memory 501, a processor 502, and a communication component 503.
[0087] Memory 501 is used to store computer programs and can be configured to store various other data to support operations on the server. Examples of this data include instructions for any application or method used to operate on the server.
[0088] In some embodiments, Figure 5The server shown can be used to execute natural language processing methods. The processor 502, coupled to the memory 501, is used to execute the computer program in the memory 501 for: acquiring corpus data to be processed for a target task via a communication component 503; inputting the corpus data to be processed into a first language model to obtain a processing result; wherein the first language model is obtained by adding a task adaptation network to the feature transformation network of a second language model; the second language model is trained using historical corpus samples; the first language model includes parameters of the task adaptation network and main model parameters inherited from the second language model; the parameters of the task adaptation network are obtained by incrementally learning the target task using incremental corpus samples.
[0089] Optionally, when the processor 502 inputs the corpus data to be processed into the first language model and obtains the processing result, it specifically performs the following steps: in the first language model, it determines the target task adaptation network corresponding to the target task based on the type of the target task; and processes the corpus data to be processed based on the main model parameters and the model parameters of the target task adaptation network to obtain the processing result.
[0090] In other embodiments, Figure 5 The server shown can be used to execute model training methods. The processor 502, coupled to the memory 501, executes the computer program in the memory 501 for: acquiring incremental corpus samples required for incremental learning of the target task via the communication component 503; inputting the incremental corpus samples into a first language model to obtain the processing result of the incremental corpus samples; the first language model being obtained by adding a task adaptation network to the feature transformation network of a second language model; the second language model being trained using historical corpus samples; determining the training loss of the first language model for the target task based on the processing result; and optimizing the parameters of the task adaptation network with the goal of satisfying a set convergence condition for the training loss.
[0091] Optionally, when the processor 502 optimizes the parameters of the task adaptation network based on the training loss, it is specifically configured to: determine the task type of the target task; determine the target task adaptation network corresponding to the task type from at least one task adaptation network of the first language model; and optimize the parameters of the target task adaptation network based on the training loss.
[0092] Optionally, before inputting the incremental corpus samples into the first language model, the processor 502 is further configured to: in the pre-training sub-stage, obtain corpus samples related to the target task from the historical corpus samples of the second language model as initialization corpus samples; input the initialization corpus samples into the first language model to obtain the initialization processing result of the initialization corpus samples; determine the initialization training loss of the first language model based on the initialization processing result; determine the first type of parameters in the task adaptation network; the first type of parameters are used to learn the knowledge of the target task; and update the first type of parameters based on the initialization training loss.
[0093] Optionally, when the processor 502 acquires incremental corpus samples required for incremental learning of the target task, it includes: in the incremental training sub-stage, obtaining real-time application data associated with the target task from the application data during the operation of the first language model, as the incremental corpus samples; when the processor 502 optimizes the parameters of the task adaptation network according to the training loss, it is specifically used to: determine a second type of parameter in the task adaptation network; the second type of parameter is used to learn multi-task fusion knowledge; and update the second type of parameter according to the training loss.
[0094] Optionally, when the processor 502 determines the training loss of the first language model for the target task based on the processing result, it is specifically configured to: determine the task learning loss of the target task based on the feature expression differences of the first language model for positive sample pairs and negative sample pairs in the incremental corpus samples; and determine the distillation loss of the first language model based on the feature expression differences of the first language model and the second language model for the incremental corpus samples; and determine the training loss of the target task based on the task learning loss and the distillation loss.
[0095] Optionally, when the processor 502 determines the task learning loss for the target task based on the feature expression differences between positive and negative sample pairs in the incremental corpus samples according to the first language model, it specifically performs the following steps: determining the positive and negative samples of any sample from the incremental corpus samples; the sample and its positive sample constitute a positive sample pair, and the sample and its negative sample constitute a negative sample pair; determining a first similarity of the positive sample pair based on the features extracted from the sample and the positive sample respectively by the first language model; and determining a second similarity of the negative sample pair based on the features extracted from the sample and the negative sample respectively by the first language model; and constructing a contrastive learning loss based on the first similarity and the second similarity as the task learning loss for the target task.
[0096] Optionally, when the processor 502 determines the distillation loss of the first language model based on the differences in feature representation of the incremental corpus samples by the first language model and the second language model, it specifically performs the following steps: determining a third similarity of the positive sample pair based on the features extracted by the second language model from the sample and the positive sample respectively; and determining a fourth similarity of the negative sample pair based on the features extracted by the second language model from the sample and the negative sample respectively; calculating the entropy of the similarity of the positive sample based on the first similarity and the third similarity, and calculating the entropy of the similarity of the negative sample based on the second similarity and the fourth similarity; and determining the distillation loss based on the cumulative value of the entropy of the similarity of the positive sample and the entropy of the similarity of the negative sample.
[0097] Furthermore, such as Figure 5 As shown, the server also includes other components such as the power supply component 504. Figure 5 The diagram only shows some components and does not mean that the server only includes... Figure 5 The components shown.
[0098] The memory 501 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0099] The communication component 503 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as Wi-Fi, 2G (e.g., Global System for Mobile Communications (GSM)), 3G (e.g., Wideband Code Division Multiple Access (WCDMA), 4G (e.g., Long Term Evolution (LTE)), 4G+ (e.g., LTE-Advanced (LTE-A)), or 5G (5th Generation Mobile Communication Technology), or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component may be implemented based on Near Field Communication (NFC), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.
[0100] The power supply component 504 is used to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0101] In this embodiment, the first language model used for incremental learning of the target task is obtained by adding a task adaptation network to the feature transformation network of the second language model. Incremental corpus samples required for incremental learning are input into the first language model to obtain the processing results of the incremental corpus samples. Based on the processing results and the training ground truth of the incremental corpus samples, the training loss of the first language model is determined, and the parameters of the task adaptation network can be optimized according to the training loss. In this incremental learning process, without updating the main model parameters of the base model, the knowledge of new data is learned by updating the parameters of the task adaptation network, thereby enabling the first language model to adaptively learn new features and reducing the probability of catastrophic forgetting.
[0102] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be executed by the server in the above method embodiments.
[0103] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code.
[0104] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0105] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0106] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0107] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.
[0108] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0109] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0110] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0111] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A natural language processing method, characterized in that, include: Obtain the corpus data to be processed for the target task; The corpus data to be processed is input into the first language model to obtain the processing result; The first language model is obtained by adding a task adaptation network to the feature transformation network of the second language model; the second language model is obtained by training with historical corpus samples; the first language model includes the parameters of the task adaptation network and the main model parameters inherited from the second language model. The target task adaptation network includes a first type of parameters and a second type of parameters. The first type of parameters is used to learn knowledge about the target task, and the second type of parameters is used to learn knowledge already learned by the task adaptation networks for multiple tasks. The first type of parameters are obtained in the pre-training sub-stage by determining the training loss based on the processing result obtained by inputting the initial corpus samples associated with the target task from the historical corpus samples into the first language model and updating it according to the determined training loss. The second type of parameters are obtained in the incremental learning sub-stage by determining the training loss based on the processing result obtained by inputting the incremental corpus samples associated with the target task into the first language model and updating it according to the determined training loss.
2. The method according to claim 1, characterized in that, The corpus data to be processed is input into the first language model to obtain the processing results, including: In the first language model, the target task adaptation network corresponding to the target task is determined according to the type of the target task; Based on the main model parameters and the model parameters of the target task adaptation network, the corpus data to be processed is processed to obtain the processing result.
3. A model training method, characterized in that, The first language model is obtained by adding a task adaptation network to the feature transformation network of the second language model. The second language model is trained using historical corpus samples. The parameters of the task adaptation network for the target task include first-type parameters and second-type parameters. The first-type parameters are used to learn knowledge about the target task, and the second-type parameters are used to learn knowledge already learned by task adaptation networks for multiple tasks, including: In the pre-training sub-stage, the training loss is determined based on the processing result obtained by inputting the initial corpus samples associated with the target task from the historical corpus samples into the first language model for processing, and the first type of parameters are updated based on the determined training loss. In the incremental learning sub-stage, the training loss is determined based on the processing result obtained by inputting the incremental corpus samples associated with the target task into the first language model for processing, and the second type of parameters are updated based on the determined training loss.
4. The method according to claim 3, characterized in that, In the pre-training sub-stage, the training loss is determined based on the processing result obtained by inputting corpus samples associated with the target task into the first language model, and the first type of parameters are updated based on the determined training loss, including: In the pre-training sub-stage, corpus samples related to the target task are obtained from the historical corpus samples of the second language model and used as initial corpus samples. The initialization corpus sample is input into the first language model to obtain the initialization processing result of the initialization corpus sample; Based on the initialization processing results, determine the initialization training loss of the first language model; Determine the first type of parameters in the task adaptation network for the target task; The first type of parameters are updated based on the initial training loss.
5. The method according to claim 3, characterized in that, In the incremental learning sub-stage, the training loss is determined based on the processing result obtained by inputting corpus samples associated with the target task into the first language model, and the second type of parameters are updated based on the determined training loss, including: In the incremental training sub-stage, real-time application data related to the target task obtained from the application data during the operation of the first language model is used as incremental corpus samples. The incremental corpus sample is input into the first language model to obtain the processing result of the incremental corpus sample; Based on the processing results, the training loss of the first language model is determined; Determine the second type of parameters in the task adaptation network for the target task; The second type of parameters are updated based on the training loss.
6. The method according to claim 4 or 5, characterized in that, Based on the processing result, the training loss of the first language model for the target task is determined, including: Based on the difference in feature representation of positive and negative sample pairs in the incremental corpus samples by the first language model, the task learning loss for the target task is determined; and, Based on the differences in feature representation of the incremental corpus samples by the first language model and the second language model, the distillation loss of the first language model is determined; The training loss for the target task is determined based on the task learning loss and the distillation loss.
7. The method according to claim 6, characterized in that, Based on the difference in feature representation of positive and negative sample pairs in the incremental corpus samples by the first language model, the task learning loss for the target task is determined, including: From the incremental corpus samples, determine the positive and negative samples of any sample; the sample and its positive sample form a positive sample pair, and the sample and its negative sample form a negative sample pair; Based on the features extracted from the sample and the positive sample by the first language model, a first similarity of the positive sample pair is determined; and based on the features extracted from the sample and the negative sample by the first language model, a second similarity of the negative sample pair is determined. Based on the first similarity and the second similarity, a contrastive learning loss is constructed as the task learning loss for the target task.
8. The method according to claim 7, characterized in that, Based on the differences in feature representation of the incremental corpus samples by the first language model and the second language model, the distillation loss of the first language model is determined, including: Based on the features extracted from the sample and the positive sample respectively by the second language model, a third similarity of the positive sample pair is determined; and based on the features extracted from the sample and the negative sample respectively by the second language model, a fourth similarity of the negative sample pair is determined. Based on the first similarity and the third similarity, calculate the entropy of the similarity of positive samples, and based on the second similarity and the fourth similarity, calculate the entropy of the similarity of negative samples. The distillation loss is determined based on the cumulative value of the entropy of the similarity of the positive samples and the entropy of the similarity of the negative samples.
9. A server, characterized in that, include: Memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute one or more computer instructions for performing the steps of the method according to any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it is able to perform the steps of the method described in any one of claims 1-8.