A training method and device for a natural language processing model

By integrating and updating the prompt matrix of multiple natural language processing tasks in the first layer of the pre-trained language model, multi-task joint learning is realized, solving the problem of insufficient model representation ability and training speed in the prior art, and significantly improving the effect of the prompt adjustment method.

CN114896371BActive Publication Date: 2025-05-30ZHONGKE DINGFU BEIJING TECH DEV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210594190.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-05-30
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

In the prior art, the effect of using prompt adjustment methods to train pre-trained language models needs to be further improved, especially in terms of improving the model's representation ability and training speed.

Method used

By obtaining the pre-trained language model and determining the prompt matrix corresponding to multiple natural language processing tasks at its first layer, the prompt matrix is ​​fused and updated to achieve joint learning and implicit data enhancement of multiple tasks. The specific steps include obtaining the pre-trained language model, determining the prompt matrix for each task, updating the prompt matrix for a single task based on these matrices, and using the updated prompt matrix for model training.

Benefits of technology

Through joint learning of the prompt matrix of multiple NLP tasks, the model's representation ability and training speed are improved, and the effect of the prompt adjustment method is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114896371B_ABST
    Figure CN114896371B_ABST
Patent Text Reader

Abstract

The present application provides a training method and apparatus for a natural language processing model. Based on a pre-trained language model using a self-attention mechanism, this method updates the prompt matrix of a single natural language processing task by fusing the prompt matrices of multiple natural language processing tasks, and then inputs the training sample data of the natural language processing task and the updated prompt matrix into the model to train the updated prompt matrix. By jointly learning multiple natural language processing tasks, this method performs implicit data augmentation and improves the representation ability of the model. Since there is a progressive or similar relationship between natural language processing tasks, the effect of the prompt adjustment method can be improved through the joint learning of the prompt matrices of multiple tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular, to a method and device for training a natural language processing model. Background Art

[0002] In the field of natural language processing (NLP), "prompting" is a technique that gives artificial rules to a pre-trained language model so that the model can better understand human instructions. It can be simply understood as adding supplementary text to the input of a task to better utilize the pre-trained language model.

[0003] Compared with general fine-tuning, prompt tuning adds a prompt to the fine-tuning process and can train only the parameters of the prompt part while keeping the parameters of the entire pre-trained model fixed. This flexibility cannot be achieved by general fine-tuning.

[0004] Therefore, how to further improve the effect of training a pre-trained language model using the prompt tuning method is worthy of research. Summary of the Invention

[0005] The present application provides a method and device for training a natural language processing model, which can further improve the effect of training a pre-trained language model using the prompt tuning method.

[0006] In a first aspect, a method for training a natural language processing model is provided, including:

[0007] Obtain a pre-trained language model, where the first layer of the pre-trained language model is a layer structure using a self-attention mechanism;

[0008] Determine a first prompt matrix corresponding to a first task in the first layer and a second prompt matrix corresponding to a second task in the first layer. The first prompt matrix and the second prompt matrix are learnable vector matrices used as continuous prompts, and the first task and the second task belong to natural language processing tasks;

[0009] Determine a first coefficient matrix of the first layer according to the second prompt matrix;

[0010] Update the first prompt matrix according to the first coefficient matrix, the first prompt matrix, and the second prompt matrix;

[0011] Train the updated first prompt matrix according to the training sample data of the first task and the updated first prompt matrix, where

[0012] The input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, which is obtained by concatenating the first vector matrix of the first layer and the updated first prompt matrix, and the first vector matrix is the key vector matrix or value vector matrix corresponding to the first task.

[0013] Before training the model in the embodiments of the present application, the prompt matrices corresponding to multiple NLP tasks are fused to update the prompt matrix of a single task, and multiple NLP tasks are jointly learned, performing implicit data augmentation and improving the representation ability of the model. Since there is a progressive or similar relationship between NLP tasks, the effect of the prompt adjustment method can be improved by jointly learning the prompt matrices of multiple tasks. Further, since the prompt adjustment method in the embodiments of the present application is based on the Transformers model and the basic parameters of Transformers are shared, the prompt matrices between different tasks are equivalent to perturbations to the model on the same basic parameters. Therefore, these perturbations have certain commonalities, and explicit information can be mutually transmitted between the prompt matrices of different tasks. An NLP task can refer to the prompt matrix trained by other NLP tasks to modify its own values, which accelerates the convergence speed of the prompt matrix and thus speeds up the training speed.

[0014] In one example, the method further includes:

[0015] Determine the second coefficient matrix of the first layer according to the first prompt matrix;

[0016] Update the second prompt matrix according to the second coefficient matrix, the first prompt matrix and the second prompt matrix;

[0017] Train the updated second prompt matrix according to the training sample data of the second task and the updated second prompt matrix, where

[0018] The input of the self-attention mechanism operation corresponding to the second task in the first layer includes a second concatenated vector matrix, which is obtained by concatenating the second vector matrix of the first layer and the updated second prompt matrix, and the second vector matrix is the key vector matrix or value vector matrix corresponding to the second task.

[0019] In one example, determining the first coefficient matrix of the first layer according to the second prompt matrix includes:

[0020] Initialize the first weight matrix corresponding to the second prompt matrix, and the first weight matrix is a learnable matrix based on the training sample data of the first task and the training sample data of the second task;

[0021] Determine the first activation function and the first bias matrix corresponding to the second prompt matrix;

[0022] Take the first weight matrix, the second prompt matrix, and the first bias matrix as the inputs of the first activation function, and determine the output of the first activation function as the first coefficient matrix of the first layer.

[0023] In one example, the second layer of the pre-trained language model is a layer structure using the self-attention mechanism, and the method further includes:

[0024] Determine the third prompt matrix corresponding to the second task in the second layer, where the third prompt matrix is a learnable vector matrix used as a continuous prompt.

[0025] In one example, updating the first prompt matrix according to the first coefficient matrix, the first prompt matrix, and the second prompt matrix includes:

[0026] Determine the third coefficient matrix of the second layer according to the first prompt matrix and the third prompt matrix;

[0027] Update the first prompt matrix according to the first coefficient matrix, the third coefficient matrix, the first prompt matrix, the second prompt matrix, and the third prompt matrix;

[0028] Among them, determining the first coefficient matrix of the first layer according to the second prompt matrix includes:

[0029] Determine the first coefficient matrix according to the first prompt matrix and the second prompt matrix.

[0030] In one example, determining the first coefficient matrix according to the first prompt matrix and the second prompt matrix includes:

[0031] Determine the first Euclidean distance between the second prompt matrix and the first prompt matrix;

[0032] Based on the land moving distance algorithm, determine the first transfer amount according to the first Euclidean distance, where the first transfer amount is used to represent the proportion of the information transferred from the second prompt matrix to the first prompt matrix;

[0033] Determine the first coefficient matrix according to the first transfer amount.

[0034] In one example, determining the third coefficient matrix of the second layer according to the first prompt matrix and the third prompt matrix includes:

[0035] Determine the second Euclidean distance between the third prompt matrix and the first prompt matrix;

[0036] Based on the land moving distance algorithm, determine the second transfer amount according to the second Euclidean distance, where the second transfer amount is used to represent the proportion of the information transferred from the third prompt matrix to the first prompt matrix;

[0037] Determine the third coefficient matrix according to the second transfer amount.

[0038] In one example, updating the first hint matrix according to the first coefficient matrix, the third coefficient matrix, the first hint matrix, the second hint matrix, and the third hint matrix includes:

[0039] Determine a first proportion according to the first task of the first layer;

[0040] Determine a second proportion according to the number of layers of the pre-trained language model and the number of remaining tasks other than the first task;

[0041] Update the first hint matrix according to the first proportion, the second proportion, the first coefficient matrix, the third coefficient matrix, the first hint matrix, the second hint matrix, and the third hint matrix.

[0042] In a second aspect, a training device for a natural language processing model is provided, including:

[0043] A model acquisition module, configured to acquire a pre-trained language model, where the first layer of the pre-trained language model is a layer structure using a self-attention mechanism;

[0044] A hint matrix determination module, configured to determine a first hint matrix corresponding to the first task in the first layer and a second hint matrix corresponding to the second task in the first layer. The first hint matrix and the second hint matrix are learnable vector matrices used as continuous hints, and the first task and the second task belong to natural language processing tasks;

[0045] A coefficient matrix determination module, configured to determine a first coefficient matrix of the first layer according to the second hint matrix;

[0046] A hint matrix update module, configured to update the first hint matrix according to the first coefficient matrix, the first hint matrix, and the second hint matrix. The first coefficient matrix is related to the second hint matrix;

[0047] A model training module, configured to train the updated first hint matrix according to the training sample data of the first task and the updated first hint matrix, where

[0048] The input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, and the first concatenated vector matrix is obtained by concatenating the first vector matrix of the first layer and the updated first hint matrix. The first vector matrix is a key vector matrix or a value vector matrix corresponding to the first task.

[0049] In one example, the device further includes:

[0050] The coefficient matrix determination module is further configured to determine a second coefficient matrix according to the first hint matrix;

[0051] The hint matrix module is further configured to update the second hint matrix according to the second coefficient matrix, the first hint matrix, and the second hint matrix;

[0052] The model training module is further configured to train the updated second prompt matrix according to the training sample data of the second task and the updated second prompt matrix, where

[0053] the input of the self-attention mechanism operation corresponding to the second task in the first layer includes a second concatenated vector matrix, which is obtained by concatenating the second vector matrix in the first layer and the updated second prompt matrix, and the second vector matrix is a key vector matrix or a value vector matrix corresponding to the second task.

[0054] Before training the model, the device according to the embodiments of the present application fuses the prompt matrices corresponding to multiple NLP tasks to update the prompt matrix of a single task, jointly learns multiple NLP tasks, performs implicit data augmentation, and improves the representation ability of the model. Since there is a progressive relationship or a similar relationship between NLP tasks, the effect of the prompt adjustment method can be improved by jointly learning the prompt matrices of multiple tasks. Further, since the prompt adjustment method in the embodiments of the present application is based on the Transformers model and the basic parameters of the Transformers are shared, the prompt matrices between different tasks are equivalent to perturbations to the model on the same basic parameters. Therefore, these perturbations have certain commonalities, and explicit information can be transmitted between the prompt matrices of different tasks. An NLP task can refer to the prompt matrix trained by other NLP tasks to modify its own value, which accelerates the convergence speed of the prompt matrix and thus speeds up the training speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 is a schematic flowchart of a method for training a natural language processing model provided by an embodiment of the present application;

[0057] Figure 2 is a schematic flowchart of another method for training a natural language processing model provided by an embodiment of the present application;

[0058] Figure 3 is a schematic diagram of a pre-trained language model architecture provided by an embodiment of the present application;

[0059] Figure 4 is a schematic diagram of a combination of prompt matrices provided by an embodiment of the present application;

[0060] Figure 5 is a schematic diagram of a method for obtaining the coefficient matrix g provided by an embodiment of the present application;

[0061] Figure 6 It is a schematic diagram of a training device for a natural language processing model provided by an embodiment of the present application. Detailed implementation manners

[0062] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0063] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0064] For the convenience of understanding the solutions in the present application, the following briefly introduces some technical concepts:

[0065] Prompt learning: A technology that gives artificial rules to a pre-trained model so that the model can better understand human instructions. It can be simply understood as adding supplementary text to the input of a task to better utilize the pre-trained model. In prompt learning, the supplementary text can be in the form of a prompt template as the input of the model. The production of the prompt template is divided into manually creating a template and automatically generating a template, and the automatic generation of the template is further divided into discrete prompts (also called hard prompts) and continuous prompts (also called soft prompts).

[0066] Hard Prompt: In a hard prompt, the prompt is an actual text string. For example, the input text x = "I love this movie". First, design a prompt template: Overall it was a[z]movie. In actual research, [z] is the position that needs to be filled by the model, and the position and number of [z] determine the type of prompt. For example, according to the different positions of [z], the prompt can be divided into cloze prompt ([z] in the sentence) and prefix prompt ([z] at the end of the sentence). Which one to choose specifically depends on the task form and model category.

[0067] Soft Prompt: In a soft prompt, the prompt is directly described in the embedding space of the underlying language model. For example, in the "prompt tuning" method, continuous prompts are learned by inserting trainable variables into the embedded input.

[0068] Prompt Tuning: Add the prompt to the fine-tuning process, and it is possible to train only the parameters of the prompt part while keeping the parameters of the entire pre-trained model fixed.

[0069] NLP Tasks: Deep learning models can use NLP tasks to learn language knowledge. NLP tasks include NLP basic tasks and NLP upper-layer tasks (i.e., NLP downstream tasks). Among them, deep learning models can use NLP basic tasks to learn basic language knowledge. NLP basic tasks include, for example, part-of-speech tagging (POS) tasks, chunking (CHUNK) tasks, and dependency parsing (DEP) tasks, etc.; NLP upper-layer tasks are tasks learned by deep learning models in specific applications, such as text semantic relatedness (Relatedness) tasks, text entailment (Entailment) tasks, and named entity recognition (NER) tasks, etc.

[0070] Earth Mover's Distance (EMD) is an image similarity metric method proposed in the IJCV journal article "The Earth Mover's Distance as a Metric for Image Retrieval". Initially, the concept of EMD was used for image retrieval, and later, due to its various advantages, it was gradually used for similarity metrics in other aspects.

[0071] Transformer Layer: Transformer is a model architecture proposed in the 2017 paper "Attention is All You Need", which presents a Transformers network consisting of stacked Transformer-Blocks. Each Transformer-Block has the same structure and includes a self-attention mechanism. The purpose of this attention operation is to calculate the "relevance" between the current representation (token) and each position, thereby determining the proportion of the vector of each position in the context at the final time step. The attention formula used in the Transformer layer is where q is the query vector matrix (abbreviated as the Q vector matrix), k is the key vector matrix (abbreviated as the K vector matrix), and v is the value vector matrix (abbreviated as the V vector matrix).

[0072] Bidirectional Encoder Representation from Transformers (BERT) Language Model: BERT uses a Masked Language Model (MLM) for pre-training and constructs the entire model using deep bidirectional Transformer components. In recent years, research on pre-trained language models (PLMs) has been abundant, and natural language processing has also made great progress with this trend. Especially between 2017 and 2019, researchers gradually shifted their focus from supervised models with traditional task features to pre-training. The research idea based on pre-trained language models is usually "pre-train, fine-tune", that is, applying the PLM to downstream tasks and designing training objects according to the downstream tasks during the pre-training and fine-tuning stages and adjusting the PLM itself.

[0073] As the size of the PLM continues to increase, the hardware requirements, data needs, and actual costs for fine-tuning it are also rising. In addition, the rich and diverse downstream tasks make the design of the pre-training and fine-tuning stages cumbersome and complex. Therefore, researchers hope to explore more compact, lightweight, and general-purpose and efficient methods, and Prompt Tuning is an attempt in this direction. Prompt Tuning can adapt the model to downstream tasks by adjusting only a few parameters.

[0074] At present, the effect of training PLMs using prompt tuning methods needs to be further improved.

[0075] It should be noted that there is currently no general Chinese explanation in the art for the Transformers layer. Therefore, in this application, Transformers is used to refer to such a layer structure or model.

[0076] To further improve the performance of the model and its few-shot learning ability, an embodiment of this application provides a method for training a natural language processing model, as Figure 1 shown in Method 100 below, where Figure 1 FIG. 11 is a schematic flowchart of a method for training a natural language processing model provided by an embodiment of this application. Method 100 includes:

[0077] S110, obtaining a pre-trained language model.

[0078] Among them, the first layer of the pre-trained language model is a layer structure using the self-attention mechanism.

[0079] Exemplarily, the pre-trained language model includes transformers layers, and the pre-trained language model is T5, RoBERTAa, DeBERTa, etc.

[0080] It should be understood that the pre-trained language model may include other layers, and the other layers may be layer structures using the self-attention mechanism or other types of layer structures.

[0081] S120, determining a first prompt matrix corresponding to a first task in the first layer and a second prompt matrix corresponding to a second task in the first layer.

[0082] Among them, the first prompt matrix and the second prompt matrix are learnable vector matrices used as continuous prompts, and the first task and the second task belong to natural language processing tasks.

[0083] Among them, the first prompt matrix corresponding to the first task in the first layer can be understood as setting a prompt matrix for each task in each layer, and the prompt matrix is trained and updated by the sample data of its corresponding task to possess the characteristics of that task.

[0084] Exemplarily, natural language processing tasks include part-of-speech tagging tasks, chunking tasks, dependency parsing tasks, named entity recognition (NER) tasks, relation extraction tasks, etc.

[0085] Exemplarily, the initial first prompt matrix and second prompt matrix are matrices after random initialization.

[0086] S130. Update the first hint matrix according to the first coefficient matrix of the first layer, the first hint matrix, and the second hint matrix.

[0087] Among them, the first coefficient matrix is determined according to the second hint matrix.

[0088] Exemplarily, add the product of the first coefficient matrix and the second hint matrix to the first hint matrix to obtain the updated first hint matrix.

[0089] The method for updating the first hint matrix includes: First, randomly initialize the first weight matrix and the first bias matrix. Among them, the first weight matrix and the first bias matrix are learnable matrices based on the training sample data of the first task and the training sample data of the second task, and the first weight matrix and the first bias matrix are used to linearly map the second hint matrix. Then, input the training sample data of the first task and the training sample data of the second task into the pre-trained language model, and update the first weight matrix of the first layer after model training. Then, determine the first activation function. Optionally, the first activation function is the sigmoid function. Finally, add the product of the first weight matrix and the second hint matrix plus the first bias matrix as the input of the first activation function, and the output of the first activation function is the first coefficient matrix.

[0090] In one example, the model is also used to train a third task. The method includes: Determine the fourth hint matrix corresponding to the third task in the first layer. The fourth hint matrix is a learnable vector matrix used as a continuous hint, and the third task belongs to the natural language processing task. Among them, the method for updating the first hint matrix further includes:

[0091] Update the first hint matrix according to the first coefficient matrix, the second hint matrix, the fourth coefficient matrix, the fourth hint matrix, and the first hint matrix of this layer. Among them, the fourth coefficient matrix is determined according to the fourth hint matrix. For the method of determining the fourth coefficient matrix, refer to the method of determining the first coefficient matrix, which will not be elaborated here.

[0092] Further optionally, add the product of the first coefficient matrix and the second hint matrix to the first hint matrix, and continue to add the product of the fourth coefficient matrix and the fourth hint matrix to the first hint matrix to obtain the updated first hint matrix.

[0093] Further, the model is used to train multiple tasks. The multiple tasks at least include the above-mentioned first task and second task, and may also include the third task. Of course, other tasks may also be included. Taking the example of including the first task, the second task, and the third task, the effect will be described below in combination with the above embodiments:

[0094] The first task is a basic language task such as a part-of-speech analysis task, the second task is a basic language task such as a chunk analysis task, and the third task is a downstream task such as a named entity recognition task. According to the above embodiments, it can be summarized that first, an initial first prompt matrix corresponding to the first task, an initial second prompt matrix corresponding to the second task, and an initial fourth prompt matrix corresponding to the third task are determined. The initial first prompt matrix, second prompt matrix, and fourth prompt matrix can be randomly initialized. Then, according to the prompt matrix update method of the above embodiments, the initial first prompt matrix is updated according to the second prompt matrix and the fourth prompt matrix, and so on. The initial first prompt matrix can be updated according to the prompt matrices corresponding to all other tasks in the current layer. Then, the updated first prompt matrix is trained according to the training sample data; the initial second prompt matrix is updated according to the first prompt matrix and the fourth prompt matrix, and so on. The initial second prompt matrix can be updated according to the prompt matrices corresponding to all other tasks in the current layer. Then, the updated second prompt matrix is trained according to the training sample data; the initial fourth prompt matrix is updated according to the first prompt matrix and the second prompt matrix, and so on. The initial fourth prompt matrix can be updated according to the prompt matrices corresponding to all other tasks in the current layer. Then, the updated fourth prompt matrix is trained according to the training sample data. It can be seen that when training multiple tasks with a pre-trained language model, after determining the initial prompt matrix for each task, the prompt matrix after being processed by all other tasks in the same layer is used to update according to the above method (since the proportion of the prompt matrix corresponding to each task itself is 1 to participate in its own update, it can also be said that the initial prompt matrix for each task is updated according to the prompt matrices after being processed by all tasks in the same layer). The multiple tasks include NLP basic tasks and downstream tasks. According to the above method of updating the initial prompt matrix for each task, the information of multiple tasks can be integrated, and the learning effect of a single task can be improved when learning the features of the training sample data, which is beneficial to improving the effect of the prompt adjustment method.

[0095] Further, the update method of the second hint matrix is similar to that of the first hint matrix, specifically as follows. The second hint matrix is updated according to the second coefficient matrix, the first hint matrix, and the second hint matrix of the first layer, and the second coefficient matrix is determined according to the first hint matrix. In one example, first, the second weight matrix and the second bias matrix are randomly initialized, where the second weight matrix and the second bias matrix are learnable matrices. Then, the training sample data of the first task and the training sample data of the second task are input into the pre-trained language model, and the second weight matrix of the first layer is updated after model training. Then, the first activation function is determined. Optionally, the first activation function is the sigmoid function. Finally, the product of the second weight matrix and the first hint matrix plus the second bias matrix is used as the input of the first activation function, and the output of the first activation function is the second coefficient matrix.

[0096] In one example, the second layer of the pre-trained language model is a layer structure adopting the self-attention mechanism, and the method further includes:

[0097] Determine the third hint matrix corresponding to the second task in the second layer, where the third hint matrix is a learnable vector matrix used as continuous hints.

[0098] It should be noted that the "first" in the first layer or the "second" in the second layer in the examples of this application are only used to distinguish two different layer structures and can be understood as a certain layer, rather than specifically referring to the first few layers in the model.

[0099] The method for updating the first hint matrix further includes:

[0100] The first hint matrix is updated according to the first coefficient matrix of the first layer, the third coefficient matrix of the second layer, the first hint matrix, the second hint matrix, and the third hint matrix, where the third coefficient matrix is determined according to the first hint matrix and the third hint matrix, and the first coefficient matrix is determined according to the first hint matrix and the second hint matrix.

[0101] Among them, the methods for determining the first coefficient matrix and the third coefficient matrix include:

[0102] Determine the first Euclidean distance between the second hint matrix and the first hint matrix, and the second Euclidean distance between the third hint matrix and the first hint matrix;

[0103] Based on the earth mover's distance algorithm, obtain the first transfer amount according to the first Euclidean distance. The first transfer amount is used to represent the proportion of the information transferred from the second hint matrix to the first hint matrix, and determine the first coefficient matrix according to the first transfer amount;

[0104] The second transfer amount is obtained according to the second Euclidean distance based on the land movement distance algorithm. The second transfer amount is used to characterize the proportion of the information transmitted from the third hint matrix to the first hint matrix, and the third coefficient matrix is determined according to the second transfer amount.

[0105] It should be understood that when the first transfer amount and the second transfer amount are each a numerical value, the first coefficient matrix and the third coefficient matrix can each be a numerical value.

[0106] Exemplarily, the method for updating the first hint matrix further includes:

[0107] Determining a first proportion according to the first task of the first layer;

[0108] Determining a second proportion according to the number of layers of the pre-trained language model and the number of remaining tasks other than the first task;

[0109] Updating the first hint matrix according to the first proportion, the second proportion, the first coefficient matrix, the third coefficient matrix, the first hint matrix, the second hint matrix, and the third hint matrix.

[0110] For example, the updated first hint matrix = the initial first hint matrix × the first proportion + the second proportion × the second hint matrix × the first coefficient matrix + the second proportion × the third hint matrix × the third coefficient matrix.

[0111] Optionally, the value of the first proportion is 1.

[0112] It can be seen from the above embodiments that the multiple tasks at least include the above-mentioned first task and second task. Of course, other tasks can also be included. Taking the inclusion of the first task and the second task as an example, the effects are described in combination with the above embodiments:

[0113] The first task is a basic language task such as a chunk analysis task, and the second task is a downstream task such as a named entity recognition task. According to the above embodiments, it can be summarized that first, an initial first prompt matrix corresponding to the first task and an initial second prompt matrix corresponding to the second task are determined. The initial first prompt matrix and the second prompt matrix can be randomly initialized. Then, according to the prompt matrix update method of the above embodiments, the initial first prompt matrix of the first layer is updated according to the second prompt matrix corresponding to the second task in the first layer and the third prompt matrix corresponding to the second task in the second layer, and so on. The initial first prompt matrix can be updated according to the prompt matrices corresponding to all other tasks in all layers in the above manner, and then the updated first prompt matrix is trained according to the training sample data. Therefore, the initial prompt matrix corresponding to each task in each layer can be updated according to the prompt matrices corresponding to all other tasks in all layers. The prompt matrix corresponding to each task in each layer can integrate the information of multiple tasks, which can improve the learning effect of a single task when learning the features of the training sample data, and is beneficial to improving the effect of the prompt adjustment method.

[0114] S140. Train the updated first prompt matrix according to the training sample data of the first task and the updated first prompt matrix.

[0115] Among them, the input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, and the first concatenated vector matrix is obtained by concatenating the first vector matrix of the first layer and the updated first prompt matrix. The first vector matrix is a key vector matrix or a value vector matrix corresponding to the first task.

[0116] Specifically, input the training sample data of the first task and the updated first prompt matrix into the pre-trained language model, train the updated first prompt matrix, input the output of the last output layer into the loss function to calculate the loss, and then use the above method to update the trained first prompt matrix. Then, use the updated first prompt matrix and the training sample data of the first task as the input of the pre-trained language model. After iterating a certain number of times until the derivative of the loss function is 0, the final first prompt matrix is determined.

[0117] In an example, train the updated second prompt matrix according to the training sample data of the second task and the updated second prompt matrix, where

[0118] the input of the self-attention mechanism operation corresponding to the second task in the first layer includes a second concatenated vector matrix, and the second concatenated vector matrix is obtained by concatenating the second vector matrix of the first layer and the updated second prompt matrix. The second vector matrix is a key vector matrix or a value vector matrix corresponding to the second task.

[0119] Before training the model, the embodiment of the present application fuses the prompt matrices corresponding to multiple NLP tasks to update the prompt matrix of a single task, and jointly learns multiple NLP tasks to perform implicit data enhancement, thereby improving the representation ability of the model. Since there is a progressive relationship or similarity relationship between NLP tasks, the effect of the prompt adjustment method can be improved by jointly learning the prompt matrices of multiple tasks. Furthermore, since the prompt adjustment method in the embodiment of the present application is based on the Transformers model, the basic parameters of Transformers are shared, and the prompt matrices between different tasks are equivalent to perturbing the model on the same basic parameters, so these perturbations have certain commonalities, and the prompt matrices of different tasks can explicitly transfer information to each other. An NLP task can refer to the prompt matrices trained by other NLP tasks to modify its own values, which accelerates the convergence speed of the prompt matrix, thereby accelerating the training speed.

[0120] Based on method 100, this application combines specific pre-training models and training steps to illustrate method 100 in detail through the following embodiments. Figure 2 FIG. 1 is a schematic flow chart of another example of a natural language processing model training method provided in an embodiment of the present application. Figure 2 As shown in method 200 in , taking the BERT model based on the transformers architecture as an example, the method 200 may include:

[0121] S210, determining training sample data corresponding to a plurality of tasks respectively.

[0122] For example, the training sample data for the text classification task is: "The weather is really nice today". In the BERT model, the above text will become ["[CLS]", "today", "day", "weather", "air", "true", "no", "wrong", "[SEP]"]. Among them, [CLS] and [SEP] are symbols in BERT that indicate the beginning and end of the text.

[0123] S220, determining prompt matrices corresponding to the multiple tasks.

[0124] Specifically, the prompt matrix of the K vector matrix corresponding to the original model when calculating self-attention is P k , P k It is used to concatenate with the K vector matrix to obtain a new K vector matrix. The prompt matrix of the V vector matrix corresponding to the original model when calculating the self-attention is P v , P v It is used to concatenate with the V vector matrix to obtain a new V vector matrix. The new K vector matrix and the new V vector matrix participate in the operation of the self-attention mechanism, so that P k and P v Got trained.

[0125] Figure 3 This is a schematic diagram of a pre-trained language model architecture provided in an embodiment of the present application. Figure 3 P k and P v For introduction. Figure 3 As shown, taking a task as an example, P is set in each layer of the BERT model. k and P v , P k It can be seen as h 0 ,h 1 ,…,h i The vector matrix composed of i can be adjusted according to the task, where h 0 to h i It is the parameter vector matrix related to the training task. Only these parameters are updated during training. The shape of each vector matrix is ​​1*768. v It can be seen as h 0’ ,h 1’ ,…,h i’ The vector matrix composed of 0’ to h i’ See h 0 to h i Assume that input x = "Amazing!", after being processed by the embedding layer, the input x is represented by a vector, which is e([CLS]), e(Amazing) and e(!) (e is the embedding function of the model). The word vector is input into the BERT model. By annotating the data to P k and P v For training, since the parameters of the BERT model are frozen, that is, they do not participate in training, only P k and P v Conduct training.

[0126] Combination Figure 3 , the following example illustrates the meaning of "splicing":

[0127] Assume that the text input to the BERT model is "You are a good student", a K vector matrix of the transformer layer corresponds to [CLS] you are a good student [SEP], K vector and P k The spliced ​​form is as follows 0 ,h 1 ,…,h i [CLS]You are a good student [SEP].

[0128] P k The concatenation with the K vector matrix can be expressed as:

[0129] K’ = contact(P k , K) Equation (1)

[0130] Where K’ represents the concatenated K vector matrix.

[0131] Based on Figure 3 , taking the training of a task as an example, the transformers model in this application is introduced, including (1) Embedding layer (embedding layer); (2) Transformers-Block (Transformers-block structure), generally multiple; (3) Output layer.

[0132] Among them, the Embedding layer is used to map the text into a matrix. Define the input text as X, and X has z words. Then the input of the Embedding layer is the index corresponding to the text of length z (this index is the index of each word in the model vocabulary, and this model vocabulary is the vocabulary trained according to the open-source bert model of Google. This model vocabulary is used to encode the text to obtain the index), and the output E is a matrix of size [z, d], where d is the length of the matrix converted from the index corresponding to each word by the Embedding layer. Among them, z = 512 and d = 768.

[0133] Multiple Transformers-Blocks are stacked to form the transformers layer. The Transformers Block will split the word vectors, and the number of splits is called "head". For example, each original word vector is 300-dimensional and there are 5 heads in total. Then each head sequentially takes the h-th part of the 300 dimensions that are split into 5 parts (each part has 60 dimensions), and puts the h parts after splitting into different Transformers Blocks respectively. In the subsequent embodiments of this application, the number of heads is 1 for illustration, and the number of layers in the model is L. For example, L = 6. The vector matrices Q, K, and V of each layer affect the output of each layer through the operation of the self-attention mechanism.

[0134] The output layer outputs the content corresponding to the corresponding task according to different tasks. For example, the output of text classification is the category probability of the text, the output of the named entity recognition task is the probability of each word classification, and the relation extraction task needs to extract the probabilities of the subject, object, and event of the text, etc.

[0135] It should be noted that this application does not limit the pre-trained language model, as long as it is a model based on the transformer layer, such as T5, RoBERTAa, DeBERTa, etc.

[0136] Based on the above transformers model, assume that the prompt matrices for Task 1 in the m-th layer are P 1,m,k (for concatenating with the K vector matrix) and P 1,m,v (for concatenating with the V vector matrix). First, randomly initialize the prompt matrices for each task. Then, during the training process of this round, determine the new prompt matrix for Task 1 in the current layer based on the prompt matrices of other tasks in this layer. Assume there are T tasks, and the re-determined P 1,m,k is as shown in the following formula:

[0137]

[0138] where,

[0139]

[0140] Figure 4 is a schematic diagram of an example of a combination of prompt matrices provided in an embodiment of the present application. Figure 4 One circle in it represents an h i . Next, in combination with Figure 4 , introduce formula (2). Among them, P′ 1,m,k is the prompt matrix of Task 1 in the m-th layer updated based on the prompt matrices of other tasks. P 1,m,k to P T,m,k are the prompt matrices of Task 1 to Task T in the m-th layer. is the weight for weighted summation (a coefficient matrix), and its subscript (1,2) represents the number of this weight. This number is directional (i.e., Task 2 transfers a certain proportion of information to Task 1), and so on to It can be seen that P′ 1,m,k is formed by the interaction between the prompt matrices of multiple tasks including itself, that is, the prompt matrices of different tasks are added to the original P 1,m,k according to a certain weight, thereby updating the prompt matrix of Task 1 in the m-th layer.

[0141] Figure 5 is a schematic diagram of a method for obtaining an example of the coefficient matrix g provided in an embodiment of the present application. Next, in combination with Figure 5 , introduce formula (3). Among them, i and j represent different tasks. and are the weight matrix and bias matrix corresponding to the linear mapping in the m-th layer respectively, and σ is the sigmoid activation function. Taking as an example, after being transposed, it performs matrix multiplication with P 2,m,k , then adds the parameter and finally is activated by the sigmoid function. It should be understood that the activation function can also be other types of functions, and the present application does not limit this.

[0142] Parameter matrix W ij and After random initialization, the values in the subsequent matrices are trained using the training sample data of tasks i and j until the updated P′ 1,m,k meets the preset conditions. For example, the preset condition is that the reciprocal of the loss function corresponding to the task is 0.

[0143] Optionally, we share the parameter matrix W of the linear mapping in multiple layers ij and b ij , that is The length of the vector matrix is the same as that of P 1,m,k , Each value in the matrix is between 0 and 1 ( Figure 5 in a circle in i (h i represents a token that serves as a hint)) represents the proportion of usage of each h

[0144] The re-determined P 1,m,v is as shown in the following formula:

[0145]

[0146] where

[0147]

[0148] It can be seen that P 1,m,k and P 1,m,v share Since the connection between multiple tasks does not change whether it is the hint matrix corresponding to the K vector matrix or the hint matrix corresponding to the V vector matrix, sharing can be done This can also reduce the parameters for model training.

[0149] And so on, the update of the hint matrices of multiple tasks in the m-th layer is as shown in the following formula:

[0150]

[0151]

[0152] Finally, the hint matrices of other layers are determined in the above manner.

[0153] S230, input the hint matrix into the model and train the model.

[0154] Specifically, after the multi-task text is input into the model, the output of each layer has a K vector matrix and a V vector matrix. After the original K vector matrix and the original V vector matrix are respectively concatenated with the corresponding updated prompt matrix, a new K vector matrix and a new V vector matrix are obtained. Then, the new K vector matrix and the new V vector matrix are used as the input for the self-attention mechanism operation of this layer. After the self-attention mechanism operation, the output of this layer is obtained, and the output of this layer is used as the input for the next layer until the final output layer of the model outputs the content corresponding to each task. Finally, the loss is calculated, and all the prompt matrices are updated according to the loss.

[0155] Among them, the concatenated K vector matrix corresponding to multiple tasks in the m-th layer is shown in the following formula:

[0156]

[0157] Among them, K T,m represents the original K vector matrix corresponding to task T in the m-th layer, and K′ T,m represents the new concatenated K vector matrix corresponding to task T in the m-th layer.

[0158] Among them, the concatenated V vector matrix corresponding to multiple tasks in the m-th layer is shown in the following formula:

[0159]

[0160] Among them, V T,m represents the original V vector matrix corresponding to task T in the current layer, and V′ T,m represents the new concatenated V vector matrix corresponding to task T in the current layer.

[0161] It should be understood that S220 and S230 can be repeated to continuously update the prompt matrix to obtain a better prompt matrix.

[0162] It should be noted that the multiple tasks in this method can be tasks with similar learning objectives, such as multiple sentiment classification tasks, or tasks with a progressive relationship (such as multiple tasks including NLP basic tasks and NLP downstream tasks like named entity recognition tasks, dependency parsing tasks, relation extraction tasks, etc.).

[0163] The model's learning on multiple NLP tasks can improve its learning effect on a single task. In the prompt tuning method, since the basic parameters of the original model are frozen and only the parameters of the prompt matrix are trained, task-related specific parameters can be learned by adjusting fewer parameters. Starting from this, in Method 200, the prompt matrix for each task in the same layer is formed by the weighted sum of the prompt matrices of other tasks, which performs implicit data augmentation and improves the model's representation ability. Since there is a progressive or similar relationship between NLP tasks, the effect of the prompt tuning method can be improved by jointly learning the prompt matrices of multiple tasks. Further, since the prompt tuning method of this application embodiment is based on the Transformers model and the basic parameters of Transformers are shared, the prompt matrices between different tasks are equivalent to perturbations to the model on the same basic parameters. Therefore, these perturbations have certain commonalities, and explicit information can be mutually transmitted between the prompt matrices of different tasks. One task can refer to the prompt matrix trained by other tasks to modify its own value, which accelerates the convergence speed of the prompt matrix and thus speeds up the training speed.

[0164] It should be noted that in Method 200, the example of training multiple tasks simultaneously is used for illustration. Each prompt matrix will be bound to its corresponding task, and the binding method can be, for example, assigning the same task label, etc. This application does not make any restrictions on this. It is also possible to train only one task at a time and update the prompt matrices of these multiple tasks after training multiple tasks.

[0165] Method 200 considers the information transmission between multiple tasks in the same layer. This application also provides an example of the information transmission method between multiple tasks in different layers. Method 300 will be introduced below with reference to Method 200.

[0166] S310. Determine the training sample data of multiple tasks.

[0167] For the specific content, refer to S210 and will not be elaborated here.

[0168] S320. Determine the prompt matrices of multiple tasks.

[0169] Taking the prompt matrix P for Task 1 in the first layer used for splicing with the K vector matrix as an example, the prompt matrix is updated through the following steps: 1,1,k For example, the prompt matrix is updated through the following steps:

[0170] S321. Calculate the Euclidean distance between Task 1 in the first layer and other tasks in multiple layers.

[0171] d 1,1,k =‖P 1,1,k - [P 2,1,k , …, P 2,L,k , P3,1,k ,…,P 3,L,k ,…,P T,1,k ,…,P T,L,k ‖ Formula (10)

[0172] where ‖X - Y‖ represents calculating the Euclidean distance between X and Y, and d 1,1,k is an array of length (T - 1)*L, which can also be regarded as a one-dimensional matrix of length (T - 1)*L.

[0173] S322. Calculate the set of transfer amounts from other tasks in multiple layers to Task 1 in the first layer according to the Euclidean distance.

[0174] Specifically, use the EMD algorithm to calculate the set of transfer amounts f 1,1,k from other tasks in multiple layers (assuming there are L layers) to Task 1 in the first layer. 1,1,k (i.e., the coefficient matrix). Each element of the set of transfer amounts f 1,1,k is the transfer amount from a task in other tasks in a certain layer to Task 1 (each transfer amount is a numerical value, and the transfer amount can also be understood as the proportion of the information that a task in other tasks in a certain layer can transfer to Task 1). For example, the transfer amount from Task 2 in the first layer to Task 1 in the first layer is denoted as Then f 1,1,k is as shown in the following formula:[[]]END]]

[0175]

[0176] where all elements in f 1,1,k sum to 1, and it is a one-dimensional matrix of length (T - 1)*L.

[0177] S323. Update the hint matrix of Task 1 in the first layer according to the set of transfer amounts.

[0178] The hint matrix of Task 1 in the first layer corresponding to the K vector matrix after update is denoted as P′ 1,1,k , and P′ 1,1,k is as shown in the following formula:[[]]END]]

[0179] P′ 1,1,k = P 1,1,k + f 1,1,k * P other,k Formula (12)

[0180] P other,k = [P 2,1,k ,…,P 2,L,k ,P 3,1,k ,…,P 3,L,k ,…,P T,1,k ,…,P T,L,k Formula (13)

[0181] Optionally, to reduce the influence of other hint matrices on the updated hint matrix, P′ 1,1,k can also be as shown in the following formula:

[0182] P′ 1,1,k = P 1,1,k + α 1,1,k f 1,1,k * P other,k Formula (14)

[0183] where the weight α 1,1,k is a value between 0 and 1 (including 0 and 1).

[0184] Optionally, the value of α 1,1,k is

[0185] It should be understood that formula (12) can also be transformed as shown in the following formula:

[0186] P′ 1,1,k = α 1,1,k * f′ 1,1,k * P all,k Formula (15)

[0187]

[0188] P all,k = [P 1,1,k , …, P 1,L,k , P 2,1,k , …, P 2,L,k , P 3,1,k , …, P 3,L,k , …, P T,1,k , …, P T,L,k Formula (17)

[0189] where the part of f′ 1,1,k that is 0 indicates that the transfer amount of task 1 of other layers to task 1 of the first layer is 0, and the part of 1 / α 1,1,k indicates that the transfer amount of task 1 of the first layer to itself is 1 (i.e., α 1,1,k * 1 / α 1,1,k = 1), and P all,k is the matrix corresponding to all K vector matrices.

[0190] According to the above content, it can be analogized to the update method of the hint matrix corresponding to the V vector matrix. The updated hint matrix of task 1 of the first layer corresponding to the V vector matrix is denoted as P′ 1,1,v , P′ 1,1,v as shown in the following formula:

[0191] P′1,1,v = P 1,1,v + f 1,1,v * P other,v Formula (18)

[0192] P other,v = [P 2,1,v , …, P 2,L,v , P 3,1,v , …, P 3,L,v , …, P T,1,v , …, P T,L,v Formula (19)

[0193] Optionally, in order to reduce the influence of other hint matrices on the hint matrix to be updated, P′ 1,1,v can also be as shown in the following formula:

[0194] P′ 1,1,v = P 1,1,v + α 1,1,v f 1,1,v * P other,v Formula (20)

[0195] where α 1,1,v is a value between 0 and 1 (including 0 and 1).

[0196] It should be understood that Formula (12) can also be transformed as shown in the following formula:

[0197] P′ 1,1,v = α 1,1,v * f′ 1,1,v * P all,v Formula (21)

[0198]

[0199] P all,v = [P 1,1,v , …, P 1,L,v , P 2,1,v , …, P 2,L,v , P 3,1,v , …, P 3,L,v , …, P T,1,v , …, P T,L,v Formula (23)

[0200] where f′ 1,1,v the part that is 0 indicates that the transfer amount of task 1 of other layers to task 1 of the first layer is 0, and the part of 1 / α 1,1,v indicates that the transfer amount of task 1 of the first layer to itself is 1 (i.e., α 1,1,v * 1 / α 1,1,v = 1), and P all,v is the matrix corresponding to all V vector matrices.

[0201] And so on, the update of the prompt matrix for the t-th task in the m-th layer is shown in the following formula:

[0202] P′ t,m,k = α t,m,k * f′ t,m,k * P all,k Formula (24)

[0203]

[0204] P′ t,m,v = α t,m,v * f′ t,m,v * P all,v Formula (26)

[0205]

[0206] S330, input the prompt matrix into the model and train the model.

[0207] For specific content, refer to S230, which will not be elaborated here.

[0208] It should be noted that this application does not limit the above algorithm for calculating the transfer amount, as long as the transfer amount from other tasks in each layer to the task to be updated can be obtained.

[0209] In Method 300, since the transfer amount can be regarded as a weight, the prompt matrix of a task is composed of the weighted sum of the prompt matrices of multiple tasks in different layers, which performs implicit data augmentation and improves the representation ability of the model. Since there is a progressive or similar relationship between NLP tasks, the effect of the prompt adjustment method can be effectively improved by jointly learning the prompt matrices of multiple tasks between different layers. Further, since the prompt adjustment method of this application embodiment is based on the Transformers model and the basic parameters of Transformers are shared, the prompt matrices between different tasks are equivalent to perturbing the model on the same basic parameters. Therefore, these perturbations have certain commonalities, and explicit information can be mutually transmitted between the prompt matrices of multiple tasks in different layers. A task can refer to the prompt matrix trained by other tasks to modify its own value, which accelerates the convergence speed of the prompt matrix and thus speeds up the training speed.

[0210] Figure 6 This is a schematic diagram of a training device for a natural language processing model provided by an embodiment of this application. Based on the above natural language processing model training method, this application also provides a training device for a natural language processing model. The following will describe this device in conjunction with Figure 6 and illustrate this device as Figure 6 shown. This device includes:

[0211] A model acquisition module 410, configured to acquire a pre-trained language model, wherein the first layer of the pre-trained language model is a layer structure adopting a self-attention mechanism;

[0212] A prompt matrix determination module 420, configured to determine a first prompt matrix corresponding to a first task in the first layer and a second prompt matrix corresponding to a second task in the first layer. The first prompt matrix and the second prompt matrix are learnable vector matrices used as continuous prompts, and the first task and the second task belong to natural language processing tasks;

[0213] A coefficient matrix determination module 430, configured to determine a first coefficient matrix of the first layer according to the second prompt matrix;

[0214] A prompt matrix update module 440, configured to update the first prompt matrix according to the first coefficient matrix and the first prompt matrix of the first layer. The first coefficient matrix is related to the second prompt matrix;

[0215] A model training module 450, configured to train the updated first prompt matrix according to the training sample data of the first task and the updated first prompt matrix, wherein

[0216] The input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, which is obtained by concatenating the first vector matrix of the first layer and the updated first prompt matrix. The first vector matrix is a key vector matrix or a value vector matrix corresponding to the first task.

[0217] For other implementation manners of the device, refer to the descriptions in Method 100 to Method 300, which will not be elaborated herein.

[0218] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps does not have a strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments. Their execution order does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0219] The above are only some implementation manners of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A training method for a natural language processing model, characterized in that, comprising: Obtaining a pre-trained language model, wherein the first layer of the pre-trained language model is a layer structure adopting a self-attention mechanism; Determining a first prompt matrix corresponding to the first task in the first layer and a second prompt matrix corresponding to the second task in the first layer, where the first prompt matrix and the second prompt matrix are learnable vector matrices used as continuous prompts, and the first task and the second task belong to natural language processing tasks; Initializing a first weight matrix corresponding to the second prompt matrix, where the first weight matrix is a learnable matrix based on the training sample data of the first task and the training sample data of the second task; Determining a first activation function and a first bias matrix corresponding to the second prompt matrix; Taking the first weight matrix, the second prompt matrix, and the first bias matrix as inputs of the first activation function, and determining the output of the first activation function as the first coefficient matrix of the first layer; Obtaining the product of the first coefficient matrix and the second prompt matrix; Adding the product to the first prompt matrix to obtain an updated first prompt matrix; Training the updated first prompt matrix according to the training sample data of the first task and the updated first prompt matrix, where, The input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, which is obtained by concatenating the first vector matrix of the first layer and the updated first prompt matrix, and the first vector matrix is a key vector matrix or a value vector matrix corresponding to the first task.

2. The method according to claim 1, characterized in that, the method further comprises: Determining a second coefficient matrix of the first layer according to the first prompt matrix; Updating the second prompt matrix according to the second coefficient matrix, the first prompt matrix, and the second prompt matrix; Training the updated second prompt matrix according to the training sample data of the second task and the updated second prompt matrix, where, The input of the self-attention mechanism operation corresponding to the second task in the first layer includes a second concatenated vector matrix, which is obtained by concatenating the second vector matrix of the first layer and the updated second prompt matrix, and the second vector matrix is a key vector matrix or a value vector matrix corresponding to the second task.

3. A training method for a natural language processing model, characterized in that, comprising: Obtaining a pre-trained language model, wherein the first layer of the pre-trained language model is a layer structure adopting a self-attention mechanism, and the second layer of the pre-trained language model is a layer structure adopting a self-attention mechanism; Determining a first prompt matrix corresponding to the first task in the first layer and a second prompt matrix corresponding to the second task in the first layer, where the first prompt matrix and the second prompt matrix are learnable vector matrices used as continuous prompts, and the first task and the second task belong to natural language processing tasks; Determine the third prompt matrix corresponding to the second task in the second layer, where the third prompt matrix is a learnable vector matrix used for continuous prompting; Determine the first Euclidean distance between the second prompt matrix and the first prompt matrix; Determine a first transfer amount based on the land moving distance algorithm according to the first Euclidean distance, where the first transfer amount is used to characterize the proportion of information transferred from the second prompt matrix to the first prompt matrix; Determine a first coefficient matrix according to the first transfer amount; Determine the second Euclidean distance between the third prompt matrix and the first prompt matrix; Determine a second transfer amount based on the land moving distance algorithm according to the second Euclidean distance, where the second transfer amount is used to characterize the proportion of information transferred from the third prompt matrix to the first prompt matrix; Determine a third coefficient matrix according to the second transfer amount; Update the first prompt matrix according to the first coefficient matrix, the third coefficient matrix, the first prompt matrix, the second prompt matrix, and the third prompt matrix; Train the updated first prompt matrix according to the training sample data of the first task and the updated first prompt matrix, where, The input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, which is obtained by concatenating the first vector matrix in the first layer and the updated first prompt matrix, and the first vector matrix is a key vector matrix or a value vector matrix corresponding to the first task.

4. The method according to claim 3, wherein, The updating the first prompt matrix according to the first coefficient matrix, the third coefficient matrix, the first prompt matrix, the second prompt matrix, and the third prompt matrix includes: Determine a first proportion according to the first task in the first layer; Determine a second proportion according to the number of layers of the pre-trained language model and the number of remaining tasks other than the first task; Update the first prompt matrix according to the first proportion, the second proportion, the first coefficient matrix, the third coefficient matrix, the first prompt matrix, the second prompt matrix, and the third prompt matrix.

5. A training device for a natural language processing model, wherein, comprises: A model acquisition module for acquiring a pre-trained language model, where the first layer of the pre-trained language model is a layer structure using a self-attention mechanism; A prompt matrix determination module for determining a first prompt matrix corresponding to the first task in the first layer and a second prompt matrix corresponding to the second task in the first layer, where the first prompt matrix and the second prompt matrix are learnable vector matrices used for continuous prompting, and the first task and the second task belong to natural language processing tasks; A coefficient matrix determination module for initializing a first weight matrix corresponding to the second prompt matrix, where the first weight matrix is a learnable matrix based on the training sample data of the first task and the training sample data of the second task; Determine the first activation function and the first bias matrix corresponding to the second hint matrix; use the first weight matrix, the second hint matrix, and the first bias matrix as the input of the first activation function, and determine the output of the first activation function as the first coefficient matrix of the first layer; A hint matrix update module, configured to obtain the product of the first coefficient matrix and the second hint matrix; Add the product to the first hint matrix to obtain an updated first hint matrix; A model training module, configured to train the updated first hint matrix according to the training sample data of the first task and the updated first hint matrix, where The input of the self-attention mechanism operation corresponding to the first task in the first layer includes a first concatenated vector matrix, which is obtained by concatenating the first vector matrix of the first layer and the updated first hint matrix, and the first vector matrix is a key vector matrix or a value vector matrix corresponding to the first task.

6. The apparatus according to claim 5, wherein, The apparatus further includes: The coefficient matrix determination module is further configured to determine a second coefficient matrix according to the first hint matrix; The hint matrix update module is further configured to update the second hint matrix according to the second coefficient matrix, the first hint matrix, and the second hint matrix; The model training module is further configured to train the updated second hint matrix according to the training sample data of the second task and the updated second hint matrix, where The input of the self-attention mechanism operation corresponding to the second task in the first layer includes a second concatenated vector matrix, which is obtained by concatenating the second vector matrix of the first layer and the updated second hint matrix, and the second vector matrix is a key vector matrix or a value vector matrix corresponding to the second task.

Citation Information

Patent Citations

  • Rotary machine missing fault feature recovery method and system

    CN113935252A

  • Natural language processing method and device based on knowledge guidance prefix fine tuning, computing equipment, and storage medium

    CN113987209A