A training method and device for a natural language processing model

By setting the learnable vector matrix and key-value vector splicing in each layer of the pre-trained language model, the problem of insufficient effect and small sample learning ability in natural language processing is solved, and the performance of the model in basic and upper-level NLP tasks is improved.

CN114625840BActive Publication Date: 2025-07-08ZHONGKE DINGFU BEIJING TECH DEV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210272019.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-07-08
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

The existing Prompt method has room for improvement in the effect and small sample learning ability in natural language processing.

Method used

By setting the first learnable vector matrix and the second learnable vector matrix for prompting in each layer of the pre-trained language model, and splicing it with the key-value vectors of the current layer, participating in training and updating, reducing training parameters, and embedding basic language knowledge to improve the small sample learning ability of the model.

Benefits of technology

It improves the small sample learning ability of the natural language processing model and enhances the performance of the model in basic and upper-level NLP tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114625840B_ABST
    Figure CN114625840B_ABST
Patent Text Reader

Abstract

The present application provides a training method and apparatus for a natural language processing model. The method includes: obtaining a pre-trained language model in which each layer is a layer structure adopting a self-attention mechanism, obtaining a first learnable vector matrix and a second learnable vector matrix for each layer to learn a first task in an NLP task, then generating a first concatenated key vector matrix and a first concatenated value vector matrix for each layer according to the first learnable vector matrix and the second learnable vector matrix, and finally training the first learnable vector matrix and the second learnable vector matrix using the training sample data of the first task. Through the first concatenated key vector matrix and the first concatenated value vector matrix, the first learnable vector matrix and the second learnable vector matrix are involved in the training. Since the pre-trained language model is fixed, the training parameters are greatly reduced; enabling the learnable vector matrix to first learn the NLP basic task and then learn the NLP upper-level task can improve the learning ability of small samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a method and apparatus for training a natural language processing model. Background Art

[0002] In the field of Natural Language Processing (NLP), "Prompt" is a technique that gives artificial rules to a pre-trained model so that the model can better understand human instructions. It can be simply understood as adding supplementary text to the input of a task to better utilize the pre-trained model.

[0003] Compared with general Fine-Tuning, Prompt Tuning adds Prompts to the fine-tuning process and can train only the parameters of the Prompt part while keeping the parameters of the entire pre-trained model fixed. This flexibility cannot be achieved by general Fine-Tuning.

[0004] However, there is still room for improvement in the effectiveness and few-shot learning ability of the Prompt method. Summary of the Invention

[0005] This application provides a method and apparatus for training a natural language processing model, which can further improve the effectiveness of the Prompt method and the few-shot learning ability.

[0006] In a first aspect, a method for training a natural language processing model is provided, including:

[0007] Obtain a pre-trained language model, where each layer of the pre-trained language model is a layer structure using a self-attention mechanism;

[0008] Obtain a first learnable vector matrix and a second learnable vector matrix for each layer. Both the first learnable vector matrix and the second learnable vector matrix are used to learn a first task in NLP;

[0009] Generate a first concatenated key vector matrix and a first concatenated value vector matrix for each layer. The first concatenated key vector matrix is obtained by concatenating the first learnable vector matrix of the current layer with the key vector of the current layer, and the first concatenated value vector matrix is obtained by concatenating the second learnable vector matrix of the current layer with the value vector of the current layer;

[0010] Input the training sample data of the first task into the pre-trained language model, calculate the output of the current layer according to the first concatenated key vector matrix, the first concatenated value vector matrix of each layer and the query vector of the current layer, and obtain the optimal first learnable vector matrix and second learnable vector matrix of each layer according to the output of the last layer and the loss function of the first task.

[0011] In the embodiment of the present application, a first learnable vector matrix and a second learnable vector matrix for prompting are set in each layer of the pre-trained language model, and then the first concatenated key vector matrix obtained by concatenating the first learnable vector matrix with the key vector of the current layer, and the first concatenated value vector matrix obtained by concatenating the second learnable vector matrix with the value vector of the current layer are used as the input of the current layer, so that the first learnable vector matrix and the second learnable vector matrix participate in the training and update of each layer. Since the parameters of the original pre-trained language model are fixed, the number of training parameters is greatly reduced. After inputting the training samples of the NLP task into the model, the first learnable vector matrix and the second learnable vector matrix can learn the NLP task, so that the prompt of each layer has the characteristics of the NLP task. In practical applications, first embed the learnable vector matrix used to learn the basic tasks in the NLP task (such as part-of-speech analysis task, chunk analysis task, dependency parsing task, etc.) to inject basic language knowledge into the model, and then embed the learnable vector matrix used to learn the upper-layer tasks of the NLP, which is very helpful for learning the upper-layer tasks of the NLP (such as named entity recognition task, text semantics-related task, text entailment task, classification task, etc.), and further improves the small-sample learning ability of the model, that is, improves the effect of prompt learning.

[0012] In a possible implementation manner, the first learnable vector matrix of each layer is obtained by concatenating the third learnable vector matrix of the current layer with the learnable vector matrices obtained from all NLP tasks participated in before the current layer, and the second learnable vector matrix of each layer is obtained by concatenating the fourth learnable vector matrix of the current layer with the learnable vector matrices obtained from all NLP tasks participated in before the current layer, where the third learnable vector matrix and the fourth learnable vector matrix are both vector matrices set to learn the first task.

[0013] In a possible implementation manner, after obtaining the optimal first learnable vector matrix and second learnable vector matrix of each layer, the method further includes:

[0014] Set the fifth learnable vector matrix and the sixth learnable vector matrix of each layer, where the fifth learnable vector matrix and the sixth learnable vector matrix are both used to learn the second task in the NLP task, and the first task is different from the second task;

[0015] Generate the first concatenated learnable vector matrix and the second concatenated learnable vector matrix for each layer. The first concatenated learnable vector matrix is obtained by concatenating the fifth learnable vector matrix of the current layer with the first learnable vector matrix of the current layer, and the second concatenated learnable vector matrix is obtained by concatenating the sixth learnable vector matrix of the current layer with the second learnable vector matrix of the current layer;

[0016] Generate the second concatenated key vector matrix and the second concatenated value vector matrix for each layer. The second concatenated key vector matrix is obtained by concatenating the first concatenated learnable vector matrix of the current layer with the key vector of the current layer, and the second concatenated value vector matrix is obtained by concatenating the second concatenated learnable vector matrix of the current layer with the value vector of the current layer;

[0017] Input the training sample data of the second task into the pre-trained language model, calculate the output of the current layer according to the second concatenated key vector matrix, the second concatenated value vector matrix of each layer and the query vector of the current layer, obtain the optimal first concatenated learnable vector matrix and the second concatenated learnable vector matrix of each layer according to the output of the last layer and the loss function of the second task, and restrict the parameter changes of the first learnable vector matrix and the parameter changes of the second learnable vector matrix.

[0018] In a possible implementation, the NLP task includes NLP basic tasks and NLP upper-layer tasks. The NLP basic tasks include part-of-speech analysis tasks, chunk analysis tasks and dependency syntax analysis tasks. The NLP upper-layer tasks include named entity recognition tasks, text semantics-related tasks, text entailment tasks and classification tasks.

[0019] In a possible implementation, all NLP tasks participated in before the current layer are NLP basic tasks, and the first task is an NLP upper-layer task.

[0020] In a possible implementation, the output h1 of the first layer of the pre-trained language model is implemented by the following formula:

[0021] Or,

[0022]

[0023] Among them, q1 is the query vector matrix input to the model, k1 is the key vector matrix input to the model, v1 is the value vector matrix input to the model, represents concatenation, prompt_K is the learnable vector matrix corresponding to k1, prompt_V is the learnable vector matrix corresponding to v1, is the vector matrix after concatenating k1 and prompt_K, is the vector matrix after concatenating v1 and prompt_V, f is the function used in the current layer, and h1 is the output of the current layer.

[0024] In a possible implementation, the output h of the nth layer of the pre-trained language model n is implemented using the following formula, where n is an integer and n > 1:

[0025] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v ;

[0026] Or,

[0027]

[0028] where h n-1 is the output of the (n - 1)th layer, W q is the linear transformation matrix of the query vector in the current layer, W k is the linear transformation matrix of the key vector in the current layer, W v is the linear transformation matrix of the value vector in the current layer, q n is the query vector in the current layer, k n is the key vector in the current layer, v n is the value vector in the current layer, represents concatenation, prompt_K is the learnable vector matrix corresponding to k n , prompt_V is the learnable vector matrix corresponding to v n , is the vector matrix after concatenating k n and prompt_K, is the vector matrix after concatenating v n and prompt_V, f is the function used in the current layer, h n is the output of the current layer.

[0029] In a possible implementation, prompt_K and prompt_V are implemented using the following formula:

[0030] Or,

[0031] Or,

[0032] wherein, prompt c1 _K represents a learnable vector matrix corresponding to the key vector of the current layer, which is set to learn the current NLP task, prompt c2 _K represents a learnable vector matrix corresponding to the key vector of the current layer obtained by training all NLP tasks participated in before the current layer, prompt c1 _V represents a learnable vector matrix corresponding to the value vector of the current layer, which is set to learn the current NLP task, prompt c2 _V represents a learnable vector matrix corresponding to the value vector of the current layer obtained by training all NLP tasks participated in before the current layer.

[0033] In a possible implementation, when the pre-trained language model is learning the first NLP task, the loss function includes the sum of the loss part of the current NLP task and the part restricting the parameters of the current NLP task; or,

[0034] When the pre-trained language model is learning a non-first NLP task, the loss function includes the sum of the loss part of the current NLP task, the part restricting the parameters of the current NLP task, and the part restricting the parameters that have been trained for all NLP tasks participated in before.

[0035] In a possible implementation, before generating the first concatenated learnable vector matrix and the second concatenated learnable vector matrix of each layer, the method further includes: performing layer normalization on the first learnable vector matrix and the second learnable vector matrix, and / or, performing layer normalization on the fifth learnable vector matrix and the sixth learnable vector matrix.

[0036] In a second aspect, a training device for a natural language processing model is provided, including:

[0037] A first module, configured to obtain a pre-trained language model, and each layer of the pre-trained language model is a layer structure adopting a self-attention mechanism;

[0038] A second module, configured to obtain a first learnable vector matrix and a second learnable vector matrix of each layer, and both the first learnable vector matrix and the second learnable vector matrix are used to learn the first task in the NLP task;

[0039] A third module, configured to generate a first concatenated key vector matrix and a first concatenated value vector matrix of each layer, where the first concatenated key vector matrix is obtained by concatenating the first learnable vector matrix of the current layer and the key vector of the current layer, and the first concatenated value vector matrix is obtained by concatenating the second learnable vector matrix of the current layer and the value vector of the current layer;

[0040] The fourth module is used to input the training sample data of the first task into the pre-trained language model, calculate the output of the current layer according to the first concatenated key vector matrix, the first concatenated value vector matrix of each layer and the query vector of the current layer, and obtain the optimal first learnable vector matrix and second learnable vector matrix of each layer according to the output of the last layer and the loss function of the first task.

[0041] A training device for a natural language processing model provided by an embodiment of the present application sets a first learnable vector matrix and a second learnable vector matrix for prompt in each layer, and then uses the first concatenated key vector matrix obtained by concatenating the first learnable vector matrix with the key vector of the current layer, and the first concatenated value vector matrix obtained by concatenating the second learnable vector matrix with the value vector of the current layer as the input of the current layer, so that the first learnable vector matrix and the second learnable vector matrix participate in the training and update of each layer. Since the parameters of the original pre-trained language model are fixed, the number of training parameters is greatly reduced. After inputting the training samples of the NLP task into the model, the first learnable vector matrix and the second learnable vector matrix can learn the NLP task, so that the prompt of each layer has the characteristics of the NLP task. In practical applications, first embed the learnable vector matrix for learning basic tasks in NLP (such as part-of-speech analysis tasks, chunk analysis tasks, dependency syntax analysis tasks, etc.) to inject basic language knowledge into the model, and then embed the learnable vector matrix for learning upper-layer tasks in NLP. This is very helpful for learning upper-layer tasks in NLP (such as named entity recognition tasks, text semantics-related tasks, text entailment tasks, classification tasks, etc.), and further improves the small-sample learning ability of the model, that is, improves the effect of prompt learning. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a schematic flowchart of a training method for a natural language processing model provided by an embodiment of the present application;

[0044] Figure 2 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of the present application;

[0045] Figure 3 It is a schematic diagram of the architecture of a pre-trained language model provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of the present application;

[0047] Figure 5 It is a schematic diagram of the architecture of another pre-trained language model provided by an embodiment of the present application;

[0048] Figure 6 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of the present application;

[0049] Figure 7 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of the present application. Detailed implementation manners

[0050] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0051] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0052] To facilitate the understanding of the solutions in the present application, the following briefly introduces some technical concepts:

[0053] Prompt Learning: A technique that gives artificial rules to a pre-trained model to enable the model to better understand human instructions. It can be simply understood as adding supplementary text to the input of a task to better utilize the pre-trained model. In prompt learning, the supplementary text can be in the form of a prompt template as the input to the model. The creation of prompt templates is divided into manually creating templates and automatically generating templates, and the automatic generation of templates is further divided into discrete prompts (also called hard prompts) and continuous prompts (also called soft prompts).

[0054] Hard Prompt: In a hard prompt, the prompt is an actual text string. For example, the input text x = "I love this movie". First, design a prompt template: Overall it was a[z]movie. In actual research, [z] is the position that needs to be filled by the model, and the position and number of [z] determine the type of prompt. For example, according to the different positions of [z], the prompt can be divided into cloze prompt ([z] in the sentence) and prefix prompt ([z] at the end of the sentence). Which one to choose specifically depends on the task form and model type.

[0055] Soft Prompt: In a soft prompt, the prompt is directly described in the embedding space of the underlying language model. For example, in the "prompt tuning" method, continuous prompts are learned by inserting trainable variables into the embedded input.

[0056] Prompt Tuning (P-tuning): Add the prompt to the fine-tuning process, and it can be done to only train the parameters of the prompt part while keeping the parameters of the entire pre-trained model fixed.

[0057] NLP Basic Tasks: Deep learning models can use NLP basic tasks to learn basic language knowledge. NLP basic tasks include, for example, part-of-speech tagging (POS) tasks, chunking (CHUNK) tasks, and dependency parsing (DEP) tasks, etc.

[0058] NLP Advanced Tasks: Tasks learned by deep learning models in specific applications, such as text semantic relatedness (Relatedness) tasks, text entailment (Entailment) tasks, and named entity recognition (NER) tasks, etc.

[0059] NLP Multi-Task Learning: In the field of NLP, there is often a hierarchical relationship among various tasks. For example, from lexical analysis to syntactic analysis to upper-level practical application tasks. Specific tasks can be: Part-of-Speech tagging (POS) -> Chunking (CHUNK) -> Dependency Parsing (DEP) -> Text Semantic Relatedness -> Text Entailment. That is, the model can be trained layer by layer, first learning part-of-speech, then chunking, and so on, and finally applying it to upper-level tasks.

[0060] Transformer Layer: Transformer is a model architecture proposed in a 2017 paper "Attention is All You Need". It proposed the self-attention mechanism. The purpose of this attention operation is to calculate the "relatedness" between the current representation (token) and each position, so as to determine the proportion of the vector of each position in the context of the final time step (timestep). The attention formula used in the transformer model is where q is the query vector matrix, k is the key vector matrix, and v is the value vector matrix.

[0061] Bidirectional Encoder Representation from Transformers (BERT) Language Model: BERT uses the Masked Language Model (MLM) for pre-training and adopts deep bidirectional Transformer components to build the entire model.

[0062] Since 2018, the scale of pre-trained models has been increasing continuously, such as T5, GPT-3, Wudao, etc. Large models have become a very important technological breakthrough in the field of NLP. At the same time, the hardware and data requirements during the fine-tuning process of pre-trained models are also increasing continuously. Rich downstream tasks also make the design of the pre-training and fine-tuning stages more complex. Currently, Prompt Tuning can train only the parameters of the Prompt part while keeping the parameters of the entire pre-trained model fixed, providing a new idea for training large models. However, its performance on downstream tasks and its learning ability with few samples need to be improved.

[0063] It should be noted that there is currently no general Chinese explanation for the Transformer layer in this field. Therefore, in this application, "Transformer" is used to refer to such a layer structure or model.

[0064] To further improve the performance of the model and its few-shot learning ability, the embodiments of this application provide a method for training a natural language processing model, as Figure 1 shown in Method 100 in Figure 1 which is a schematic flowchart of a method for training a natural language processing model provided by the embodiments of this application. Method 100 includes:

[0065] S110, obtain a pre-trained language model, where each layer of the pre-trained language model is a layer structure using the self-attention mechanism.

[0066] In one example, the layer structure using the self-attention mechanism is the Transformer layer, and the pre-trained language model is the BERT model.

[0067] According to the above introduction of the self-attention mechanism, the layer using the self-attention mechanism calculates the output by taking the query vector, key vector, and value vector as the input of the hidden layer.

[0068] S120, obtain the first learnable vector matrix and the second learnable vector matrix for each layer.

[0069] Among them, both the first learnable vector matrix and the second learnable vector matrix are used to learn the first task in the NLP task.

[0070] In a possible implementation, the NLP task includes NLP basic tasks and NLP upper-layer tasks. Among them, the NLP basic tasks include part-of-speech analysis tasks, chunk analysis tasks, and dependency parsing tasks; the NLP upper-layer tasks include named entity recognition tasks, text semantics-related tasks, text entailment tasks, and classification tasks.

[0071] The first learnable vector matrix and the second learnable vector matrix embedded in each layer of the model can be regarded as prompt matrices. In practical applications, first embed the learnable vector matrix used to learn NLP basic tasks (such as part-of-speech analysis tasks, chunk analysis tasks, dependency parsing tasks, etc.) to inject basic language knowledge into the model, and then embed the learnable vector matrix used to learn NLP upper-layer tasks. This is very helpful for the model to learn NLP upper-layer tasks (such as named entity recognition tasks, text semantics-related tasks, text entailment tasks, classification tasks, etc.), and further improves the few-shot learning ability of the model, that is, improves the effect of prompt learning.

[0072] Since the first task can be the first task learned by the model, or other tasks may have been learned and trained before learning the first task, in this method, the first learnable vector matrix and the second learnable vector matrix of each layer have two meanings: one is the matrix after splicing processing, and the other is the matrix that has not been spliced and is first embedded in the model.

[0073] If the first learnable vector matrix and the second learnable vector matrix of each layer are matrices after splicing processing, for the sake of the flexibility of the solution, when the first learnable vector matrix and the second learnable vector matrix are obtained by splicing matrices that have learned NLP basic tasks, they can also be used to learn NLP upper-layer tasks, such as named entity recognition tasks, text semantics-related tasks, text entailment tasks, classification tasks, etc.

[0074] In an example, when the first learnable vector matrix and the second learnable vector matrix of each layer are matrices after splicing processing, there are the following implementation methods:

[0075] The first learnable vector matrix of each layer is obtained by splicing the third learnable vector matrix of the current layer with the learnable vector matrices obtained by training all NLP tasks participated in before the current layer. The second learnable vector matrix of each layer is obtained by splicing the fourth learnable vector matrix of the current layer with the learnable vector matrices obtained by training all NLP tasks participated in before the current layer. Among them, both the third learnable vector matrix and the fourth learnable vector matrix are set to learn the first task. Since the first learnable vector matrix contains the third learnable vector matrix, and the second learnable vector matrix contains the fourth learnable vector matrix, we can also say that the first learnable vector matrix and the second learnable vector matrix are also matrices used to learn the first task. It should be noted that in actual applications, although the learnable vector matrices obtained by training all NLP tasks participated in before the current layer participate in the training of the new task, the change of their parameters will be restricted compared with their states when they were first trained well. Further optionally, there are two splicing orders for the third learnable vector matrix of the current layer and the learnable vector matrices obtained by training all NLP tasks participated in before the current layer: one is that the third learnable vector matrix of the current layer is in the front and the learnable vector matrices obtained by training all NLP tasks participated in before the current layer are in the back; the other is that the third learnable vector matrix of the current layer is in the back and the learnable vector matrices obtained by training all NLP tasks participated in before the current layer are in the front. Different splicing orders will make the relative positions of the matrices different, affecting the results of the layer output.

[0076] The following is an example to illustrate "front" or "back". Suppose there are matrix A and matrix B that need to be concatenated. A = [a, b, c, d], B = [e, f, g, h], where a, b, c, d, e, f, g, h are all 1*15 dimensional vector matrices. Then, when A is in the front and B is in the back, the concatenated matrix is [a, b, c, d, e, f, g, h]; when A is in the back and B is in the front, the concatenated matrix is [e, f, g, h, a, b, c, d]. Therefore, "front" or "back" should be understood as the logical order of concatenating matrices, rather than the spatial position of the matrices.

[0077] S130. Generate the first concatenated key vector matrix and the first concatenated value vector matrix for each layer.

[0078] Specifically, the first concatenated key vector matrix is obtained by concatenating the first learnable vector matrix of the current layer with the key vector of the current layer, and the first concatenated value vector matrix is obtained by concatenating the second learnable vector matrix of the current layer with the value vector of the current layer.

[0079] Optionally, there are two concatenation orders for the first learnable vector matrix of the current layer and the key vector of the current layer, and there are two concatenation orders for the second learnable vector matrix of the current layer and the value vector of the current layer. For specific content, refer to the description of the concatenation order scheme in S120, which will not be elaborated here.

[0080] S140. Use the first task training sample data to train and update the first learnable vector matrix and the second learnable vector matrix of each layer.

[0081] Specifically, input the training sample data of the first task into the pre-trained language model, calculate the output of the current layer according to the first concatenated key vector matrix, the first concatenated value vector matrix of each layer and the query vector of the current layer, and obtain the optimal first learnable vector matrix and the second learnable vector matrix of each layer according to the output of the last layer and the loss function of the first task.

[0082] In the specific implementation of S110 to S140, the present application provides the following formula for easy understanding of the above scheme.

[0083] In one example, the output h1 of the first layer of the pre-trained language model is implemented by the following formula:

[0084] Or,

[0085]

[0086] where q1 is the query vector matrix input to the model, k1 is the key vector matrix input to the model, and v1 is the value vector matrix input to the model. Denotes concatenation, prompt_K is a learnable vector matrix corresponding to k1, and prompt_V is a learnable vector matrix corresponding to v1. is the vector matrix after concatenating k1 and prompt_K. is the vector matrix after concatenating v1 and prompt_V, f is the function used in the current layer, and h1 is the output of the current layer.

[0087] In one example, the size of prompt_K in the depth direction and the size of prompt_V in the depth direction are both the same as the size of the hidden layer of the pre-trained language model in the depth direction.

[0088] Furthermore, the output h of the nth layer of the pre-trained language model n is implemented using the following formula, where n is an integer and n > 1:

[0089] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v ;

[0090] Or,

[0091]

[0092] where h n-1 is the output of the (n - 1)th layer, W q is the linear transformation matrix of the query vector of the current layer, W k is the linear transformation matrix of the key vector of the current layer, W v is the linear transformation matrix of the value vector of the current layer, q n is the query vector of the current layer, k n is the key vector of the current layer, v n is the value vector of the current layer, Denotes concatenation, prompt_K is a learnable vector matrix corresponding to k n , and prompt_V is a learnable vector matrix corresponding to v n . is the vector matrix after concatenating k n and prompt_K, is the vector matrix after concatenating v n and prompt_V, f is the function used in the current layer, and h n is the output of the current layer.

[0093] In one example, both prompt_K and prompt_V are matrices after splicing processing, and there are two ways of splicing order. prompt_K and prompt_V are implemented using the following formulas:

[0094] Or,

[0095] Or,

[0096] where prompt c1 _K represents a learnable vector matrix corresponding to the key vector of the current layer that is set to learn the current NLP task, and prompt c2 _K represents a learnable vector matrix corresponding to the key vector of the current layer obtained from the training of all NLP tasks participated in before the current layer, and prompt c1 _V represents a learnable vector matrix corresponding to the value vector of the current layer that is set to learn the current NLP task, and prompt c2 _V represents a learnable vector matrix corresponding to the value vector of the current layer obtained from the training of all NLP tasks participated in before the current layer.

[0097] In one example, after obtaining the optimal first learnable vector matrix and second learnable vector matrix for each layer, the method further includes:

[0098] Setting a fifth learnable vector matrix and a sixth learnable vector matrix for each layer. Both the fifth learnable vector matrix and the sixth learnable vector matrix are used to learn a second task in the NLP task, and the first task is different from the second task;

[0099] Generating a first spliced learnable vector matrix and a second spliced learnable vector matrix for each layer. The first spliced learnable vector matrix is obtained by splicing the fifth learnable vector matrix of the current layer with the first learnable vector matrix of the current layer, and the second spliced learnable vector matrix is obtained by splicing the sixth learnable vector matrix of the current layer with the second learnable vector matrix of the current layer;

[0100] Generating a second spliced key vector matrix and a second spliced value vector matrix for each layer. The second spliced key vector matrix is obtained by splicing the first spliced learnable vector matrix of the current layer with the key vector of the current layer, and the second spliced value vector matrix is obtained by splicing the second spliced learnable vector matrix of the current layer with the value vector of the current layer;

[0101] Input the training sample data of the second task into the pre-trained language model, calculate the output of the current layer according to the second concatenated key vector matrix, the second concatenated value vector matrix of each layer and the query vector of the current layer, obtain the optimal first concatenated learnable vector matrix and the second concatenated learnable vector matrix of each layer according to the output of the last layer and the loss function of the second task, and limit the parameter changes of the first learnable vector matrix and the parameter changes of the second learnable vector matrix.

[0102] It should be noted that although the first concatenated learnable vector matrix and the second concatenated learnable vector matrix participate in the training of the second task, in actual applications, the parameter changes of the first learnable vector matrix and the parameter changes of the second learnable vector matrix after being trained by the first task will be restricted to prevent forgetting.

[0103] Optionally, there are two ways to concatenate the fifth learnable vector matrix of the current layer and the first learnable vector matrix of the current layer, and there are two ways to concatenate the sixth learnable vector matrix of the current layer and the second learnable vector matrix of the current layer. For the specific implementation method, please refer to the above description of the concatenation order scheme, which will not be elaborated here.

[0104] In one example, first perform the embedding of the corresponding learnable vector matrices and the learning of their respective corresponding basic tasks in the order of part-of-speech analysis task, chunk analysis task, and dependency syntax analysis task, and then embed the learnable vector matrices for learning NLP upper-layer tasks. After the learnable vector matrices have basic language knowledge, then perform the learning of upper-layer tasks, which is beneficial to improving the ability of few-shot learning and further improving the prompt effect.

[0105] In a possible implementation manner, when the pre-trained language model is learning the first NLP task, the loss function includes the sum of the loss part of the current NLP task and the part restricting the parameters of the current NLP task; or,

[0106] When the pre-trained language model is learning a non-first NLP task, the loss function includes the sum of the loss part of the current NLP task, the part restricting the parameters of the current NLP task, and the part restricting the parameters that have been trained for all previous NLP tasks participated in.

[0107] In one example, before generating the first concatenated learnable vector matrix and the second concatenated learnable vector matrix of each layer, the method further includes: performing layer normalization on the first learnable vector matrix and the second learnable vector matrix, and / or performing layer normalization on the fifth learnable vector matrix and the sixth learnable vector matrix.

[0108] Among them, Layer Normalization normalizes all neuron nodes of a single sample in each layer, which can improve the training speed and accuracy of the model and make the model more robust.

[0109] In method 100, a first learnable vector matrix and a second learnable vector matrix for prompt are set in each layer. Then, the first concatenated key vector matrix obtained by concatenating the first learnable vector matrix with the key vector of the current layer, and the first concatenated value vector matrix obtained by concatenating the second learnable vector matrix with the value vector of the current layer are used as the input of the current layer, so that the first learnable vector matrix and the second learnable vector matrix participate in the training and update of each layer. Since the parameters of the original pre-trained language model are fixed, the number of training parameters is greatly reduced. After inputting the training samples of the NLP task into the model, the first learnable vector matrix and the second learnable vector matrix can learn the NLP task, so that the learnable vector matrix of each layer has the characteristics of the NLP task. In practical applications, first embed the learnable vector matrix for learning the basic tasks in the NLP task (such as part-of-speech analysis task, chunk analysis task, dependency parsing task, etc.) to inject basic language knowledge into the model, and then embed the learnable vector matrix for learning the upper-layer tasks of NLP. This is very helpful for learning the upper-layer tasks of NLP (such as named entity recognition task, text semantics-related task, text entailment task, classification task, etc.), and further improves the small-sample learning ability of the model, that is, improves the effect of prompt learning.

[0110] Based on method 100, this application combines a specific pre-trained model and training steps, and details method 100 through the following embodiments. Figure 2 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of this application. As Figure 2 shown in method 200 in, taking the pre-trained language model bert with 12 layers (where the hidden layer is 768 dimensions) as an example, this method 200 may include:

[0111] S210, set two vector matrices for learning part-of-speech (pos) features in each layer of the bert model.

[0112] Specifically, assume that the two learnable vector matrices are denoted as prompt_pos_K and prompt_pos_V respectively. Among them, the shape of prompt_pos_K is 15 * 768, and the shape of prompt_pos_V is 15 * 768. prompt_pos_K and prompt_pos_V can be randomly generated vector matrices, and prompt_pos_K and prompt_pos_V can be the same or different when set.

[0113] Optionally, the settings of the learnable vector matrices between different layers are independent.

[0114] Figure 3 It is a schematic diagram of a pre-trained language model architecture provided by an embodiment of the present application. Next, in combination with Figure 3 prompt_pos_K and prompt_pos_V will be introduced. Taking the training of the part-of-speech (pos) features of prompt_pos_K and prompt_pos_V as an example. As Figure 3 shown, prompt_pos_K and prompt_pos_V are set in each layer of the bert model. prompt_pos_K can be regarded as a vector matrix composed of h0, h1,..., h i (i = 15), where h0 to h i can be regarded as virtual representations of the vector matrix, and the shape of each virtual representation is 1 * 768; prompt_pos_V can be regarded as a vector matrix composed of h 0’ , h 1’ ,..., h i’ (i = 15), where h 0’ to h i’ can be regarded as virtual representations of the vector matrix, and the shape of each virtual representation is 1 * 768. Assume that the input x = "Amazing!", after being processed by the embedding layer, the input x is represented by a vector, and the vectors are e([CLS]), e(Amazing), and e(!) (e is the embedding function of the model). The word vector is input into the bert model, and prompt_pos_K and prompt_pos_V are trained through pos annotation data. Since the parameters of the bert model are frozen, that is, they do not participate in the training, only prompt_pos_K and prompt_pos_V are trained.

[0115] It can be seen that the set prompt_pos_K and prompt_pos_V are equivalent to the prompts of each transformer layer.

[0116] It should be noted that this application does not limit the pre-trained language model, as long as it uses the attention mechanism.

[0117] S220, use the training samples of the part-of-speech analysis task to train the learnable vector matrices of all layers.

[0118] In this method, it mainly includes two settings:

[0119] Setting 1, the output of each layer is set in the following way:

[0120] The output of the first layer is implemented using the following formulas (1) and (2):

[0121]

[0122] Among them, q1 is the query vector matrix determined according to the input x, k1 is the key vector matrix determined according to the input x, and v1 is the value vector matrix determined according to the input x. denotes concatenation, is the vector matrix after concatenating k1 and the prompt_pos_K of the first layer. For example, if the shape of k1 is 200*768, then has a shape of 215*768, is the vector matrix after concatenating v1 and the prompt_pos_V of the first layer. For example, if the shape of v1 is 200*768, then has a shape of 215*768. Transformer-Layer() is the function of the current transformer layer, and h1 is the output after q1, and are processed by Transformer-Layer(), which is the output of the current transformer layer.

[0123] The following is an example to illustrate the meaning of "concatenation":

[0124] Suppose the text input to the bert model is "You are a good student", the k vector of a transformer layer corresponds to [CLS]You are a good student[SEP], and the form after concatenating the k vector and prompt_pos_K is like h0, h1,..., h i (i = 15)[CLS]You are a good student[SEP], with a shape of 215*768.

[0125] Since the concatenation order of the vector matrices is different, it will make the relative positions different, affecting the output result of the transformer layer. Therefore, formula (1) can also be replaced by formula (3), and formula (3) is as follows:

[0126]

[0127] Combined with the above description, the output of the nth layer is implemented using the following formulas (4), (5), and (6), where n > 1 and n is an integer:

[0128] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v (4)

[0129]

[0130] Among them, W q , W k , and W v are the linear transformation matrices corresponding to q n , k n , and v n of the current layer respectively, h n-1 represents the output of the (n - 1)th layer, q n is the vector matrix after mapping h n-1 to W q (that is, q n is the matrix product of h n-1 and W q ), k n is the vector matrix after mapping h n-1 to W k , and v n is the vector matrix after mapping h n-1 to W v . For other content, refer to formulas (1) and (2), which will not be elaborated here.

[0131] Similarly, according to the concatenation order of the vector matrices, formula (5) can also be replaced by formula (7), and formula (7) is as follows:

[0132]

[0133] It should be noted that although prompt_pos_K and prompt_pos_V in formulas (5) and (7) do not indicate the layer position, they should be understood as referring to two learnable vector matrices set for the current layer.

[0134] Combine and As the input of the current Transformer layer, it enables prompt_pos_K and prompt_pos_V to participate in the training, so that learning and updating can be carried out.

[0135] Setting Two, the loss function is set as follows:

[0136] Calculate the loss according to the output of the last layer of the model. prompt_pos_K, prompt_pos_V, and the fully connected layer of the current task are trained (it should be noted that the fully connected layers for different tasks are different), while the parameters in the BERT model are fixed. Among them, the calculation of loss is shown in the following formula (8):

[0137]

[0138] where, θ pos =(W pos , b pos , prompt_pos_K, prompt_pos_V), represents the set of model parameters associated with the pos task, c represents the category, W pos is the weight matrix of the fully connected layer, b pos is the bias vector. J1(θ pos ) is the output of the loss, is the loss, which is the learning target of pos, λ||θ pos || 2 is the regularization term of the L2 norm, which is a restriction on the parameters of the pos layer. λ is the L2-norm regularization hyperparameter and can be configured as needed.

[0139] In one example, after prompt_pos_K and prompt_pos_V are processed by layer normalization (LayerNormalization), they are respectively concatenated with the k vector matrix and v vector matrix of the current layer.

[0140] It should be noted that the above description of the vector matrix shape is for illustrative purposes, and this application does not limit it.

[0141] In method 200, the learnable vector matrix of the current layer (i.e., the learnable vector matrix for the prompt function) is concatenated with the k vector matrix and the v vector matrix of the current layer respectively, and then used as the input of the current layer, so that the learnable vector matrix of the current layer is trained. Since only the parameters of the pos layer are trained, the parameters in the bert model are frozen, reducing the training parameters. At the same time, due to the use of the training samples of the pos task, the learnable vector matrix has part-of-speech features, and the model has basic language knowledge, which is beneficial to the learning of small samples in subsequent upper-layer NLP tasks and improves the effect of prompt learning.

[0142] In a possible implementation, the pos data in method 200 can also be replaced with chunk data, or data of other NLP basic tasks, so that the learnable vector matrix can learn chunk features or features of other basic tasks. The following takes the training samples of the chunk task as an example to illustrate method 300 in combination with method 200. Method 300 may include:

[0143] S310, set two vector matrices for learning chunk analysis features in each layer of the bert model.

[0144] Specifically, assume that the two learnable vector matrices are denoted as prompt_chunk_K and prompt_chunk_V respectively. Among them, the shape of prompt_chunk_K is 18*768, and the shape of prompt_chunk_V is 18*768. prompt_chunk_K and prompt_chunk_V can be randomly generated vector matrices, and prompt_chunk_K and prompt_chunk_V can be the same or different when set.

[0145] For the schematic diagram of the pre-trained language model architecture under this task, see Figure 3 , which will not be elaborated here.

[0146] S320, use the training samples of the chunk analysis task to train the learnable vector matrices of all layers.

[0147] In this method, it mainly includes two settings:

[0148] Setting 1, the output of each layer is set in the following way:

[0149] The output of the first layer is implemented by the following formulas (9) and (10):

[0150]

[0151] Since the splicing order of the vector matrices is different, the relative positions will be different, affecting the output results of the transformer layer. Therefore, formula (9) can also be replaced by formula (11), as shown in formula (11) below:

[0152]

[0153] Combined with the above description, the output of the nth layer is implemented using the following formulas (12), (13), and (14), where n > 1 and n is an integer:

[0154] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v (12)

[0155]

[0156] Similarly, according to the splicing order of the vector matrices, formula (13) can also be replaced by formula (15), as shown in formula (15) below:

[0157]

[0158] Setting two, the loss function is set as follows:

[0159] Calculate the loss according to the output of the last layer of the model. The bert parameters of the 12th layer are fixed, and prompt_chunk_K, prompt_chunk_V, and the fully connected layer of the current task are trained (it should be noted that the fully connected layers for different tasks are different). Among them, the calculation of loss is shown in the following formula (16):

[0160]

[0161] Among them, θ chk = (W chk , b chk , prompt_chunk_K, prompt_chunk_V). For the other contents of the above formulas, refer to the descriptions in the above formulas (1) to (8), which will not be elaborated here. Generally speaking, is the learning objective of the chunk, +λ||θ chk || 2 is the parameter restriction on the chunk layer.

[0162] In one example, after processing prompt_chunk_K and prompt_chunk_V through layer normalization, they are respectively concatenated with the k vector matrix and v vector matrix of the current layer.

[0163] Figure 4 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of the present application. As shown in method 400 in Figure 4 In a possible implementation manner, based on the learnable vector matrix trained with pos data in method 200, a learnable vector matrix for learning chunks is embedded. This method 400 includes:

[0164] The content of S410 is the same as that of S210, and the content of S420 is the same as that of S220. That is, after performing S210 and S220, S430 and S440 are performed.

[0165] S430, set two vector matrices for learning chunk features in each layer of the model.

[0166] Specifically, the two learnable vector matrices are respectively denoted as prompt_chunk_K and prompt_chunk_V. Among them, the shape of prompt_chunk_K is 18 * 768, and the shape of prompt_chunk_V is 18 * 768. Prompt_chunk_K and prompt_chunk_V can be randomly generated vector matrices, and prompt_chunk_K and prompt_chunk_V during setting can be the same or different.

[0167] Optionally, the setting of the learnable vector matrices between different layers is independent.

[0168] Since the model trained in method 200 already contains prompt_pos_K and prompt_pos_V with dimensions of 15 * 768 respectively, embedding prompt_chunk_K and prompt_chunk_V on this basis is equivalent to migrating the parameters of the trained model.

[0169] Figure 5 It is a schematic diagram of the pre-trained language model architecture provided by another embodiment of the present application. The following combines Figure 5 to introduce prompt_chunk_K and prompt_chunk_V. As shown in Figure 5 , prompt_chunk_K can be regarded as composed of h i+1 , h i+2 , …, h j(j = i + 18) forms a vector matrix, and prompt_chunk_V can be regarded as a vector matrix composed of h i+1’ , h i+2’ , …, h j’ . In this method, only prompt_chunk_K and prompt_chunk_V are trained. For other content, please refer to Figure 3 , which will not be elaborated here.

[0170] Define the vector matrix formed by combining prompt_chunk_K and prompt_pos_K as prompt_K, that is, As Figure 5 shown, the shape of prompt_K is 33 * 768 (or the dimension is [33, 768]). Due to different splicing orders, the relative positions of the vector matrix will be different, which will affect the output of the transformer layer. Therefore, another calculation method of prompt_K is Similarly, define the vector matrix formed by combining prompt_chunk_V and prompt_pos_V as prompt_V, and the calculation methods of prompt_V are or

[0171] In an example, after prompt_chunk_K and prompt_chunk_V are processed by layer normalization, they are respectively spliced with prompt_pos_K and prompt_pos_V.

[0172] In an example, after prompt_pos_K and prompt_pos_V are processed by layer normalization, they are respectively spliced with prompt_chunk_K and prompt_chunk_V.

[0173] S440. Use the training samples of the chunk task to train the learnable vector matrices of all layers.

[0174] Note that the learnable vector matrices at this time refer to the spliced prompt_K and prompt_V.

[0175] In this method, there are mainly two settings:

[0176] Setting 1. The output of each layer is set as follows:

[0177] The output of the first layer is implemented by the following formulas (17) and (18):

[0178]

[0179] It should be noted that, different from when only training the pos task or the chunk task and at this time is the concatenation of k1 and prompt_K, is the concatenation of V1 and prompt_V. For other content, please refer to the above description of the relevant formula, which will not be elaborated here.

[0180] Since the concatenation order of the vector matrices is different, the relative positions will be different, affecting the output results of the transformer layer. Therefore, formula (17) can also be replaced by formula (19), and formula (19) is as follows:

[0181]

[0182] Combined with the above description, the output of the nth layer is implemented using the following formulas (20), (21) and (22), where n > 1 and n is an integer:

[0183] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v (20)

[0184]

[0185] Similarly, according to the concatenation order of the vector matrices, formula (21) can also be replaced by formula (23), and formula (23) is as follows:

[0186]

[0187] For the content of the above formulas, please refer to the description of the relevant formulas in method 200, which will not be elaborated here.

[0188] Taking and as the input of the current transformer layer enables prompt_chunk_K and prompt_chunk_V to participate in the training, so that learning and updating can be carried out.

[0189] Setting two, the loss function is set as follows:

[0190] Calculate the loss based on the output of the last layer of the model, and train prompt_chunk_K, prompt_chunk_V, and the fully connected layer of the current task. Among them, the calculation of the loss is shown in formula (24):

[0191]

[0192] where θ chk =(W chk , b chk , prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V), W chk is the weight matrix of the fully connected layer of the current task, b chk is the bias vector of the current task, θ pos represents the parameters of the pos layer of the current task, θ' pos represents the parameters of the pos layer after being trained in the previous task (the pos task in method 200 in this embodiment), δ||θ pos -θ' pos || 2 is the continuous regularization term, and the magnitudes of the λ and δ parameters can be configured as needed. Generally speaking, the first part in formula (24) is the learning objective of the current chunk task, the second part is the parameter restriction on the chunk layer of the current task, that is, prompt_chunk_K, prompt_chunk_V, and the last fully connected layer, and the third part is the restriction on the parameters of the pos layer to prevent the parameters of the pos from changing too much, which is a means to prevent forgetting and retains the training results of the pos task. For the content of other parts of formula (24), please refer to the description of the relevant formulas in method 200 and will not be elaborated here. It can be seen that prompt_pos_K, prompt_pos_V, prompt_chunk_K, and prompt_chunk_V are all involved in the training of the chunk task, but in order to prevent forgetting, the changes in the parameters of prompt_pos_K and prompt_pos_V are restricted.

[0193] In method 400, a learnable vector matrix prompt_chunk_K and prompt_chunk_V for learning the chunk task are further embedded on the learnable vector matrix that already has part-of-speech features. Then, prompt_chunk_K is concatenated with prompt_pos_K trained in the previous pos task, and prompt_chunk_V is concatenated with prompt_pos_V trained in the previous pos task. After the new learnable vector matrices of the current layer (i.e., prompt_K and prompt_V) are respectively concatenated with the k vector matrix and v vector matrix of the current layer, they are used as the input of the current transformer layer. Finally, the loss function is used to train and update prompt_chunk_K and prompt_chunk_V. Since only the parameters of the learnable vector matrix are trained, the parameters in the bert model are frozen, greatly reducing the training parameters. At the same time, the learnable vector matrix has part-of-speech features and chunk features, and the model has basic language knowledge, which is beneficial to the learning of few-shot samples in subsequent upper-layer NLP tasks and improves the effect of prompt learning.

[0194] In one example, a learnable vector matrix prompt_pos_K and prompt_pos_V for learning the pos task can also be further embedded on the prompt with chunk features, that is, the pos task is learned on the model trained by method 300. Then, prompt_pos_K is concatenated with prompt_chunk_K trained in the previous chunk task, and prompt_pos_V is concatenated with prompt_chunk_V trained in the previous chunk task. After the new prompt of the current layer is respectively concatenated with the k vector matrix and v vector matrix of the current layer, they are used as the input of the current transformer layer. Finally, the loss function is used to train and update prompt_pos_K and prompt_pos_V. The specific settings can be analogized according to method 400 and will not be elaborated here.

[0195] Figure 6 It is a schematic flowchart of another training method for a natural language processing model provided by an embodiment of the present application. As shown in method 500 in Figure 6 In a possible implementation manner, on the basis of the learnable vector matrix trained by using chunk data in method 400, a learnable vector matrix for learning dependency parsing (dep) features is further embedded. This method 400 includes:

[0196] The contents of S510 to S540 are respectively the contents of S410 to S440.

[0197] S550, set two vector matrices for learning dependency parsing (dep) features at each layer in the model.

[0198] Specifically, these two learnable vector matrices are denoted as prompt_dp_K and prompt_dp_V respectively. Among them, the shape of prompt_dp_K is 20 * 768, and the shape of prompt_chunk_V is 20 * 768. prompt_dp_K and prompt_dp_V can be randomly generated vector matrices, and prompt_dp_K and prompt_dp_V can be the same or different when set.

[0199] Optionally, the setting of the learnable vector matrices between different layers is independent.

[0200] Since the model trained in method 400 already contains prompt_pos_K, prompt_pos_V, prompt_chunk_K, and prompt_chunk_V, with dimensions of 15 * 768, 15 * 768, 18 * 768, and 18 * 768 respectively, embedding prompt_dp_K and prompt_dp_V on this basis is equivalent to migrating the parameters of the trained model.

[0201] For the schematic diagram of the pre-trained language model architecture in method 500, see Figure 3 and Figure 5 , which will not be elaborated here.

[0202] Based on method 400, concatenate prompt_dp_K with the prompt_K trained in the previous chunk task to generate a new prompt_K, whose shape is 53 * 768, that is, Or Similarly, concatenate prompt_dp_V with the prompt_V trained in the previous chunk task to generate a new prompt_V, whose shape is 53 * 768, that is, Or

[0203] In one example, perform layer normalization on prompt_dp_K and prompt_dp_V and then perform concatenation processing.

[0204] In one example, perform layer normalization on the original prompt_K and prompt_V and then perform concatenation processing.

[0205] S560, use the training samples of the dependency parsing task to train the learnable vector matrices of all layers.

[0206] In this method, it mainly includes two settings:

[0207] Setting 1, the output of each layer is set in the following way:

[0208] The output of the first layer is implemented using the following formulas (25) and (26):

[0209]

[0210] Since the concatenation order of the vector matrices is different, the relative positions will be different, affecting the output result of the transformer layer. Therefore, formula (25) can also be replaced by formula (27), and formula (27) is as follows:

[0211]

[0212] Combined with the above description, the output of the nth layer is implemented using the following formulas (28), (29), and (30), where n > 1 and n is an integer:

[0213] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v (28)

[0214]

[0215] Similarly, according to the concatenation order of the vector matrices, formula (29) can also be replaced by formula (31), and formula (31) is as follows:

[0216]

[0217] It can be seen that due to the proposal of prompt_K and prompt_V, the expressions of formulas (25) to (31) are the same as those of formulas (17) to (23), while the data for each item is different in different tasks. For example, prompt_K in formulas (25) to (31) includes prompt_dp_K, while prompt_K in formulas (17) to (23) does not include prompt_dp_K.

[0218] For the content of the above formulas, please refer to the description of the relevant formulas in method 400, which will not be elaborated here.

[0219] Put and As the input of the current Transformer layer, it enables prompt_dp_K and prompt_dp_V to participate in the training, so that they can be learned and updated.

[0220] Setting Two, the loss function is set as follows:

[0221] Calculate the loss based on the output of the last layer of the model, and train prompt_dp_K, prompt_dp_V and the fully connected layer of the current task. Among them, the calculation of the loss is shown in Equation (32):

[0222]

[0223] where θ dep =(W dep , b dep , prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V, prompt_dp_K, prompt_dp_V), W dep is the weight matrix of the fully connected layer of the current task, b dep is the bias vector of the current task, θ chk represents the parameters of the chunk layer of the current task, and θ' chk represents the parameters of the chunk layer after being trained for the previous task (in this embodiment, it is the chunk task in Method 400). For other details, please refer to the relevant descriptions in Equation (8) and Equation (16), which will not be elaborated here. Generally speaking, the first part in Equation (32) is the learning objective of the current task dep. The second part is the parameter restriction on the dep layer of the current task, that is, prompt_dep_K, prompt_dep_V and the last fully connected layer. The third part is the restriction on the parameters trained for the previous task, which prevents the parameters from changing too much and is a means to prevent forgetting. It can be seen that prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V, prompt_dp_K, prompt_dp_V all participate in the training of the dependency parsing task, but in order to prevent forgetting, the changes in the parameters of prompt_pos_K, prompt_pos_V, prompt_chunk_K and prompt_chunk_V are restricted.

[0224] In method 500, on the learnable vector matrix that already has part-of-speech features and chunk features, a learnable vector matrix prompt_dp_K and prompt_dp_V for learning the dependency parsing (dep) task are embedded again. Then, prompt_dp_K is concatenated with prompt_K after being trained in the previous chunk task to generate a new prompt_K. prompt_dep_V is concatenated with prompt_V after being trained in the previous chunk task to generate a new prompt_V. After the new prompts in the current layer are respectively concatenated with the k vector matrix and v vector matrix in the current layer, they are used as the input of the current transformer layer. Finally, the loss function is used to train and update prompt_dp_K and prompt_dp_V. Since only the parameters of the learnable vector matrix are trained, the parameters in the bert model are frozen, greatly reducing the training parameters. At the same time, the learnable vector matrix has part-of-speech features, chunk features, and dependency parsing features, and the model has basic language knowledge, which is beneficial to the few-shot learning in subsequent upper-layer NLP tasks and improves the effect of prompt learning.

[0225] It can be seen that the order of the model learning tasks in method 500 is part-of-speech analysis (POS) -> chunk analysis (CHUNK) -> dependency parsing analysis (DEP). The present application does not limit the order of the learning tasks. For example, it can also be chunk analysis (CHUNK) -> part-of-speech analysis (POS) -> dependency parsing analysis (DEP). Further, based on method 500, other NLP tasks can continue to be learned, such as text relatedness and text entailment.

[0226] In a possible implementation manner, after the prompt in the model has learned the NLP basic task, the upper-layer task can continue to be trained. The following takes the classification task as an example for illustration. Figure 7 It is a schematic flowchart of another training method of the natural language processing model provided by the embodiment of the present application. Figure 7 The method 600 in it includes:

[0227] The content of S610 to S660 is the same as the content of S510 to S560.

[0228] S670, set two vector matrices for learning classification task features in each layer of the model.

[0229] Specifically, the two learnable vector matrices are denoted as prompt_cls_K and prompt_cls_V respectively. Among them, the shape of prompt_cls_K is 15 * 768, and the shape of prompt_cls_V is 15 * 768. prompt_cls_K and prompt_cls_V can be randomly generated vector matrices, and prompt_cls_K and prompt_cls_V can be the same or different when set.

[0230] Optionally, the settings of the learnable vector matrices between different layers are independent.

[0231] Since the model trained in Method 500 already contains prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V, prompt_dp_K, and prompt_dp_V, with dimensions of 15 * 768, 15 * 768, 18 * 768, 18 * 768, 20 * 768, and 20 * 768 respectively, embedding prompt_cls_K and prompt_cls_V on this basis is equivalent to migrating the parameters of the trained model.

[0232] For the schematic diagram of the pre-trained language model architecture in Method 600, see Figure 3 and Figure 5 , which will not be elaborated here.

[0233] Based on Method 500, prompt_cls_K is concatenated with the prompt_K trained in the previous dependency parsing task to generate a new prompt_K, whose shape is 68 * 768, that is, prompt_K = [prompt_K ○ prompt_cls_K], or prompt_K = [prompt_cls_K ○ prompt_K]. Similarly, prompt_cls_V is concatenated with the prompt_V trained in the previous dependency parsing task to generate a new prompt_V, whose shape is 68 * 768, that is, prompt_V = [prompt_V ○ prompt_cls_V], or prompt_V = [prompt_cls_V ○ prompt_V].

[0234] In one example, prompt_cls_K and prompt_cls_V are subjected to layer normalization processing and then concatenated.

[0235] S680. Use the training samples of the classification task to train the learnable vector matrices of all layers.

[0236] In this method, it mainly includes two settings:

[0237] Setting 1, the output of each layer is set as follows:

[0238] The output of the first layer is implemented using the following formulas (33) and (34):

[0239]

[0240] Since the concatenation order of the vector matrices is different, the relative positions will be different, affecting the output result of the transformer layer. Therefore, formula (33) can also be replaced by formula (35), and formula (35) is as follows:

[0241]

[0242] Combined with the above description, the output of the nth layer is implemented using the following formulas (36), (37), and (38), where n > 1 and n is an integer:

[0243] q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v (36)

[0244]

[0245] Similarly, according to the concatenation order of the vector matrices, formula (37) can also be replaced by formula (39), and formula (39) is as follows:

[0246]

[0247] It can be seen that due to the proposal of prompt_K and prompt_V, the expression forms of formulas (33) to (39) are the same as those of formulas (17) to (23), while the data for each item is different in different tasks. For example, prompt_K in formulas (33) to (39) includes prompt_cls_K, while prompt_K in formulas (17) to (23) does not include prompt_cls_K.

[0248] For the content of the above formulas, refer to the description of the relevant formulas in Method 400, which will not be elaborated here.

[0249] Put and As the input of the current transformer layer, it enables prompt_cls_K and prompt_cls_V to participate in the training, so that they can learn and be updated.

[0250] Setting Two, the loss function is set as follows:

[0251] Calculate the loss according to the output of the last layer of the model, and train prompt_cls_K, prompt_cls_V, and the fully connected layer of the current task. Among them, the calculation of the loss is shown in formula (40):

[0252]

[0253] Where, θ cls =(W cls , b cls , prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V, prompt_dp_K, prompt_dp_V, prompt_cls_K, prompt_cls_V), W cls is the weight matrix of the fully connected layer of the current task, and b cls is the bias vector of the current task. The first part in formula (40) is the learning objective of the current classification task, the second part is the parameter restriction for the classification task, that is, prompt_cls_K, prompt_cls_V, and the last fully connected layer, and the third part is the restriction on the parameters trained in the previous task to prevent the parameters from changing too much, which is a means to prevent forgetting. For the other content of the formula, please refer to the description of the relevant formula in method 400, which will not be elaborated here. It can be seen that prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V, prompt_dp_K, prompt_dp_V, prompt_cls_K, and prompt_cls_V all participate in the training of the classification task, but in order to prevent forgetting, the changes in the parameters of prompt_pos_K, prompt_pos_V, prompt_chunk_K, prompt_chunk_V, prompt_dp_K, and prompt_dp_V are restricted.

[0254] In method 600, on the learnable vector matrix that already has part-of-speech features, chunk features, and dependency syntactic features, a learnable vector matrix prompt_cls_K and prompt_cls_V for learning classification tasks are embedded again. Then, prompt_cls_K is concatenated with prompt_K after the previous dependency syntactic task training to generate a new prompt_K. Prompt_cls_V is concatenated with prompt_V after the previous dependency syntactic task training to generate a new prompt_V. After the new prompts of the current layer are concatenated with the k vector matrix and v vector matrix of the current layer respectively, they are used as the input of the current transformer layer. Finally, the loss function is used to train and update prompt_cls_K and prompt_cls_V. This method reduces the training parameters while enabling the model to learn classification tasks based on basic language knowledge features such as part-of-speech features, chunk features, and dependency syntactic features, improving the learning ability of few-shot samples, that is, improving the effect of prompt learning.

[0255] Based on the above training method of the natural language processing model, the present application also provides a training device for the natural language processing model. The device includes:

[0256] A first module, configured to obtain a pre-trained language model, and each layer of the pre-trained language model is a layer structure using the self-attention mechanism;

[0257] A second module, configured to obtain a first learnable vector matrix and a second learnable vector matrix for each layer. Both the first learnable vector matrix and the second learnable vector matrix are used to learn the first task in the NLP task;

[0258] A third module, configured to generate a first concatenated key vector matrix and a first concatenated value vector matrix for each layer. The first concatenated key vector matrix is obtained by concatenating the first learnable vector matrix of the current layer with the key vector of the current layer, and the first concatenated value vector matrix is obtained by concatenating the second learnable vector matrix of the current layer with the value vector of the current layer;

[0259] A fourth module, configured to input the training sample data of the first task into the pre-trained language model, calculate the output of the current layer according to the first concatenated key vector matrix, the first concatenated value vector matrix, and the query vector of the current layer of each layer, and obtain the optimal first learnable vector matrix and second learnable vector matrix for each layer according to the output of the last layer and the loss function of the first task.

[0260] For other implementation manners of this device, refer to the descriptions in methods 100 to 600, which will not be elaborated here.

[0261] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this text, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0262] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A training method for a natural language processing model, characterized in that, Including: Obtain a pre-trained language model, where each layer of the pre-trained language model is a layer structure using a self-attention mechanism; Obtain a first learnable vector matrix and a second learnable vector matrix for each layer. Both the first learnable vector matrix and the second learnable vector matrix are used to learn a first task in NLP tasks. The first learnable vector matrix for each layer is obtained by concatenating the third learnable vector matrix of the current layer with the learnable vector matrix obtained from all NLP tasks participated in before the current layer. The second learnable vector matrix for each layer is obtained by concatenating the fourth learnable vector matrix of the current layer with the learnable vector matrix obtained from all NLP tasks participated in before the current layer. Among them, both the third learnable vector matrix and the fourth learnable vector matrix are vector matrices set to learn the first task. The NLP tasks include NLP basic tasks and NLP upper-layer tasks. The NLP basic tasks include part-of-speech analysis tasks, chunk analysis tasks, and dependency syntax analysis tasks. The NLP upper-layer tasks include named entity recognition tasks, text semantics-related tasks, text entailment tasks, and classification tasks; Generate a first concatenated key vector matrix and a first concatenated value vector matrix for each layer. The first concatenated key vector matrix is obtained by concatenating the first learnable vector matrix of the current layer with the key vector of the current layer. The first concatenated value vector matrix is obtained by concatenating the second learnable vector matrix of the current layer with the value vector of the current layer; Input the training sample data of the first task into the pre-trained language model, calculate the output of the current layer according to the first concatenated key vector matrix, the first concatenated value vector matrix, and the query vector of the current layer for each layer, and obtain the optimal first learnable vector matrix and second learnable vector matrix for each layer according to the output of the last layer and the loss function of the first task.

2. The method according to claim 1, wherein After obtaining the optimal first learnable vector matrix and second learnable vector matrix for each layer, the method further includes: Set a fifth learnable vector matrix and a sixth learnable vector matrix for each layer. Both the fifth learnable vector matrix and the sixth learnable vector matrix are used to learn a second task in the NLP tasks, and the first task is different from the second task; Generate a first concatenated learnable vector matrix and a second concatenated learnable vector matrix for each layer. The first concatenated learnable vector matrix is obtained by concatenating the fifth learnable vector matrix of the current layer with the first learnable vector matrix of the current layer. The second concatenated learnable vector matrix is obtained by concatenating the sixth learnable vector matrix of the current layer with the second learnable vector matrix of the current layer; Generate a second concatenated key vector matrix and a second concatenated value vector matrix for each layer. The second concatenated key vector matrix is obtained by concatenating the first concatenated learnable vector matrix of the current layer with the key vector of the current layer. The second concatenated value vector matrix is obtained by concatenating the second concatenated learnable vector matrix of the current layer with the value vector of the current layer; Input the training sample data of the second task into the pre-trained language model, calculate the output of the current layer according to the second concatenated key vector matrix, the second concatenated value vector matrix of each layer and the query vector of the current layer, obtain the optimal first concatenated learnable vector matrix and the second concatenated learnable vector matrix of each layer according to the output of the last layer and the loss function of the second task, and limit the parameter changes of the first learnable vector matrix and the parameter changes of the second learnable vector matrix.

3. The method according to claim 1, wherein The output h1 of the first layer of the pre-trained language model is implemented by the following formula: or, Among them, q1 is the query vector matrix input to the model, k1 is the key vector matrix input to the model, and v1 is the value vector matrix input to the model. denotes concatenation, prompt_K is the learnable vector matrix corresponding to k1, and prompt_V is the learnable vector matrix corresponding to v1. is the vector matrix after concatenating k1 and prompt_K. is the vector matrix after concatenating v1 and prompt_V, f is the function used in the current layer, and h1 is the output of the current layer. The output h of the n-th layer of the pre-trained language model n is implemented using the following formula, where n is an integer and n > 1: q n , k n , v n = h n-1 W q , h n-1 W k , h n-1 W v ; or, Among them, h n-1 is the output of the (n - 1)-th layer, W q is the linear transformation matrix of the query vector of the current layer, W k is the linear transformation matrix of the key vector of the current layer, W v is the linear transformation matrix of the value vector of the current layer, q n is the query vector of the current layer, k n is the key vector of the current layer, v n is the value vector of the current layer, denotes concatenation, prompt_K is the learnable vector matrix corresponding to k n and prompt_V is the learnable vector matrix corresponding to v n respectively. is the vector matrix after concatenating k n with prompt_K, is the vector matrix after concatenating v n with prompt_V, f is the function used in the current layer, and h n is the output of the current layer.

4. The method according to claim 3, wherein The prompt_K and prompt_V are implemented by the following formula: or, or, Among them, prompt c1 _K represents a learnable vector matrix corresponding to the key vector of the current layer that is set to learn the current NLP task, prompt c2 _K represents a learnable vector matrix corresponding to the key vector of the current layer obtained from the training of all NLP tasks participated in before the current layer, prompt c1 _V represents a learnable vector matrix corresponding to the value vector of the current layer that is set to learn the current NLP task, prompt c2 _V represents a learnable vector matrix corresponding to the value vector of the current layer obtained from the training of all NLP tasks participated in before the current layer.

5. The method according to claim 1, wherein When the pre-trained language model is learning the first NLP task, the loss function includes the sum of the loss part of the current NLP task and the part for restricting the parameters of the current NLP task; or, When the pre-trained language model is learning a non-first NLP task, the loss function includes the sum of the loss part of the current NLP task, the part for restricting the parameters of the current NLP task and the part for restricting the parameters that have been trained for all previous NLP tasks participated in.

6. The method according to claim 1, wherein All NLP tasks participated in before the current layer are NLP basic tasks, and the first task is an NLP upper-layer task.

7. The method according to claim 2, wherein Before generating the first concatenated learnable vector matrix and the second concatenated learnable vector matrix of each layer, the method further includes: Performing layer normalization on the first learnable vector matrix and the second learnable vector matrix, and / or, Performing layer normalization on the fifth learnable vector matrix and the sixth learnable vector matrix.

8. A training device for a natural language processing model, characterized in that, Including: A first module for obtaining a pre-trained language model, where each layer of the pre-trained language model is a layer structure using the self-attention mechanism; A second module for obtaining the first learnable vector matrix and the second learnable vector matrix of each layer. Both the first learnable vector matrix and the second learnable vector matrix are used to learn the first task in the NLP task. The first learnable vector matrix of each layer is obtained by concatenating the third learnable vector matrix of the current layer and the learnable vector matrix trained by all previous NLP tasks participated in by the current layer. The second learnable vector matrix of each layer is obtained by concatenating the fourth learnable vector matrix of the current layer and the learnable vector matrix trained by all previous NLP tasks participated in by the current layer. Among them, both the third learnable vector matrix and the fourth learnable vector matrix are vector matrices set to learn the first task. The NLP task includes NLP basic tasks and NLP upper-layer tasks. The NLP basic tasks include part-of-speech analysis tasks, chunk analysis tasks, and dependency parsing tasks. The NLP upper-layer tasks include named entity recognition tasks, text semantics-related tasks, text entailment tasks, and classification tasks; The third module is used to generate the first concatenated key vector matrix and the first concatenated value vector matrix for each layer. The first concatenated key vector matrix is obtained by concatenating the first learnable vector matrix of the current layer with the key vector of the current layer. The first concatenated value vector matrix is obtained by concatenating the second learnable vector matrix of the current layer with the value vector of the current layer; The fourth module is used to input the training sample data of the first task into the pre-trained language model, calculate the output of the current layer according to the first concatenated key vector matrix, the first concatenated value vector matrix of each layer and the query vector of the current layer, and obtain the optimal first learnable vector matrix and the second learnable vector matrix of each layer according to the output of the last layer and the loss function of the first task.

Citation Information

Patent Citations

  • Method for quickly extracting fault information in power grid equipment fault report

    CN112632972A

  • Rapid picture classification method based on prompt learning

    CN114090780A