A text generation method, device, and system in a low-resource scenario
Through consistent semi-supervised learning and adapter fine-tuning pre-training parameter freezing methods, the problem of text generation model training in low-resource scenarios is solved, and efficient text generation performance and low training overhead are achieved.
Patent Information
- Application Number
- CN202210308980.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The prior art is difficult to effectively train text generation models in low-resource scenarios, especially in the problems of few labeling samples and high training overhead.
The consistent semi-supervised learning method is adopted to reduce the overhead of model training by reducing the number of labeled samples and using the adapter fine-tuned pre-trained parameter freezing method.
In low-resource scenarios, the dependence on massive amounts of manual annotation data is significantly reduced, the good text generation performance is maintained, and the overhead of model training is reduced.
Smart Images

Figure CN114611472B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and mainly relates to a text generation method, device and system in a low-resource scenario. Background Art
[0002] With the development of Internet technology, a large amount of text information on the World Wide Web has grown rapidly. In the existing scenario of information explosion, for reading news and other content, there is an urgent need for a method that can automatically condense and generate simple text, such as automatically generating a title, automatically generating a summary of a news, or automatically generating a timeline narrative document of a news. And with the popularization of mobile Internet devices, the screens of mobile devices also require the content and display of news to be presented in a summary form. The automatic text generation method is the only way to solve the extraction and generation of core content from a large amount of information such as a large number of news.
[0003] The traditional mode to implement this method is to use a large amount of manually annotated data to train a text generation model, and let the trained model automatically generate text for new news data. However, in many real scenarios, annotating a large amount of target text data requires a lot of manpower and material resources, which is time-consuming and inefficient. For example, the annotation scale of the LCSTS data for generating Chinese news titles reaches more than 2.1 million, and the annotation scale of the THUCNews data for Chinese news summaries reaches more than 830,000. Existing methods do not discuss how to train a text generation model in a low-resource scenario with few annotated samples. Secondly, the existing pre-trained models perform excellently in text generation tasks, but due to the large number of model parameters of the pre-trained models themselves, it brings a large training cost (such as large GPU video memory cost and long model training time). How to reduce the training cost of the model is also an urgent problem to be solved in the lightweight aspect. The present invention relates to a text generation method, device and system in a low-resource scenario. It is suitable for extractive text generation, such as extracting keywords for generation, and generative text generation, such as generating the target text word by word. The present invention uses consistent semi-supervised learning to solve the scenario of few annotated samples, and can reduce the number of annotated samples in the LCSTS Chinese news title generation dataset of 2.1 million to 10%, and ensure that the model performance under 10% of the labeled data is comparable to the text generation performance of about 50% of the labeled data with a large amount of unlabeled data. The present invention also uses the pre-trained parameter freezing method of adapter fine-tuning. For example, freezing the pre-trained BERT model can reduce about 110M parameters from participating in the gradient backpropagation calculation, reducing the training cost of the text generation model. Summary of the Invention
[0004] In response to the requirements of the current text generation method in low-resource scenarios, the present invention conducts in-depth research and practice to achieve automatic text generation in scenarios with few annotations, greatly reducing the dependence of the text generation method on a large amount of manually annotated data and maintaining good text generation performance.
[0005] To achieve the above object, the present invention adopts the following technical solutions.
[0006] It includes three steps:
[0007] Step 1: Input a small amount of supervised training samples, the corresponding embedding vectors of the input documents, into the supervised network. At the same time, input a large amount of unsupervised training samples, that is, a large amount of source document data without manual annotation obtained from the open corpus, into the unsupervised network. Duplicate the unsupervised documents twice, and then perform dropout on their corresponding embedding vectors respectively to obtain two sets of embedding vectors.
[0008] Step 2: Parallelly integrate small neural modules (Adapters) of the adapter into the large pre-trained text generation network (Pre-trained model) to form an adapter fine-tuning pre-training learning component. In the supervised network T and the two consistent unsupervised networks A and B, the adapter fine-tuning pre-training learning component with the same network architecture is adopted. In the adapter fine-tuning pre-training learning component, the additional small adapter neural module participates in the model training, while the original large pre-trained text generation module needs to keep the parameters frozen. Specifically,
[0009] Among them, supervised training is carried out in the supervised network T. The input during the training process is the supervised source document-target text pair (x * , y * ). Unsupervised consistency learning is carried out in the unsupervised networks A and B. The input during the training process is x, and the outputs of A and B are their predicted labels. The consistency learning is to make their predicted labels consistent.
[0010] Among them, in the pre-training learning component based on adapter fine-tuning, the input of the network is the embedding vector: H input , and the output is For the supervised network T, H input is the embedding vector of x * . For the unsupervised network A, H input is the embedding vector of x corresponding to the dropout. For the unsupervised network B, H input is the embedding vector of the duplicated x after another dropout. H inputIt will be input into a large pre-trained text generation network and a small adapter network simultaneously. During the training process, this component keeps the parameters of the large pre-trained text generation model part frozen, that is, the parameters do not participate in the parameter learning and update process of backpropagation. Only the parameters of the small Adapter network participate in the update calculation, so as to achieve the purpose of reducing the model training cost.
[0011] Among them, in the pre-trained learning component based on adapter fine-tuning, the adopted adapter small neural network (Adapter), the updated parameter of its forward part is W in , through a non-linear activation function Relu function to non-linearly optimize the embedded vector, and then input it into the latter part of the adapter, and use its updated parameter W out to train the adapter, and the output representation vector of the adapter is
[0012]
[0013] Among them, in the pre-trained learning component based on adapter fine-tuning, combine the output of the large pre-trained text generation model and add it linearly to obtain the final output representation vector of the adapter fine-tuning pre-trained learning component
[0014]
[0015] Step 3: Based on the consistency learning of the unsupervised network and combined with the supervised learning of the supervised network, train and optimize the text generation model.
[0016] Among them, the unsupervised networks A and B perform the above-mentioned unsupervised consistency learning to make the prediction targets of the two unsupervised pre-trained text generation neural networks consistent. The unsupervised loss function is:
[0017]
[0018] Among them, S A and S B respectively represent the specific unsupervised network A and unsupervised network B, which are a pair of twin networks. In extractive text generation, it is BERT parallel integrated with Adapter, and in generative text generation, it is BART parallel integrated with Adapter. X u is the input unsupervised text generation dataset, and represent the input values after enhanced data augmentation. In the present invention, they are two groups of different embedded vector representations obtained after dropout respectively;
[0019] Meanwhile, jointly optimize the supervised learning of the supervised network to train the supervised text generation model, and the supervised loss function is;
[0020]
[0021] Among them, T(x * ) represents the supervised network, which is the twin network of the S A and S B . In extractive text generation, it is BERT parallel integrated with Adapter; in generative text generation, it is BART parallel integrated with Adapter. X l is the input text generation dataset with manual annotations, and x * and y * represent the source document and its corresponding manually annotated generated target text respectively:
[0022] Finally, combine the consistency learning of the unsupervised network with the supervised learning of the supervised network to obtain the final loss function l final , which is used for the training and optimization of the model:
[0023] l final (θ, X) = λl unsup (θ, X u ) + l sup (θ, X l ), X = X u + X l
[0024]
[0025]
[0026]
[0027] Among them, λ is a hyperparameter, representing the importance of the training part of unsupervised consistency learning in the entire model training process. θ A is the parameter of the small neural module of the adapter, θ B is the parameter of the large pre-trained text generation model (BERT or BART), θ is the parameter of the entire model, and m is the number of epochs; it can be seen that during the gradient backpropagation calculation, θ B is not updated in each epoch that is, it does not participate in the gradient backpropagation calculation. And θ A will update the learning parameters in each epoch during the training process;
[0028] Finally, use the optimized model for text generation prediction.
[0029] A text generation device in a low-resource scenario, comprising:
[0030] A source document input module, configured to input a small amount of labeled source documents and target texts for text generation, as well as a large amount of unlabeled source documents;
[0031] A text generation module in a low-resource scenario, applying the text generation method in the low-resource scenario;
[0032] A target output module, configured to output the automatically generated target text through an interface program.
[0033] A text generation system in a low-resource scenario, the system comprising at least one server, and a text generation device in a low-resource scenario connected to the server. When the server executes the process of generating a target text, the above-mentioned text generation method in the low-resource scenario is executed through the device.
[0034] The advantages of the present invention over the prior art are as follows:
[0035] 1. The present invention proposes a set of text generation methods for low-resource scenarios, which uses consistency learning to improve the robustness of neural networks for unsupervised learning under unlabeled data, and further improves the text generation prediction performance of the overall network under few labeled samples. It greatly reduces the dependence of text automatic generation methods on a large amount of manually labeled data. For example, for the LCSTS Chinese news title generation dataset with 2.1 million samples, the number of labeled samples is reduced to 10%, and the model performance under 10% labeled data and a large amount of unlabeled data is made comparable to the text generation performance of about 50% labeled data.
[0036] 2. The present invention uses a pre-training method based on adapters to freeze the parameters of a large pre-trained text generation neural network, and updates a small adapter neural network module, thus alleviating the overhead of model training. For example, when using BERT-base as the model basic framework for text generation, which requires 110M parameters, when these parameters are frozen and do not participate in the gradient backpropagation calculation, the model calculation efficiency can be greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is the overall flowchart (model framework diagram) of the present invention; DETAILED DESCRIPTION
[0038] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0039] The present invention proposes a method for generating generated text in a low-resource scenario. In response to the current scenario requirement of text generation methods with few labeled samples, the present invention conducts in-depth research and practice to achieve automatic target text generation for source texts in scenarios with few labeled samples, greatly reducing the dependence of text generation methods on a large amount of manually labeled data. Secondly, in response to the problem of high training overhead of current pre-trained text generation models, an adapter pre-training method is adopted to reduce the training overhead. The specific technical solutions include:
[0040] Step 1: Input a small amount of supervised training samples into the supervised network, that is, the embedding vectors of the corresponding source texts and the generated target text labels. At the same time, input a large amount of unsupervised training samples into the unsupervised network, duplicate the unlabeled documents twice, and then perform dropout on their corresponding embeddings respectively to obtain two sets of embedding vectors;
[0041] (1) Extract a small amount of text generation data with manual annotations from the labeled corpus, that is, a small amount of source documents and the corresponding manually annotated target texts;
[0042] (2) Obtain a large amount of unsupervised source document data from the open corpus, that is, a large amount of source documents without manual annotations;
[0043] (3) Duplicate the same unlabeled data twice, and then perform dropout respectively to obtain two sets of embedding vector representations;
[0044] Step 2: Parallelly integrate small neural modules (Adapters) of adapters into a large pre-trained text generation network (Pre-trained model) to form an adapter fine-tuning pre-training learning component. In the supervised network T, and the two consistent unsupervised networks A and B, an adapter fine-tuning pre-training learning component with the same network architecture is adopted. In the adapter fine-tuning pre-training learning component, the added small adapter neural module participates in model training, while the original large pre-trained text generation module needs to keep the parameters frozen. Specifically,
[0045] (1) The input is a supervised source document-target text pair (x * , y *) To the supervised network, the unsupervised document x is input into the unsupervised networks A and B. The outputs of A and B are their predicted labels, and consistency learning is to make their predicted labels consistent.
[0046] (2) In the pre-training learning component based on adapter fine-tuning, the input embedding vector is: H input , and the output representation vector is For the supervised network T, H input is the embedding vector of x * For the unsupervised network A, H input is the embedding vector of x corresponding to dropout. For the unsupervised network B, H input is the embedding vector of x after replication and another dropout. H input will be input into the large pre-trained text generation network and the small adapter network at the same time. During the training process, the parameters of the large pre-trained text generation model part in this component are frozen, that is, the parameters do not participate in the parameter learning and update process of backpropagation. Only the parameters of the small Adapter network participate in the update calculation, so as to achieve the purpose of reducing the model training cost.
[0047] First, in the pre-training learning component based on adapter fine-tuning, the adopted adapter small neural network (Adapter) has the updated parameter W in for its forward part. The embedding vector is non-linearly optimized through a non-linear activation function, the Relu function, and then input into the latter part of the adapter. The adapter is trained using its updated parameter W out , and the output representation vector of the adapter is Formula:
[0048] Second, in the pre-training learning component based on adapter fine-tuning, the output representation vector of the large pre-trained text generation model is linearly added to it to obtain the final output representation vector of the adapter fine-tuning pre-training learning component.
[0049] Step three, based on the consistency learning of the unsupervised network, and combined with the supervised learning of the supervised network, the text generation model is trained and optimized.
[0050] (1) For the unsupervised networks A and B, perform the above-mentioned unsupervised consistency learning to make the prediction targets of the two unsupervised pre-trained text generation neural networks consistent. The optimized loss function is:
[0051] Formula:
[0052] Among them, SA and S B respectively represent the specific unsupervised network A and unsupervised network B, which are a pair of twin networks. In extractive text generation, they are BERT parallel integrated with Adapters, and in generative text generation, they are BART parallel integrated with Adapters. X u is the input unsupervised text generation dataset, and represent the input values after data augmentation. In the present invention, they are two different sets of embedded vector representations obtained after dropout respectively;
[0053] (2) Jointly optimize the supervised learning of the supervised network and train the supervised text generation model. The optimized loss function is:
[0054] Formula:
[0055] where, T(x * ) represents the supervised network, which is the twin network of the unsupervised network S A and S B In extractive text generation, it is BERT parallel integrated with Adapters, and in generative text generation, it is BART parallel integrated with Adapters. X l is the input source document dataset with manual annotations, where x * and y * respectively represent the source document and its corresponding manually annotated target document:
[0056] (3) Combine the consistency learning of the unsupervised network and the supervised learning of the supervised network to obtain the final loss function l final , which is used for the training and optimization of the model:
[0057] Formula: l final (θ, X) = λl unsup (θ, X u ) + l sup (θ, X l ), X = X u + X l
[0058]
[0059]
[0060]
[0061] where, λ is a hyperparameter, representing the importance of the training part of unsupervised consistency learning in the entire model training process, θ AThe parameters of the small neural module (Adapter) for the adapter, θ B The parameters of the large pre-trained text generation module (BERT or BART), θ are the parameters of the entire model, and m is the number of epochs; it can be seen that during the gradient backpropagation calculation, θ B Does not participate in the update process during each epoch, that is, does not participate in the gradient backpropagation calculation. And θ A Will update the learning parameters in each epoch during the training process.
[0062] (4) Finally, use the optimized model for text generation prediction.
Claims
1. A text generation method in a low-resource scenario, characterized in that, it includes three steps: Step 1: Input a small amount of supervised training samples into the supervised network, and corresponding embedding vectors of the small-scale input training sample documents. At the same time, input a large amount of unsupervised training samples into the unsupervised network, that is, a large amount of source document data obtained from the open corpus without manual annotation. Duplicate the unsupervised documents twice, and then perform dropout on their corresponding embedding vectors respectively to obtain two sets of embedding vectors; Step 2: Parallelly integrate a small neural module of an adapter into the large pre-trained text generation network to form an adapter fine-tuning pre-trained learning component. In the supervised network T, and the two consistent unsupervised networks A and B, use the adapter fine-tuning pre-trained learning component with the same network architecture. In the adapter fine-tuning pre-trained learning component, the additional small adapter neural module participates in model training, while the original large pre-trained text generation module needs to keep its parameters frozen. Specifically, Among them, supervised training is carried out in the supervised network T, and the input of the training process is the supervised source document-target text pair (x * , y * ). Unsupervised consistency learning is carried out in the unsupervised network A and the unsupervised network B. The input of the training process is x. The unsupervised network A and the unsupervised network B output their predicted labels, and the consistency learning is to make their predicted labels consistent; Among them, in the pre-training learning component based on adapter fine-tuning, the input of the network is the embedding vector represented as: H input , and the output representation vector is: For the supervised network T, H input is the embedding vector of x * , for the unsupervised network A, H input is the embedding vector of x corresponding to dropout, for the unsupervised network B, H input is the embedding vector of x after replication and another dropout, H input is simultaneously input into the large pre-trained text generation network and the small adapter network. During the training process, this component keeps the parameters of the large pre-trained text generation model part frozen, that is, the parameters do not participate in the parameter learning and update process of backpropagation, and only the parameters of the small adapter neural network participate in the update calculation; In the pre-trained learning component based on adapter fine-tuning, the updated parameter of the forward part of the adapter small neural network is W in , the embedding vector is non-linearly optimized through a non-linear activation function, and then input into the latter part of the adapter, using its updated parameter W out The adapter is trained, and the output representation vector of the adapter is Furthermore, combined with the output representation vector of the large pre-trained text generation model After linearly adding them, the final output representation vector of the adapter fine-tuning pre-training learning component is obtained Step 3: Based on the consistency learning of the unsupervised network and combined with the supervised learning of the supervised network, train and optimize the text generation model, Perform the unsupervised consistency learning on the unsupervised network A and the unsupervised network B, so that the prediction targets of the two unsupervised pre-trained text generation neural networks are consistent, and the loss function l of the unsupervised learning unsup is as follows: Among them, S A and S B respectively represent the specific unsupervised network A and the unsupervised network B, which are a pair of twin networks. In extractive text generation, the unsupervised networks A and B are BERT parallel integrated Adapters, and in generative text generation, the unsupervised networks A and B are BART parallel integrated Adapters. X u is the input unsupervised text generation dataset, and represent the input values after data augmentation, that is, two different embedding vector representations obtained after dropout respectively; Meanwhile, jointly optimize the supervised learning of the supervised network to train the supervised text generation model, and the loss function l of the supervised learning is sup as follows; Among them, T(x * ) represents a supervised network, which is the twin network of the unsupervised networks S A and S B . In extractive text generation, it is BERT parallel integrated with Adapter; in generative text generation, it is BART parallel integrated with Adapter. X l is an input text generation dataset with manual annotations, and x * and y * respectively represent the source document data and its corresponding manually annotated target text data: Finally, by combining the consistency learning of the unsupervised network and the supervised learning of the supervised network, the final loss function l is obtained final , which is used for the training and optimization of the model: l final (θ,X) = λl unsup (θ,X u ) + l sup (θ,X l ), X = X u + X l Among them, λ is a hyperparameter representing the importance of the training part of unsupervised consistency learning in the entire model training process, and θ A is the parameter of the small adapter neural module, and θ B is the parameter of the large pre-trained text generation module BERT or BART. θ is the parameter of the entire model, and m is the number of epochs. During the gradient backpropagation calculation, θ B is not updated in each epoch, that is, it does not participate in the gradient backpropagation calculation, while θ A will update the learning parameters in each epoch during the training process; Finally, use the optimized model to perform text generation prediction.
2. A text generation method in a low-resource scenario according to claim 1, characterized in that, For the extractive text generation model, the model is a large BERT model parallelly integrated with a small adapter network; for generative text generation, the model is a large BART model parallelly integrated with a small adapter network.
3. A text generation method in a low-resource scenario according to claim 1, characterized in that, Based on the consistency learning of the unsupervised network and combined with the supervised learning of the supervised network, train and optimize the text generation model.
4. A text generation device in a low-resource scenario, including: A source document input module, used to input a small amount of labeled training source documents and target documents, and input a large amount of unlabeled training source documents; A text generation module in a low-resource scenario, applying the text generation method in a low-resource scenario according to claim 1 or 2; A target text output module, outputting the automatically generated target text through an interface program.
5. A text generation system in a low-resource scenario, the system includes at least one server, and a text generation device in a low-resource scenario according to claim 4 connected to the server. When the server executes the text generation process, it executes the text generation method in a low-resource scenario through the device.