A text generation method based on deep learning
By adjusting the loss function and integrating a scoring neural network to refine word probabilities, the method improves the quality of text generation models, addressing discrepancies and enhancing performance in tasks like text summarization.
Patent Information
- Application Number
- CN202310040513.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-13
AI Technical Summary
The existing text generation models have gaps with human natural language in terms of generation quality, grammatical accuracy, context logic and diversity, and it is difficult to achieve fine-grained conditional control.
Adjust the loss function of the text generation model, and combine it with the neural network model to score candidate words. By modifying the loss function and training the scoring model, the probability distribution of candidate words during the generation process is optimized.
The quality of text generation has been improved, especially in text summary and dialogue generation tasks, the matching and diversity of generation results with reference standards has been significantly improved.
Smart Images

Figure CN116226378B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing and relates to a text generation method based on deep learning. Background Art
[0002] Natural language processing is a branch of artificial intelligence. Its main goal is to enable machines to learn human language features and complete a series of tasks that previously required human intelligence. Text generation is an important research direction in the field of natural language processing, including machine translation, summary generation, dialogue generation, story generation and many other tasks.
[0003] The current text generation task mainly adopts the autoregressive generation method, that is, generating words one by one. When each word is generated, the probability of each word in the vocabulary appearing at the current position is calculated based on all the text above, and the generation result is selected according to the sampling algorithm.
[0004] In recent years, large-scale pre-trained language models such as GPT-2 have emerged, which have greatly improved the quality of text generation. By fine-tuning the pre-trained model on a dataset in a specific field, a variety of specific tasks can be achieved.
[0005] However, in practical applications, there are still some key issues that need to be addressed in text generation tasks. First, the text generated by the model still has a gap with human natural language. Whether it is grammatical accuracy, contextual logic, or fluency of expression, it is impossible to achieve complete authenticity. Second, for specific text generation tasks, the current model results are still lacking in indicator performance, especially in tasks such as text summarization and dialogue generation that require both accuracy and diversity. The problem is more obvious. Third, in the actual application of language models, it is usually necessary to conditionally restrict the generation results of the language model according to the specific application scenario, while traditional language models only model the conditional probability of each word appearing, and it is difficult to achieve more fine-grained conditional control. Summary of the invention
[0006] Based on the above-mentioned problems existing in the text generation process, the present invention proposes a text generation method based on deep learning, which can improve the quality of text generation. The present invention solves the problems existing in the text generation model from two aspects: first, adjust the loss function of the text generation model so that the trained language model can generate sentences that are closer to the task objectives. Second, train a neural network model for scoring, evaluate candidate words in the process of autoregressive text generation, change the probability distribution of candidate words according to the evaluation results, and finally change the generation results.
[0007] The method of the present invention is roughly divided into four parts:
[0008] (1) Initial generation model part: mainly refers to the initial text generation model G0 to be optimized. The optional range of model structures includes LSTM, Transformer, GPT, etc. The initial generation model G0 can also directly select a pre-trained model, such as the pre-trained GPT2 model, and then perform domain-specific fine-tuning on it to make it suitable for the current task. In this method, the loss function is modified to further train the initial generation model.
[0009] (2) Representation model part: includes a neural network model M r , which takes text as input and outputs a vector representation of a specified dimension, and is used to participate in the calculation of the new loss model described in this method. The model structure can select dynamic word vector models such as ELMO and Bert.
[0010] (3) Scoring model part: mainly includes a trainable neural network model M score , which takes natural language as input and outputs a number between 0 and 1 as the score. The selection of the model structure needs to consider the performance of the model as much as possible, and pre-trained models such as Bert can be used as the basic structure of the scoring model.
[0011] (4) Guided generation part: that is, combining the scoring model with the further trained initial generation model. During the generation process of the initial generation model, the scoring model is used to score each generated candidate word, and the probability distribution of the original candidate word is changed according to the score, so as to change the subsequent sampling result.
[0012] The technical solution of the present invention:
[0013] A text generation method based on deep learning, including the following steps:
[0014] Step (1): Collect the target task data set, add a loss function, and retrain the initial generation model G0 on the target task data set.
[0015] Step (2): Construct a data set for training the scoring model, including operations such as data set generation and division.
[0016] Step (3): Construct the scoring model M score , and use the data set described in step (2) to train the scoring model until convergence, and verify the performance of the scoring model through the test set.
[0017] Step (4): Combine the scoring model with the generation model, insert a scoring module in the decoding stage of the generation model, and adjust the probability distribution of the candidate words through the output of the scoring model.
[0018] Further, the step (1) includes the following steps:
[0019] (1.1) Collect the relevant data set D according to the target task gen , where each piece of data includes the input of the generation model and the standard reference of the generation result. Select the structure of the initial generation model G0, such as the pre-trained GPT2 model. It is required that the initial generation model G0 can be trained and inferred on the data set D gen .
[0020] (1.2) Determine the structure of the representation model, such as the pre-trained Bert, and construct the representation model M r . Train the representation model M on the data set D gen so that the representation model M r can complete the mapping from sentences to vectors. r
[0021] (1.3) Design the mean square error loss function L r , and the formula is as follows:
[0022]
[0023] where z i represents the vector obtained after the i-th sentence generated by the model during training is input into the representation model M r , that is, M r (y), and this sentence is sampled according to the greedy algorithm based on the output of the last layer of the representation model. represents the vector obtained by mapping the standard generation result of the i-th sentence by the representation model M r , that is
[0024] Add the mean square error loss to the loss L LM of the initial generation model G0, that is:
[0025] L = L r + L LM
[0026] Retrain the initial generation model G0 with L as the new loss function to obtain the optimized generation model G1.
[0027] Further, the step (2) includes the following steps:
[0028] (2.1) Let the optimized generation model G1 complete the generation task on the data set D gen , match the generation result y of each sentence with the corresponding original input x and the standard reference one by one to form a new data set, that is, the scoring model data set D score .
[0029] (2.2) Use the scoring model data set D scoreDivided into a training set, a validation set, and a test set, which are used for training, validating, and testing the scoring model M respectively. score respectively.
[0030] Furthermore, step (3) includes the following steps:
[0031] (3.1) Construct the scoring model M score , determine the model structure, such as a pre-trained Bert model. Set reasonable hyperparameters.
[0032] (3.2) Obtain data from the scoring model dataset D score , split each piece of data into the original input x, the generated result y, and the standard reference. Truncate the generated result y from a random position r, take the segment y[0:r] from the beginning of the sentence to the breakpoint, evaluate this segment according to the evaluation index of the current task, use the original input x and the generated result segment y[0:r] as the input of the scoring model together, and use the evaluation score S eval as the label corresponding to each input.
[0033] (3.3) Start training the scoring model M score , select an appropriate loss function and optimizer, iterate through multiple rounds of training until convergence, adjust the hyperparameters through the validation set to improve performance, and maximize the accuracy of the model in scoring the generated result segments.
[0034] (3.4) Use the test set to test the generalization ability of the scoring model M score , and ensure that the trained scoring model has generality for the target text generation task.
[0035] Furthermore, step (4) includes the following steps:
[0036] (4.1) Generate using the test set. In the decoding stage of the optimized generation model G1, according to the probability distribution output by the model, select the k (k is a hyperparameter) words with the highest probability as candidate words, obtain the probabilities of all candidate words, and form a k-dimensional vector p0.
[0037] (4.2) Concatenate the candidate words to the end of the text already generated in the current generation stage respectively, form k candidate sentence segments, use the scoring model to score the sentence segments, obtain the corresponding output scores, and form a k-dimensional vector s pre .
[0038] (4.3) Combine the probabilities of the candidate words with the scores of each corresponding candidate sentence segment to obtain a k-dimensional vector p final , representing the adjusted probability distribution. The specific formula is as follows, where α is a hyperparameter used to adjust the influence degree of the scoring model on the generation process:
[0039] p final = αp0+(1 - α)s pre
[0040] (4.4) Select the sampling strategy according to different task requirements, and adjust the probability distribution p final for sampling to obtain the generation result of the next position of the sentence.
[0041] Advantages of the present invention:
[0042] The present invention proposes a method for optimizing text generation using a neural network model, aiming to narrow the gap between the generation result of the text generation model and real human natural language. After the optimization of the present invention, the generation quality of the text generation model is significantly improved. Specifically, in the task with the matching degree between the generation result and the reference standard as the evaluation index, the index score of the generation result is significantly improved. Description of the Drawings
[0043] Figure 1 It is a flowchart of the text generation method described in the present invention. Detailed Embodiments
[0044] The following will take the Chinese text abstract task as an example to elaborate on the present invention in detail. The specific process is as Figure 1 shown.
[0045] (1) Retrain the initial generation model:
[0046] (1.1) Select the text abstract generation task on the LCSTS dataset. Select the pre-trained language model GPT2-Chinese trained using the CLUECorpusSmall corpus with the GPT2 structure as the initial generation model G0.
[0047] (1.2) Select the pre-trained Chinese Bert model as the representation model M r , and fine-tune and train this model on the LCSTS dataset. During training, modify the first token of the abstract part of each data to the [MASK] token, and let the Bert model predict this token.
[0048] (1.3) Modify the loss function during the training of the generation model. During the training of the generation model, take the generation result of each data and the corresponding standard abstract, modify the first token of the abstract part of each sentence to the [MASK] token, and input them into the trained representation model M r respectively. Calculate the mean square error loss for the obtained vector representations, and add this loss to the original loss function of the generation model to obtain a new loss function. Use the new loss function to retrain until convergence to obtain the optimized generation model G1.
[0049] (2) Generate a dataset for training the scoring model:
[0050] (2.1) Fine-tune G1 on the LCSTS dataset. After initial convergence, let it generate corresponding summaries for all texts in the dataset. Match the generated summaries with the corresponding original input texts and reference summaries one by one to form the scoring model dataset D. score 。
[0051] (2.2) Divide the scoring model dataset D score 。Since the LCSTS dataset has already completed the division of the training set, validation set, and test set, this step can be skipped.
[0052] (3) Construct and train the scoring model:
[0053] (3.1) Select Bert as the structure of the scoring model, load the parameters of the pre-trained model bert-base-chinese, and select appropriate hyperparameters such as batch_size.
[0054] (3.2) Obtain data from the scoring model dataset D score and split each piece of data into the original input text x, the generated result y, and the reference summary For each piece of data, randomly take a position r on the generated result y, truncate y, and take the segment y[0:r] from the beginning of the sentence to the truncation position. According to the reference summary Calculate the ROUGE-1 metric score S eval 。Connect the original input text x and the generated result segment y[0:r] with the sentence tokenizer [SEP] as the input of the scoring model, and use the metric score S eval as the label corresponding to each input.
[0055] (3.3) Select the mean squared error loss function and the Adam optimizer, and train the scoring model M score for multiple rounds until convergence. Adjust the hyperparameters according to the feedback of the validation set to improve the scoring performance of the model.
[0056] (3.4) Use the test set to test the generalization ability of the scoring model M score When the accuracy of the scoring meets the expected threshold, the training is completed.
[0057] (4) Combine the scoring model with the initial generation model:
[0058] (4.1) Use the test set for generation. At the decoding stage of the text generation model G1, select the top 50 words with the highest occurrence probability as candidate words, obtain the probabilities of all candidate words, and form a 50-dimensional vector p0.
[0059] (4.2) Concatenate the above candidate words to the end of the text generated in the current stage respectively to form 50 candidate sentence fragments. Use the scoring model to score these sentence fragments respectively to obtain the corresponding output scores, and form a 50-dimensional vector s pre 。
[0060] (4.3) According to the calculation formula of the probability distribution p final , combine the probability of the candidate word with the score corresponding to each word, take α as 0.5, and obtain a 50-dimensional vector p final , representing the adjusted probability distribution.
[0061] (4.4) Select a greedy sampling strategy. According to the adjusted probability distribution p final , select the word with the highest probability each time, and continuously generate autoregressively until the end token is generated or the length exceeds the preset limit, and then end the generation.
[0062] After the optimization of the initial generation model GPT2-Chinese by the present invention, according to the test results of the model on the test set, the generation quality of the optimized model is significantly higher than that of the initial generation model. The specific data is shown in Table 1:
[0063] Table 1 Comparison of evaluation index scores before and after the optimization of the initial generation model
[0064] ROUGE-1 ROUGE-2 ROUGE-L GPT2-Chinese 0.2883 0.1571 0.2579 Optimized generation model 0.3472 0.2073 0.3142
Claims
1. A text generation method based on deep learning, characterized in that It includes the following steps: Step (1): Collect the target task dataset, add a loss function, and retrain the initial text generation model G0 on the target task dataset; Step (2): Construct a dataset for training the scoring model, including the generation and division of the dataset; Step (3): Construct a scoring model M score , and use the dataset described in step (2) to train the scoring model until convergence, and verify the performance of the scoring model through the test set; Step (4): Combine the scoring model with the generation model, insert a scoring module in the decoding stage of the generation model, and adjust the probability distribution of candidate words through the output of the scoring model; Further, the step (1) includes the following steps: (1.1) Collect the relevant data set D according to the target task gen , where each piece of data includes the input of the generation model and the standard reference of the generation result; select the structure of the initial generation model G0, and require that the initial generation model G0 can be trained and inferred on the data set D gen ; (1.2) Determine the structure of the representation model and construct the representation model M r ; On the dataset D gen Train the representation model M r so that the representation model M r can complete the mapping from sentences to vectors; (1.3) Design mean squared error loss function L r , the formula is as follows: where z i represents the vector obtained after the i-th sentence input generated by the model during training is input into model M r , that is, M r (y), and this sentence is sampled according to the greedy algorithm based on the output of the last layer of the representation model; represents the vector after mapping of the standard generation result of the i-th sentence by the representation model M r , that is Add the mean squared error loss to the loss L of the initial generation model G0, i.e.: LM Add them together, that is: L = L r + L LM Retrain the initial generation model G0 with L as the new loss function to obtain the optimized generation model G1; Further, the step (2) includes the following steps: (2.1) Let the optimized generation model G1 complete the generation task on the dataset D gen and match the generation result y of each sentence with the corresponding original input x and the standard reference one by one to form a new dataset, namely the scoring model dataset D score ; (2.2) Divide the scoring model dataset D score into a training set, a validation set, and a test set, which are used for training, validation, and testing of the scoring model M score respectively; Further, the step (3) includes the following steps: (3.1) Build a scoring model M score , determine the network structure, and set reasonable hyperparameters; (3.2) Obtain data from the scoring model dataset D score and split each piece of data into the original input x, the generated result y, and the standard reference Truncate the generated result y at the random position r, take the segment y[0:r] from the beginning of the sentence to the breakpoint, evaluate this segment according to the evaluation metric of the current task, use the original input x and the generated result segment y[0:r] as the input to the scoring model, and use the evaluation score S eval as the label corresponding to each input; (3.3) Start the training of the scoring model M score Select a loss function and an optimizer, iterate through multiple rounds of training until convergence, adjust hyperparameters through the validation set to improve performance, and maximize the accuracy of the model in scoring the generated result segments; (3.4) Use the test set to test the generalization ability of model M score to ensure that the trained scoring model has versatility for the target text generation task; Further, the step (4) includes the following steps: (4.1) Use the test set for generation. In the decoding stage of the optimized generation model G1, according to the probability distribution output by the model, select the k words with the highest probability as candidate words, obtain the probabilities of all candidate words, and form a k-dimensional vector p0, where k is a hyperparameter; (4.2) Concatenate the candidate words to the end of the text generated in the current generation stage respectively to form k candidate sentence fragments, use the scoring model to score the sentence fragments, obtain the corresponding output scores, and form a k-dimensional vector s pre ; (4.3) Combine the probability of the candidate word with the score corresponding to each candidate sentence fragment to obtain a k-dimensional vector p final , representing the adjusted probability distribution. The specific formula is as follows, where α is a hyperparameter used to adjust the influence degree of the scoring model on the generation process: p final = αp0 + (1 - α)s pre (4.4) Select a sampling strategy according to different task requirements, and sample according to the adjusted probability distribution p final to obtain the generation result of the next position of the sentence.
2. The text generation method based on deep learning according to claim 1, wherein For the initial generation model G0, the model structure is selected from LSTM, Transformer or GPT.
3. A text generation method based on deep learning according to claim 1, wherein The representation model M described above r is a neural network model that takes text as input and outputs a vector representation of a specified dimension.
4. A text generation method based on deep learning according to claim 3, characterized in that The representation model M r selects ELMO or Bert for its model structure.
5. A text generation method based on deep learning according to claim 1, characterized in that, The scoring model M score is a neural network model that can be trained, takes natural language as input, and outputs a number between 0 and 1 as a score.
6. A text generation method based on deep learning according to claim 5, characterized in that The scoring model M score selects Bert as its model structure.
Citation Information
Patent Citations
Text semantic feature generation optimization method based on deep learning
CN106959946A
Automatic text generation method and device
CN108334497A