Story generation method, system, storage medium and terminal based on pre-training prompts
Through the story generation method based on pre-training prompts, the Para-Comet and question-answer model are used to generate event inference and answers, which solves the problems of logic and consistency in story generation, and realizes the generated story statements are smooth and the plot is reasonable.
Patent Information
- Application Number
- CN202210818147.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-12
AI Technical Summary
In story generation, the existing technology is difficult to ensure that the generated story sentences are smooth, logically correct, and have problems of readability, logic, consistency and coherence.
A story generation method based on pre-training prompts is adopted. By entering the beginning of the story, a Para-Comet pre-trained model is used to generate event inference, a question template is filled in according to the event type, a question-answer model is used to generate answers, a question-answer model is used to calculate the confusion of the answer, and the answer with the smallest score is selected as the story below.
The generated story sentences are smooth, logically correct and reasonable, which improves the readability, logic, consistency and coherence of the story, and reduces the dependence on large-scale data set training.
Smart Images

Figure CN115905852B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a story generation method, system, storage medium and terminal based on pre-training prompts, and belongs to the field of natural language processing in the computer field. Background Art
[0002] Open-ended story generation is a classic task in the field of natural language processing. How to ensure the consistency, coherence and logic of the story during generation has always been a very challenging problem. With the development of deep neural networks, deep neural network models with more parameters and more complex structures are constantly being proposed and applied to story generation. However, models often require a lot of data and time to train in order to achieve certain results, which is difficult to meet in many scenarios. Therefore, some people have studied the use of models that have been trained on large-scale datasets to perform different tasks. The previous mainstream approach was to fine-tune the pre-trained model, change part of the model's structure and retrain it to perform downstream tasks. But for each task, all the parameters of the model need to be saved once, which is a considerable expense in most cases.
[0003] Recently, research on Prompt has been in full swing. A lot of work has shown that Prompt has advantages and performance that general fine-tuning models cannot achieve in few-sample and zero-sample scenarios, which is very consistent with the characteristics of pre-trained language models. Prompt Learning is based on language models and directly models text probabilities. In order to perform tasks using these pre-trained models, the original input is modified into a text string with some unfilled slots using the template Prompt, and then the language model is used to probabilistically fill the unfilled information to obtain the final string, from which the final output can be derived. Summary of the invention
[0004] The present invention aims to solve the following technical problems:
[0005] The present invention is dedicated to solving the following difficulties in story generation, and generating stories with smooth sentences, correct logic and reasonable plots:
[0006] Readability: The most basic requirement for the generated story is that it should be readable by others and that they can understand the narration and development of the story. It cannot be a bunch of gibberish, and you cannot ask for a Chinese story and end up generating an English one.
[0007] Logic: The characters in the story have logical characteristics, the content of the story conforms to the social logic system, and the direction of the story adheres to logical rules.
[0008] Consistency: The context roles, environment, events, etc. are consistent.
[0009] Coherence: The stories told before and after are continuous. There will not be a situation where one story is told in the previous part and another story begins in the next part.
[0010] The present invention adopts the following technical solution to solve the technical problem: a story generation method based on pre-training prompts, comprising the following steps:
[0011] 1) Inputting the beginning of a story into a first pre-trained model, wherein the beginning of the story includes a plurality of sentences, and the first pre-trained model generates event reasoning corresponding to each sentence;
[0012] 2) Fill in different question templates according to the type of event to get questions about the story context;
[0013] 3) Use the question-answering model to answer questions, generate answers, and obtain an answer set;
[0014] 4) For each answer in the answer set, use the second pre-trained model to calculate its perplexity, and select the answer with the smallest score as the story context.
[0015] Preferably, the first pre-trained model is a Para-Comet pre-trained model.
[0016] Preferably, the implementation process of step 2) is:
[0017] 2.1) Construct different link templates and question templates according to the type of events;
[0018] 2.2) Fill the event reasoning obtained in step 1) into the corresponding link template, and splice the beginning of the story before the link template, input the link template into the RoBERTa model to generate sentences, and then splice the generated sentences after the beginning of the story as the input of the next round of RoBERTa model, obtain the role corresponding to each event reasoning, and obtain the <event, role> pair;
[0019] 2.3) Fill in the event and role into the question template to get the final question.
[0020] Preferably, the implementation process of step 3) is:
[0021] The sentences generated by each round of the RoBERTa model are concatenated at the beginning of the story as a document, and input into the ELI5QA model together with the question to generate a set of candidate answers; or the question and the document are concatenated into a training set and input into the BART model to generate a set of candidate answers.
[0022] Preferably, the implementation process of step 4) is: merge the answer sets for each question, calculate the perplexity of each element in the set, i.e. the answer, using the second pre-trained model, and select the element with the smallest perplexity as the context of the story, wherein the second pre-trained model is the GPT2 model.
[0023] Preferably, after the ELI5QA model generates an answer, the answer is cleaned to a certain extent: multiple forbidden phrases are collected to form a set, and when the answer is a single sentence, it is added to the candidate answer set; when the answer is multiple sentences, those sentences containing elements that appear in the forbidden phrase set and all sentences following it are discarded; if the length of the first sentence is less than 6, the content of the first sentence is also discarded, and the remaining sentences are added to the candidate answer set to obtain a candidate answer set corresponding to each question.
[0024] Preferably, the BART model generates a candidate answer set by splicing the question and the document into the form of Question--T--Document, and the spliced content is used as the value of the corresponding keyword text; all sentences after the current sentence are used as the value of the corresponding keyword question, and form a dictionary with the text key-value pair above, which is stored in the jsonlines file. In this way, the training set and the validation set are obtained. BART is trained. When generating, the question and the document are spliced into the form of training set data, the Top-k sampling decoding algorithm is selected, the k value of the algorithm is set, and the generated results are cleaned using the same strategy as the ELI5QA model to obtain a candidate answer set corresponding to each question.
[0025] A story generation system based on pre-trained prompts, including the following modules:
[0026] 1) Common sense reasoning module: used to generate reasoning events for the corresponding sentences using the Para-COMET pre-trained model for the input story beginning;
[0027] 2) Question generation module: used to fill in different question templates according to the type of reasoning event, and obtain multi-angle guiding questions about the above story;
[0028] 3) The answer generation module uses the ELI5QA model or the BART model to answer questions and generate a set of candidate options for the story context;
[0029] 4) The scoring selection module is used to send an answer to each element in the candidate set, calculate its perplexity score using the GPT2 model, and select the one with the smallest score as the following sentence.
[0030] A storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned story generation method based on pre-training prompts.
[0031] A terminal comprises: a processor and a memory;
[0032] The memory is used to store computer programs;
[0033] The processor is used to execute the computer program stored in the memory so that the terminal executes the above-mentioned story generation method based on pre-training prompts.
[0034] Compared with the prior art, the present invention adopts the above technical solution and has the following beneficial effects:
[0035] 1) This invention aims to solve the problem of lack of data sets for training large-scale deep neural network models. By combining the emerging prompt learning idea, this invention constructs high-quality prompt templates to fully stimulate the potential of pre-trained models and help the models recall the knowledge learned during training to better complete downstream tasks.
[0036] 2) The system proposed in this invention preliminarily explores the combination of multiple pre-trained models with different architectures and applied to different downstream tasks: these models are linked through prompt templates;
[0037] 3) The system proposed by the present invention has strong vitality. In the future, if a pre-trained model with better performance appears, when using it to replace the model in the system, it only needs to change the template style. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the overall process of the proposed story generation system framework. DETAILED DESCRIPTION
[0039] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0040] The novel story generation system proposed in this paper consists of four modules, each of which uses a different pre-trained language model. The whole system combines different models through high-quality prompt question templates to perform story generation tasks. The overall architecture of the system is as follows: Figure 1 For input S, the common sense reasoning module generates n possible reasoning events E 1 ,E 2 ,...,E n ; The question generation module uses these events to construct the corresponding question Q 1 ,Q 2 ,...,Q n , the answer generation module answers the question and obtains the answer A1 ,A 2 ,...,A n ; The scoring selection module selects the answer A' with the minimum perplexity as the next part of the story. It is merged with the beginning to form the beginning for the new round of generation.
[0041] (1) Commonsense reasoning module
[0042] The commonsense reasoning module uses the Para-Comet pre-trained model to generate corresponding inference events for each sentence based on the input story beginning. Commonsense reasoning has always been regarded as high-quality external knowledge to assist text generation. With the injection of external knowledge, the model will not be limited to the input text but can generate from a perspective more in line with human social common sense.
[0043] The Para-Comet model was proposed to solve the problem that the Comet model has only been trained in phrase commonsense reasoning and cannot perform long-sentence reasoning. It considers implicit commonsense knowledge from two aspects: semantic knowledge based on world knowledge and culture-specific social knowledge; situational knowledge based on causal understanding and cognitive reasoning. It extracts semantic knowledge from the context and uses recurrent memory enhancement to assist in generation.
[0044] Based on the Comet model, Para-Comet inputs narrative text and outputs corresponding commonsense inferences for each sentence. These inferences are in the form of events and have 9 types, namely: xIntent, xNeed, xAttr, xWant, xEffect, xReact, oWant, oEffect, oReact. The present invention selects the first six for subsequent generation.
[0045] (2) Question generation module
[0046] The events obtained from the commonsense reasoning module do not show the corresponding character roles. Therefore, it is necessary to first use a pre-trained model to obtain the most likely character related to the event, and then fill the event Inference and the character Character into the question template to generate the final question.
[0047] Select the RoBERTa model trained on SQuAD (Stanford Question Answering Dataset) provided by Huggingface. Taking xNeed and xIntent as examples, the following templates are designed:
[0048] -Who needs to Event?
[0049] -Who wants to Event?
[0050] Fill the event into the template, and splice the previously generated content before the question, and then let the RoBERTa model generate it. The generated content is filtered and cleaned, so that the role corresponding to each reasoning event is obtained.
[0051] After having the <event, role> pair, fill it into the question template according to different event types to generate corresponding questions. Still taking xNeed and xIntent as an example, design the following template:
[0052] -What does character do to needEvent?
[0053] -Why does character do Event?
[0054] Fill in the events and roles into the template to get the final question.
[0055] (3) Answer Generation Module
[0056] The questions obtained in the question generation module are input into the pre-trained model to generate stories. Two pre-trained models are selected: ELI5QA and Finetuned BART.
[0057] The ELI5QA model is a long-text question-answering model trained in the Fairseq-py framework. It is trained on the ELI5 dataset. The full name of the ELI5 dataset is Explain Like I'm Five, which is collected from the Reddit community corpus. In this dataset, people give long and easy-to-understand answers to open-ended questions, just like giving answers to five-year-old children.
[0058] The input of the ELI5QA model is a question and a document. The model generates answers to questions from the document and the knowledge obtained through training. In this system, in order to ensure the consistency and continuity of the generated stories as much as possible, the beginning of the story and the generated content are used as documents and input into the model together with the questions.
[0059] The original intention of constructing the question template is to get short and relevant sentences from the ELI5QA model as the context of the story. To ensure this, the Top-k sampling algorithm is used for generation. During the decoding process, the k tokens with the highest probability are taken and their probabilities are summed up as sum-topk. The probability redistribution transformation is as follows: the above k tokens are divided by the probability and sum-topk to get their new probabilities, and the probabilities of the remaining tokens become 0. The constant k is a given value. If k is set too small, it will be easy to generate more bland or general sentences. When k is very large, the candidate set will contain some inappropriate tokens. Select k=50.
[0060] Even if the Top-k sampling algorithm is used, there is still a certain probability that the generated answer will contain meaningless or repeated sentences, so the answer needs to be cleaned to a certain extent. The strategy adopted is: collect more than 100 prohibited phrases and form a set. When the answer is a single sentence, add it to the candidate items; when the answer is multiple sentences, discard those sentences that contain elements that appear in the prohibited phrase set and all the sentences following it. If the length of the first sentence is less than 6, the content of the first sentence is also discarded, and the remaining sentences are added to the candidate items. In this way, a set of candidate answers corresponding to each question can be obtained.
[0061] In addition to the ELI5QA model, BART was also selected as a pre-trained model for generating answers. However, the original BART model did not perform well in QA tasks because the data was not in the form of question-answering corpus during training. Therefore, it was finetune on the ROCStories dataset to improve the generation effect. The ROCStories dataset is a collection of common sense short stories, containing 100,000 stories of five sentences. Each story follows an everyday theme. These stories contain various common sense causal and temporal relationships between everyday events. The dataset was randomly divided into training set, validation set, and test set in a ratio of 70%:15%:15%. For each story in the training set and validation set, it was processed as follows:
[0062] 1) For each of the first four sentences in the story, the common sense reasoning module and the question generation module are used to generate 20 questions;
[0063] 2) The current sentence and all the sentences before it are taken as Documents and concatenated with the question in the following way:
[0064] Question--T--Document
[0065] The concatenated content is used as the value of the corresponding keyword text.
[0066] 3) All sentences following the current sentence are used as the value of the corresponding keyword question, and together with the text key-value pairs above, form a dictionary and are stored in the jsonlines file.
[0067] In this way, the training set and validation set are obtained. The learning rate is set to 2e-5, the batch size is set to 16, and BART is trained.
[0068] During generation, Question and Document are also concatenated into the form of training set data, the Top-ksampling decoding algorithm is selected, k is set to 50, and the generated results are cleaned using the same strategy as the ELI5QA model to obtain a set of candidate answers for each question.
[0069] (4) Rating selection module
[0070] After obtaining a set of candidate answers for each question, these sets are merged, the perplexity of each element in the set is calculated, and the element with the smallest perplexity is selected as the context of the story.
[0071] The language model is used to calculate the probability of a sentence. Given a sentence (word sequence) S = W 1 , W 2 , ..., W k , the probability that the model can generate this sentence is:
[0072] P(S)=P(W 1 , W 2 , …, W K )=p(W 1 )p(W 2 |W 1 )…P(W K |W 1 , W 2 , …, W K - 1 )
[0073] The language model that assigns higher probability values to the sentences in the test set is better. Perplexity is the unit that quantifies the quality of the model. Its formula is as follows:
[0074]
[0075] The formula shows that the language model that makes the sentence perplexity score smaller is better.
[0076] GPT2 is selected as the benchmark model to calculate the probability of a sentence. For a given sentence, shift it to the left by one position as the label, remove the last position as the label, remove the last position as the input, and calculate the cross entropy loss between the output of the input and the label. The loss is raised to the power of 2 to get the perplexity score of the sentence, but in the end, we need to make a comparison, so we can directly compare the size of the cross entropy loss. The smaller the loss, the smaller the perplexity score, and the better the sentence.
[0077] (5) Generation process
[0078] Input the beginning of the story, and the system will give the next sentence of the story. If the target number of sentences of the story is set, the system will put the currently generated story after the beginning of the story as the input of the next process. The system will continue to generate stories until the number of story sentences reaches the requirement.
[0079] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A story generation method based on pre-trained prompts, characterized in that, it includes the following steps: 1) Input the beginning of the story into the first pre-trained model. The beginning of the story contains multiple sentences, and the first pre-trained model generates event inferences corresponding to each sentence. Among them, the first pre-trained model is the Para-Comet pre-trained model; 2) According to the type of event, fill it into different question templates to obtain questions about the previous context of the story. The specific process is: 2.1) Construct different link templates and question templates according to the type corresponding to the event; 2.2) Fill the event inferences obtained in step 1) into the corresponding link templates, splice the beginning of the story before the link templates, and input them into the RoBERTa model for generation to obtain the characters corresponding to each event inference, and get <event, character> pairs; 2.3) Fill the events and characters into the question templates to obtain the final questions; 3) Use a question-answering model to answer the questions, generate answers, and obtain an answer set; 4) For each answer in the answer set, use the second pre-trained model to calculate its perplexity, and select the answer with the smallest score as the following context of the story. Among them, the second pre-trained model is the GPT2 model.
2. The story generation method based on pre-trained prompts according to claim 1, characterized in that, the implementation process of step 3) is: Take the beginning of the story as the document, and input it into the ELI5QA model together with the questions to generate a candidate answer set; or splice the questions and the beginning of the story and input them into the BART model to generate a candidate answer set.
3. The story generation method based on pre-trained prompts according to claim 2, characterized in that, the implementation process of step 4) is: Merge the answer sets of each question, calculate the perplexity Perplexity of each element in the set, that is, the answer, using the second pre-trained model, and select the element with the smallest perplexity as the following context of the story.
4. A story generation system based on pre-trained prompts, characterized in that, it includes the following modules: A common sense reasoning module, which is used to generate inference events corresponding to sentences for the input beginning of the story using the Para-COMET pre-trained model; A question generation module, which is used to fill the inference events into different question templates according to the type of the inference events to obtain multi-angle guiding questions about the previous story. The specific process is: Construct different link templates and question templates according to the type corresponding to the event; Fill the obtained event inferences into the corresponding link templates, splice the beginning of the story before the link templates, and input them into the RoBERTa model for generation to obtain the characters corresponding to each event inference, and get <event, character> pairs; Fill the events and characters into the question templates to obtain the final questions; An answer generation module, which answers the questions by using the ELI5QA model or the BART model to generate a set of candidate items for the following context of the story; A scoring and selection module, which is used to calculate the perplexity score of each element in the set of candidate items, that is, the answer, using the GPT2 model, and select the one with the smallest score as the following sentence.
5. A storage medium, It is characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the story generation method based on pre-training prompts described in any one of claims 1 to 3 is implemented.
6. A terminal, It is characterized in that include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the terminal executes the story generation method based on pre-training prompts as described in any one of claims 1-3.
Citation Information
Patent Citations
Construction method of question-answering system based on document set multi-hop reasoning
CN111538819A
Question-driven social network answer abstract automatic generation method and device
CN114048309A