A fully generative method for generating knowledge question-answer pairs based on pre-trained models

Through a completely generated method based on pre-trained models, combined with pointer network and multi-head attention mechanism, the incompatibility and optimization imbalance in the generation of knowledge Q&A is solved, and the consistent semantics of answers and question generation are achieved, which improves the generation effect.

CN116089576BActive Publication Date: 2025-07-08NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211398794.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-07-08
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

In the existing knowledge Q&A generation methods, extracted answer generation is not suitable for more open tasks. Pipeline methods lead to accumulation errors. The task combination method of end-to-end methods lacks cross-task information exchange, resulting in incompatible Q&A and unbalanced optimization of generated Q&A.

Method used

A fully generated method based on pre-trained models is adopted, combined with pointer network and answer-guided multi-headed attention mechanism, and unified knowledge Q&A tasks are handled to ensure the consistent semantics of answers and question generation, and the problem of unbalanced task difficulty is solved through a unified generation method.

Benefits of technology

Improves compatibility and consistency in generating Q&A pairs, and significantly improves generation effect on various data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116089576B_ABST
    Figure CN116089576B_ABST
Patent Text Reader

Abstract

A fully generative knowledge question-answer pair generation method based on a pre-trained model, comprising: selecting an original data set and processing it into the format of <text, question, answer>; learning the high-level semantic representation of each word in the text and the final output representation of the question and answer through the pre-trained model; combining the output expression of the answer with the learned high-level semantic representation of the text, and using a pointer-generator network to copy words from the source text, and finally generating the final answer through a generator; after generating the answer, integrating the already generated information into the output representation of the question through an answer-guided multi-head attention mechanism, and finally generating the question using the generator. The present invention considers the semantic compatibility of answer and question generation, uses a unified generative model to solve the cross-task communication between answers and questions during the training process, improves the comprehensive expression ability of answer generation, and alleviates the optimization imbalance problem caused by task difficulty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer application technologies, and particularly relates to text generation, knowledge Q&A, and multi-task learning, and more particularly to a fully generative knowledge Q&A pair generation method based on a pre-trained model. Background Art

[0002] In recent years, with the rapid development of deep learning, the Q&A task has also been greatly improved. A question answering system is an advanced form of an information retrieval system. It can answer questions posed by users in natural language with accurate and concise natural language. The main reason for the rise of its research is people's need to quickly and accurately obtain information. The question answering system is a research direction that has attracted much attention and has broad development prospects in the current fields of artificial intelligence and natural language processing. The Q&A task based on deep learning extremely depends on a large amount of labeled Q&A pair data. Then, in order to create these data sets, the existing methods either use manual annotation or utilize (semi)-synthetic methods. The (semi)-synthetic data sets [1] are large in scale and low in acquisition cost; however, they do not have the same characteristics as the explicit Q&A tasks. In contrast, because the labeled examples require professional knowledge and careful design, the scale of high-quality manually annotated data sets is much smaller, and the annotation process is very expensive. Therefore, there is a need for methods that can automatically generate high-quality Q&A pairs to generate a large amount of Q&A pair data. In addition, the method of automatically generating high-quality Q&A pairs can also assist in building a knowledge base, improving a search engine by generating questions from documents, and training a chatbot for smooth conversations.

[0003] Currently, much work has been carried out on methods for automatically generating high-quality Q&A pairs at home and abroad, and certain research results have been achieved. The knowledge Q&A pair task is further divided into two subtasks: question generation and knowledge Q&A. For the question generation task, it is mainly divided into two categories: template matching-based methods and neural network-based methods; for the knowledge Q&A task, the existing related research methods can mainly be divided into two categories: deep learning-based extraction methods and deep learning-based generation methods.

[0004] For the question generation task, most early work adopted template-based methods to convert the input text into questions, usually requiring the application of a series of carefully designed general rules or templates. Heilman and Smith (2010) [2] introduced an over-generation and ranking method: their system generated a set of questions and then ranked them to select the best candidates. In addition to generating questions from the original text, generating questions from symbolic representations has also been studied [3]This type relies on manually designed template rules and thus cannot be extended across datasets. In recent years, with the development of neural networks, many neural network-based methods have also been proposed. Serban et al. [4] used an encoder-decoder framework to generate knowledge Q&A pairs from knowledge base triples; Redd et al. [5] generated questions from a knowledge graph; Du [6] et al. studied how to use an attention-based sequence-to-sequence model to generate questions from sentences and investigated the effect of leveraging sentence and paragraph information. Du and Cardie [7] proposed a hierarchical neural sentence-level sequence tagging model for identifying sentences worth asking questions in a text passage.

[0005] For the knowledge Q&A task, Wang & Jiang (2016) [8] proposed a classical recurrent neural model based on the SQuAD dataset. SQuAD defines an extractive knowledge Q&A task where the answer consists of a word span from the corresponding document. Wang & Jiang (2016) demonstrated that learning to point to the answer boundary is more effective than learning to sequentially point to the tokens that make up the answer span. Many subsequent studies have adopted this boundary model and achieved near-human performance on the task. However, the boundary-pointing mechanism is not suitable for more open tasks, including generative knowledge Q&A tasks (Nguyen [9] ). Based on the current powerful pre-trained models, "forcing" the extractive boundary model onto generative datasets currently yields state-of-the-art results.

[0006] Currently, the main work on generating Q&A pairs adopts a pipeline approach. Du

[10] et al. proposed a neural network that combines coreference knowledge through a novel gating mechanism to detect answers worth asking questions and then generates an answer-aware question. Liu

[11] et al. mimicked the way humans ask questions to introduce answer clue style-aware question generation. However, the pipeline architecture brings incompatibility of Q&A pairs and produces cumulative errors during two-stage training. To overcome these drawbacks, Cui

[12] et al. introduced a one-stop method for Q&A pairs, integrating question generation and answer extraction into a unified framework. However, the joint training of answer extraction and question generation leads to optimization imbalance. In addition, simply implementing the interaction between answer generation and question generation through the same encoder does not guarantee the common semantics of the decoded output.

[0007] Although the above methods have all improved the knowledge Q&A generation task to a certain extent, they have not solved or alleviated the fundamental problems in the generation process of generating Q&A pairs. The boundary pointing mechanism in extractive answer generation is not suitable for more open tasks; the cumulative error caused by the pipeline has a great impact; the task combination method of existing end-to-end methods is not sufficient for cross-task information exchange between the two tasks; and the joint learning of the two different generation methods will cause unbalanced optimization problems. Summary of the Invention

[0008] The object of the present invention is to address some problems existing in the existing methods for generating knowledge Q&A pairs. For example, the boundary pointing mechanism in extractive answer generation is not suitable for more open tasks; the cumulative error caused by the pipeline has a great impact, resulting in an incompatible phenomenon in the generated Q&A pairs; the task combination method of existing end-to-end methods is not sufficient for cross-task information exchange between the two tasks, and cannot ensure that the generated answers and questions have consistent semantics; and the joint learning of the two different generation methods will cause unbalanced optimization problems. The present invention provides a fully generative knowledge Q&A pair generation method based on a pre-trained model.

[0009] The present invention believes that the extraction method of obtaining answers is not sufficient to generate natural Q&A pairs and cannot be applied to complex scenarios. Therefore, the traditional knowledge Q&A extraction task is transformed into a fully generative task, which can give play to the advantages of generative methods. In addition, by combining the pointer network in the generation task, the generated answers can not only be directly generated but also be copied from the input text, so that the answer generation method is not limited to the extractive or generative method. Using a unified pre-trained model to help information exchange between the two tasks of knowledge Q&A and question generation, so as to ensure that the generated answers and questions have consistent semantics and improve the compatibility of the generated answers and questions. In addition, during the training process, the same generation method is used for the two tasks, which can effectively solve or alleviate the unbalanced optimization problem caused by different task difficulties.

[0010] The technical solution of the present invention is as follows:

[0011] A fully generative knowledge Q&A pair generation method based on a pre-trained model, comprising:

[0012] Step 1) Select the original data set and process each piece of data in it into the form of <text, question, answer>;

[0013] According to the task form of the knowledge Q&A pair, a reading comprehension type data set is selected for training. According to the input and output requirements of the model, each piece of text only corresponds to one Q&A pair. Therefore, each standard piece of data in the original data set needs to be segmented into the form of <text, question, answer>;

[0014] Step 2) The "text" in each piece of data is passed through a tokenizer based on a large pre-trained model to obtain a shallow representation of each word;

[0015] Step 3) The shallow semantic encoding representation of the text in Step 2 is fed into the encoder of the pre-trained model to learn the high-level semantic encoding of each word in the "text";

[0016] Step 4) According to the high-level semantic encoding output of the "text" obtained in Step 3, this representation and the target "answer" need to be fed into the decoder of the pre-trained model for decoding. Finally, through the answer generator, the output vector representation of the "answer" on the dictionary is obtained;

[0017] Step 5) According to the vector representation of the "answer" obtained in Step 4, combined with the pointer network, the generated "answer" can directly obtain information from the input text, so as to obtain the final vector representation of the "answer" on the dictionary;

[0018] During training, the output vector representation of the "answer" on the dictionary obtained in Step 4) needs to be input into the pointer network, so that the learned output vector of the "answer" on the dictionary can both generate itself and select from the information expressed by the "text", so that the content in the "text" can be directly copied. Finally, in this way, a more comprehensive vector expression of the "answer" is obtained;

[0019] Step 6) According to the high-level semantic encoding output of the "text" obtained in Step 3), this representation needs to be fed into the decoder of the pre-trained model for decoding. Then, using the answer-guided multi-head attention mechanism, the vector expression of the "question" obtained in Step 4) is fused to obtain the decoded vector representation of the "question";

[0020] During training, the standard "question" in the dataset and the representation of the "text" learned by the encoder need to be fed into the encoder of the same pre-trained model for learning to obtain the decoded representation of the "question". Then, through the multi-head attention mechanism, the vector representation of the "answer" obtained in Step 3) is incorporated into the decoded vector representation of the "question" to obtain the final vector representation of the "question";

[0021] Step 7) According to the final vector representation of the "question" obtained in Step 6), use the question generator to generate the "question" to obtain the vector distribution on the dictionary;

[0022] The advantages and beneficial effects of the present invention are:

[0023] The present invention realizes the generation of question-answer pairs based on a unified fully generative pre-trained model, a pointer network, and an answer-guided multi-head attention mechanism. The extraction method of obtaining answers by traditional models is not sufficient to generate natural question-answer pairs and cannot be applied to complex scenarios. Therefore, the traditional knowledge question-answering extraction task is transformed into a fully generative task, which can bring into play the advantages of generative models. In addition, by combining the pointer network in the generation task, the generated answers can not only be directly generated but also be copied from the input text, so that the answer generation method is not limited to the extraction method or the generation method. The unified pre-trained model is used to facilitate the information exchange between the two tasks of knowledge question-answering and question generation, so as to ensure that the generated answers and questions have consistent semantics and improve the compatibility of the generated answers and questions. In addition, during the training process, the same generation method is adopted for the two tasks, which can effectively solve or alleviate the unbalanced optimization problem caused by different task difficulties. The generation effect of this model has been significantly improved on various data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic diagram of the pointer network model of the present patent invention.

[0025] Figure 2 It is the answer-guided multi-head attention mechanism of the present patent invention

[0026] Figure 3 It is a processing flowchart of a fully generative knowledge question-answer pair generation method based on a pre-trained model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] Embodiment 1:

[0028] The following combines the drawings and specific embodiments to detail the extreme multi-label classification data augmentation method based on label and text block attention mechanisms provided by the present invention.

[0029] The present invention mainly adopts theories and methods related to natural language processing. In order to ensure the normal operation of the method, in specific implementations, it is required that the computer platform used is equipped with at least 16G of memory, the number of CPU cores is not less than 4, the main frequency is not less than 2.6GHz, the Linux operating system, and the necessary software environments such as Python 3.6 and above and the pytorch framework are installed.

[0030] See the appendix Figure 3 , in the embodiment of the present invention, a fully generative knowledge question-answer pair generation method based on a pre-trained model, the training process of which includes the following steps:

[0031] Step 1) Select the original data set and process each piece of data in it into the form of <text, question, answer>;

[0032] According to the task form of the knowledge Q&A pairs, a reading comprehension dataset is selected for training. According to the input and output requirements of the model, each piece of text corresponds to only one Q&A pair. Therefore, each standard piece of data in the original dataset needs to be split into the form of <text, question, answer>;

[0033] Step 2) After the "text" in each piece of data passes through the tokenizer based on the large pre-trained model, a shallow representation of each word is obtained;

[0034] Step 3) The shallow semantic encoding representation of the text in Step 2 is fed into the encoder of the pre-trained model to learn the high-level semantic encoding of each word in the "text";

[0035] Step 4) According to the output of the high-level semantic encoding of the "text" obtained in Step 3, during training, this representation and the target "answer" need to be fed into the decoder of the pre-trained model for decoding. Finally, through the answer generator, the output vector representation of the "answer" on the vocabulary is obtained;

[0036] Step 5) According to the vector representation of the "answer" obtained in Step 4, combined with the pointer network, the generated "answer" can directly obtain information from the input text, so as to obtain the final vector representation of the "answer" on the vocabulary, as shown in the appendix Figure 1 shown;

[0037] During training, the output vector representation of the "answer" on the vocabulary obtained in Step 4 needs to be input into the pointer network, so that the learned output vector of the "answer" on the vocabulary can both generate itself and select from the information expressed by the "text", so that the content in the "text" can be directly copied. Finally, a more comprehensive vector representation of the "answer" is obtained in this way;

[0038] Step 6) According to the output of the high-level semantic encoding of the "text" obtained in Step 3), this representation needs to be fed into the decoder of the pre-trained model for decoding. Then, using the answer-guided multi-head attention mechanism, the vector representation of the "question" obtained in Step 4 is fused to obtain the decoded vector representation of the "question", as shown in the appendix Figure 2 shown;

[0039] During training, the standard "question" in the dataset and the representation of the "text" learned by the encoder need to be fed into the encoder of the same pre-trained model for learning to obtain the decoded representation of the "question". Then, through the multi-head attention mechanism, the vector representation of the "answer" obtained in Step 3 is incorporated into the decoded vector representation of the "question" to obtain the final vector representation of the "question";

[0040] Step 7) According to the final vector representation of the "question" obtained in Step 6), use the question generator to generate the "question" to obtain the vector distribution on the vocabulary.

[0041] The original dataset in step 1), denoted as X N :

[0042]

[0043] where C represents the number of data in the dataset, represents an input text, where N represents the length of the sentence; represents the target question, where M represents the sentence length of the target question; represents the target answer, where L represents the sentence length of the target answer;

[0044] The method for performing high-level semantic encoding in step 3) is:

[0045] By sending the vector representation of each word in the shallow text into the encoder of the pre-trained model, a high-level semantic vector representation of the "text" is obtained

[0046] h i = Encoder(d i ), h i ∈ R d

[0047] where i ∈ [1, N], i represents the i-th word of the input text D, N is the maximum number of words in the input text, d represents the dimension of the high-level semantic representation h i and Encoder represents the encoder of the pre-trained model.

[0048] The method for obtaining the distribution vector representation of the generated "answer" on the dictionary in step 4) is based on the high-level representation obtained in step 3 and the target answer representation obtained in step 1 The multi-head cross-attention mechanism of the pre-trained model decoder is used to obtain the input "text" and the target "answer" The attention matrix between them After that, the vector representation of the generated "answer" after decoding is obtained Finally, the "answer" vector representation is sent into the answer generator to generate the distribution vector P of the l-th word in the "answer" on the dictionary voc (w l ):

[0049]

[0050] where and are both learned parameters, V represents the length of the dictionary, d model represents the latent variable after decoding dimension

[0051] In step 5), obtain the vector representation P of the l-th word in the "answer" obtained in step 4 voc (w l ), and cooperate with the pointer network to enable the generated "answer" to directly obtain information from the input text. First, combine the attention vector corresponding to the l-th word in and the high-level semantics of the text obtained in step 3 to obtain its context vector c using the following formula l :

[0052]

[0053] Context vector c l can be regarded as a fixed-length information vector read from the input "text", and this vector is concatenated with the decoder layer hidden vector obtained in step 4 to generate a soft-switching probability P through the following transformation gen ∈ [0, 1]

[0054]

[0055] where W gen and b gen are both learnable parameters, and σ represents the activation function. P gen as a soft switch can determine whether to copy from the input "text" or generate from the dictionary, and obtain the final probability distribution P a (w l ) of the l-th character in the "answer". The formula is as follows

[0056] P a (w l ) = P gen P voc (w l ) + (1 - P gen )P co (w l )

[0057]

[0058] In step 6), it is necessary to send the high-level semantic encoding output of the "text" obtained in step 3 into the pre-trained model decoder for decoding to obtain the decoder layer hidden vector representation of the "question" After that, use the multi-head attention mechanism to capture the "answer" vector representation and the "question" vector representation The semantic relationship H q , as shown in the following formula

[0059]

[0060]

[0061] H q = Concat(head1,..., head z )

[0062] where are all learnable weight matrices, and d k refers to the dimension of each head. To prevent incomplete matching, an answer-guided gate mechanism G is introduced to further incorporate the information of "answer" as follows

[0063]

[0064] h ques = H q ⊙ G

[0065] where W g ∈ R M×L is a learnable parameter, and M and L refer to the lengths of "question" and "answer" respectively. Finally, the vector representation of "answer" is obtained

[0066] In step 7), the "question" vector representation h ques obtained from step 6) needs to be generated to obtain the vector distribution P q (W m ) of the m-th word in the vocabulary, where the learnable parameters

[0067]

[0068] For example, for the publicly available dataset SQuAD, the original input text statement is as follows

[0069] HarVard’s $37.6 billion financial endowment is the largest of any academic institution

[0070] Input into the model, the generated questions and answers can be obtained as follows

[0071] Question: How much money is Harvard’s financial endowment?

[0072] Answer: $37.6 billion financial endowment.

[0073] References:

[0074] [1] Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems. pages 1693–1701.

[0075] [2] Michael Heilman and Noah A. Smith. 2010. Good question!statistical ranking for question generation. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics.

[0076] [3] Xuchen Yao, Gosse Bouma, and Yi Zhang. 2012. Semantics-based question generation and implementation. Dialogue & Discourse 3(2): 11–42.

[0077] [4]Iulian Vlad Serban,Alberto García-Durán,Caglar Gulcehre,SungjinAhn,Sarath Chandar,Aaron Courville,and Yoshua Bengio.2016.Generating factoidquestions with recurrent neural networks:In Proceedings of the 54th AnnualMeeting of the Association for Computational Linguistics。

[0078] [5]Sathish Reddy,Dinesh Raghu,Mitesh M.Khapra,and SachindraJoshi.2017.Generating natural language question-answer pairs from a knowledgegraph using a rnn based question generation model.In Proceedings of the 15thConference of the Eu-ropean Chapter of the Association for ComputationalLinguistics。

[0079] [6]Xinya Du,Junru Shao,and Claire Cardie.2017.Learning to ask:Neuralquestion generation for reading comprehension.In Proceedings of the 55thAnnual Meeting of the Association for Computational Linguistics。

[0080] [7]Xinya Du and Claire Cardie.2017.Identifying where to focus inreading comprehension for neural question generation.In Proceedings of the2017 Conference on Empirical Meth-ods in Natural Language Processing。

[0081] [8]Rajpurkar,Pranav,Zhang,Jian,Lopyrev,Konstantin,and Liang,Percy.Squad:100,000+questions for machine comprehension of text.arXivpreprint arXiv:1606.05250,2016。

[0082] [9]Nguyen,Tri,Rosenberg,Mir,Song,Xia,Gao,Jianfeng,Tiwary,Saurabh,Majumder,Rangan,and Deng,Li.Ms marco:A human generated machine readingcomprehension dataset.arXiv preprint arXiv:1611.09268,2016。

[0083]

[10] Xinya Du and Claire Cardie.Harvest-ing paragraph-level question-answer pairs from wikipedia.In Proceedings of the 56th Annual Meeting of theAssociation for Computational Linguistics,pages 1907–1917,2018。

[0084]

[11] Bang Liu,Haojie Wei,Di Niu,Haolan Chen,and Yancheng He.Asking questions the human way:Scalable question-answer generation from text corpus.New York,NY,USA,2020.Association for Computing Machinery。

[0085]

[12] Shaobo Cui,Xintong Bao,and Xinxing Zu.One stop qamaker:Extract question-answer pairs from text in a one-stop approach.CoRR,abs / 2102.12128,2021。

Claims

1. A fully generative knowledge Q&A pair generation method based on a pre-trained model, and its training process includes the following steps: Step 1) Select the original data set and process each piece of data in it into the form of <text, question, answer>; According to the task form of the knowledge Q&A pair, a reading comprehension type data set is selected for training. According to the input and output requirements of the model, each piece of text only corresponds to one Q&A pair. Therefore, each standard piece of data in the original data set needs to be split into the form of <text, question, answer>; Step 2) After the "text" in each piece of data passes through the tokenizer based on the large pre-trained model, a shallow representation of each word is obtained; Step 3) Send the shallow semantic encoding representation of the text in Step 2 into the encoder of the pre-trained model to learn the high-level semantic encoding of each word in the "text"; Step 4) According to the high-level semantic encoding output of the "text" obtained in Step 3, during training, this representation and the target "answer" need to be sent into the decoder of the pre-trained model for decoding, and finally the output vector representation of the "answer" on the dictionary is obtained through the answer generator; Step 5) According to the vector representation of the "answer" obtained in Step 4, cooperate with the pointer network so that the generated "answer" can directly obtain information from the input text, so as to obtain the final vector representation of the "answer" on the dictionary; During training, the output vector representation of the "answer" on the dictionary obtained in Step 4) needs to be input into the pointer network, so that the learned output vector of the "answer" on the dictionary can both generate itself and select from the information expressed by the "text", so that the content in the "text" can be directly copied, and finally a more comprehensive "answer" vector expression is obtained in this way; Step 6) According to the high-level semantic encoding output of the "text" obtained in Step 3), this representation needs to be sent into the decoder of the pre-trained model for decoding, and then the vector expression of the "question" obtained in Step 4) is fused using the answer-guided multi-head attention mechanism, so as to obtain the decoded vector representation of the "question"; During training, the standard "question" in the data set and the "text" representation learned by the encoder need to be sent into the encoder of the same pre-trained model for learning to obtain the decoded "question" representation. Then, through the multi-head attention mechanism, the vector representation of the "answer" obtained in Step 3) is incorporated into the vector representation of the decoded "question" to obtain the final vector representation of the "question"; Step 7) According to the final vector representation of the "question" obtained in Step 6), use the question generator to generate the "question" to obtain the vector distribution on the dictionary.

2. The fully generative knowledge question-answer pair generation method based on a pre-trained model according to claim 1, wherein, The original data set in step 1), denoted as X N : Where C represents the number of data in the dataset, represents an input text, where N represents the length of the sentence; represents the target question, where M represents the sentence length of the target question; represents the sentence length of the target answer.

3. The fully generative knowledge question-answer pair generation method based on a pre-trained model according to claim 2, wherein The method for performing high-level semantic encoding in Step 3) is: By feeding the vector representations of each word in the shallow text into the encoder of the pre-trained model, a high-level semantic vector representation of the "text" is obtained h i = Encoder(d i ), h i ∈R d Among them, i ∈ [1, N], i represents the i-th word of the input text D, N is the maximum number of words in the input text, d represents the dimension of the high-level semantic representation, and Encoder represents the encoder of the pre-trained model. i Dimension, and Encoder represents the encoder of the pre-trained model.

4. The method for generating complete generative knowledge question-answer pairs based on a pre-trained model according to claim 3, wherein The method for obtaining the distribution vector representation of the generated "answer" in the dictionary in step 4) is based on the high-level representation obtained in step 3 and the target answer representation obtained in step 1 Obtain the input "text" through the multi-head cross-attention mechanism of the pre-trained model decoder And the target "answer" The attention matrix between After that, the vector representation of the generated "answer" after decoding is obtained Finally, the "answer" vector representation is fed into the answer generator to generate the distribution vector P of the l-th word in the dictionary in the "answer" voc (w l ): where and are both parameters for learning. V represents the length of the dictionary, and d model represents the dimension of the decoded latent variable .

5. The method for generating complete generative knowledge question-answer pairs based on a pre-trained model according to claim 4, wherein In step 5), obtain the vector representation P of the l-th word in the "answer" obtained in step 4 voc (w l ), cooperate with the pointer network to enable the generated "answer" to directly obtain information from the input text. First, combine the attention vector corresponding to the l-th word in and the high-level semantics of the text obtained in step 3 to obtain its context vector c using the following formula l : Context vector c l can be regarded as a fixed-length information vector read from the input "text", and this vector is concatenated with the decoding layer hidden vector obtained in step 4) to generate a soft-switching probability P through the following linear transformation gen ∈[0, 1] Where W gen and b gen are both learnable parameters, σ represents the activation function, and P gen , as a soft switch, can determine whether to copy from the input "text" or generate from the dictionary, and obtain the final probability distribution P a (w l ), and the formula is as follows: P a (w l ) = P gen P voc (w l )+(1 - P gen )P co (w l ) 6. The fully generative knowledge question-answer pair generation method based on a pre-trained model according to claim 5, wherein In step 6), the "text" high-level semantic encoding output obtained from step 3) needs to be fed into the pre-trained model decoder for decoding to obtain the decoded layer hidden vector representation of the "question". After that, the multi-head attention mechanism is used to capture the "answer" vector representation from multiple perspectives. and the "question" vector representation The semantic relationship H between them q , as shown in the following formula H q = Concat(head1,..., head z ) Among them The part is a learnable weight matrix, d k refers to the dimension of each head. To prevent incomplete matching, an answer-guided gate mechanism G is introduced to further incorporate the information of "answer" as follows where W g ∈R M×L are learnable parameters, M and L respectively refer to the lengths of "question" and "answer", and finally obtain the vector representation of "answer" 7. The fully generative knowledge question-answer pair generation method based on a pre-trained model according to claim 5, characterized in that, In step 7), it is necessary to generate the vector distribution P of the m-th word in the dictionary from the "question" vector representation h obtained in step 6 ques (W q (W m ), where the learnable parameters

Citation Information

Patent Citations

  • Method for solving video question and answer tasks needing common knowledge by using question-knowledge guided progressive space-time attention network

    CN110704601A

  • Intelligent robot integrating chat, knowledge and task questions and answers

    CN113515613A