A method for extracting answers from a question-answering system based on deep model transfer learning

By adopting deep model transfer learning methods in open domain question-and-answer systems, using question word classification and entity marking technology, combined with the self-attention mechanism of deep transfer model, the existing system's lack of capabilities in large-scale and multi-domain question-and-answer questions is solved, and higher answer extraction accuracy and user experience are achieved.

CN116432751BActive Publication Date: 2025-06-06KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310394497.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-06-06
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

The existing open domain question and answer system is insufficient in handling large-scale and multi-field question-and-answer questions, and due to insufficient label sample data, it is difficult to obtain a higher accuracy of answer extraction.

Method used

The answer extraction method of the question-and-answer system based on deep model transfer learning is adopted. By pre-processing the data set, the question classification method of the question words is used to classify the questions, the expected answer type of the question sentence is obtained, and the text paragraphs are physically marked. Then, the processed question and text paragraphs are input into the depth migration model, and the semantic similarity is calculated using the self-attention mechanism, and the final probability of the answer is finally calculated through the full connection layer and sigmoid normalization.

Benefits of technology

It improves the answer extraction accuracy of the Q&A system in multi-field Q&A questions, and can provide users with reasonable and accurate text answers faster and more accurately, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116432751B_ABST
    Figure CN116432751B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for extracting answers to a question-answering system based on deep model transfer learning, and belongs to the technical field of information retrieval. The present invention first preprocesses the Stanford data set, and when the text paragraphs and questions enter the model to calculate the vector representation, the questions in the data set are classified by the question classification method of the question word, and the expected answer type corresponding to the question is obtained; according to the classification information of the question, the text paragraph is processed, and the entity corresponding to the expected answer type in the question is specially marked; secondly, the processed question and text paragraph are input into the BERT model, and the self-attention mechanism of the BERT model is used to focus on the text fragments with special marks, and the semantic similarity between the text paragraph and the question is calculated; finally, the final probability is calculated in the similarity layer through the full connection layer and sigmoid normalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for extracting answers from a question-answering system based on deep model transfer learning, and belongs to the technical field of information retrieval. Background Art

[0002] With the rapid growth of network data, it has become a huge challenge to obtain relevant information from massive network data. Search engines have solved this problem to a certain extent. The input of search engines is a set of keywords, but sometimes it is difficult to accurately express the user's information needs with keywords. At the same time, sometimes the granularity of the information required by the user is not a document, but a descriptive paragraph, sentence, conclusion, name or number, etc., but the search engine returns a collection of documents for a query, and the user still needs to find relevant content from it. If users can interact with the system in a more natural way, users can express their information needs naturally and accurately, and the system can directly return the content that users want to know. Based on such needs, open domain question answering systems have become another hot spot in the field of information systems after search engines.

[0003] The capabilities of existing open-domain question answering systems mainly deal with questions that can be answered by extracting answers directly from a document set. Semantic-based answer extraction provides exact answers from relevant documents or statements. It usually involves candidate answer extraction and answer ranking processes. However, this task is challenging because question facts are usually brief and limited when multiple entities are involved. Building an end-to-end framework learning architecture enables the model to produce the expected answer without taking a separate process. The answer extraction method that combines question sentence and text paragraph information is conducive to more accurate extraction of answers in text paragraphs. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method for answer extraction in a question-answering system based on deep model transfer learning, which mainly solves the problem that in open domain answer extraction tasks, the current early answer extraction models constructed based on a large number of artificial rule methods are insufficient and difficult to handle large-scale, multi-field question-answering problems; and can solve the problem that in specific answer extraction tasks, it is difficult to obtain a high accuracy rate due to insufficient labeled sample data and insufficient information utilization.

[0005] The technical solution of the present invention is: a method for extracting answers from a question-answering system based on deep model transfer learning, wherein the method first preprocesses a data set, and when a text paragraph and a question enter a deep transfer model to calculate a vector representation, the questions in the data set are classified by a question classification method of question words to obtain the expected answer type corresponding to the question; the text paragraph is processed according to the classification information of the question, and the entity corresponding to the expected answer type in the question is marked; secondly, the processed question and text paragraph are input into a deep transfer model, and the self-attention mechanism of the deep transfer model is used to focus on the marked text paragraph, and the semantic similarity between the text paragraph and the question is calculated; finally, the final probability is calculated at the similarity layer through a fully connected layer and sigmoid normalization.

[0006] As a further solution of the present invention, the specific steps of the method are as follows:

[0007] Step 1: Preprocess the questions and text paragraphs in the dataset. The preprocessing includes classifying the questions according to the interrogative words to obtain the classified questions, so as to determine the expected answer type of the questions; then mark the entities of the text paragraphs according to the expected answer type obtained by the question classification to obtain the marked text paragraphs;

[0008] Step 2: The classified questions and marked text paragraphs in Step 1 are sent as input to the deep transfer model. The encoder is used in the deep transfer model BERT to process them, unify the text length, and generate question feature vectors and candidate text segment feature vectors.

[0009] Step 3: Through the self-attention mechanism of the deep transfer model BERT, focus on the marked text paragraphs, calculate the semantic similarity between the marked text paragraphs and the question, and obtain the best answer text;

[0010] Step 4: Calculate the final probability of the answer and the best answer text in the similarity layer through the fully connected layer and sigmoid normalization.

[0011] As a further solution of the present invention, the specific steps of Step 1 are:

[0012] Step 1.1: Detect the question word type in the question sentence and classify the question sentence to obtain the classified question sentence, so as to determine the expected answer type of the question sentence;

[0013] Step 1.2: Based on the expected answer type of the question, use Spacy NER to detect the relevant entity types in the text paragraph, highlight the entities in the text paragraph that contain the expected answer type of the question, reduce the influence of irrelevant entities in the answer sentence, and obtain the marked text paragraph;

[0014] Step 1.3: Finally, the processed question and text paragraph are input into the deep transfer model BERT to calculate the semantic similarity.

[0015] As a further solution of the present invention, the Step 2 includes:

[0016] The classified question and marked text paragraphs in Step 1 are sent as input to the deep transfer model BERT. The deep transfer model BERT is used to unify the text length of the input and generate an attention masking matrix attention_mask to obtain the question feature vector and the candidate text segment feature vector;

[0017] When unifying the text length, short sentences are increased to a fixed length by padding, and long sentences are truncated to a fixed length. The matrix value of the padded part of the text is 0, and the matrix value of the original part of the text is 1. In this way, when calculating the self-attention mechanism, the deep transfer model BERT will only focus on the input part, and reduce the attention weight of the padded part to almost 0.

[0018] As a further solution of the present invention, in Step 3, the self-attention mechanism of the deep transfer model BERT includes a multi-head attention module, a forward neural network and a layer normalization module; after obtaining the vector representation sequence X of the question q and the text paragraph t q and X t Finally, they are concatenated to obtain the input sequence X of the deep transfer model BERT, and X is input into the multi-head attention module to obtain the context information of each word in the input question and text paragraph. Finally, the output of the final self-attention structure is obtained through a feedforward neural network and layer normalization.

[0019] In Step 2, the classified questions and marked text paragraphs in Step 1 are fed into the deep transfer model as input, and the encoder is used in the deep transfer model BERT to process them, unify the text length, and generate question feature vectors and candidate text segment feature vectors;

[0020] The BERT model converts the input questions and text paragraphs into vector form. Each word in the question and text paragraph is obtained by adding the word vector, segment vector and position vector. The word vector is the vector representation of each word in the question and text paragraph. The segment vector is to help the model distinguish the input questions and text paragraphs. The position vector records the position information of each word in the model.

[0021] As shown in formula (1), after obtaining the input of question and text paragraph, the BERT model encodes the question q and text paragraph t:

[0022] xi =p i +e wi +s i (1)

[0023] where x i Yes X q or X t The i-th word w in i The vector representation of X q or X t are the vector representation sequences of question q and text paragraph t, respectively, i It is the word w i The position vector, e wi It is the word w i The word vector of s i It is a segment vector representation that enables the model to distinguish between questions and text paragraphs. After obtaining the matrix representation of questions and text paragraphs, they are input into the BERT model, combined with the attention masking matrix processed with the expected answer type information, and encoded using the self-attention mechanism.

[0024] The beneficial effects of the present invention are:

[0025] 1. Because the open-domain question-answering system is more inclined to handle multi-domain comprehensive basic problems than to handle professional specific fields, and can often handle a wider range of problems, the deep transfer model can be applied to the answer extraction task of the open-domain question-answering system. The present invention uses the BERT model, question classification, text tagging and other methods to process, and makes full use of the combination of transfer learning and deep learning to improve the accuracy of answer extraction of the question-answering system, and provide users with more reasonable and accurate text answers faster and better;

[0026] 2. The present invention uses a method that combines the answer extraction method with deep transfer learning (deep transfer model BERT), which can solve the problem of difficulty in obtaining a high accuracy rate in specific answer extraction tasks due to insufficient labeled sample data and insufficient information utilization. It can provide users with more accurate text answers, improve the quality of retrieved answers in answer extraction tasks, and improve the user experience to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flow chart of the steps of the present invention;

[0028] Figure 2 It is a diagram of the entity detection framework of the present invention;

[0029] Figure 3 is an example flow chart of the present invention using only the deep transfer model for the answer extraction task;

[0030] Figure 4is a flow chart of the open domain question answering system of the present invention;

[0031] Figure 5 It is a framework diagram of the deep transfer model of the present invention combined with question information in answer extraction; DETAILED DESCRIPTION

[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.

[0033] Example 1: Figure 1-Figure 5 As shown, a method for extracting answers for a question-answering system based on deep model transfer learning is described. The method first preprocesses a data set. When a text paragraph and a question enter a deep transfer model to calculate a vector representation, the questions in the data set are classified by a question classification method of question words to obtain the expected answer type corresponding to the question. According to the classification information of the question, the text paragraph is processed to mark the entity corresponding to the expected answer type in the question. Secondly, the processed question and text paragraph are input into a deep transfer model, and the self-attention mechanism of the deep transfer model is used to focus on the marked text paragraph to calculate the semantic similarity between the text paragraph and the question. Finally, the final probability is calculated at the similarity layer through a fully connected layer and sigmoid normalization.

[0034] As a further embodiment of the present invention, Figure 1 As shown in the flowchart of the steps of the present invention, the framework diagram of the deep migration model of the present invention combined with question information in answer extraction is as follows Figure 5 As shown, the specific steps of the method of the present invention are as follows:

[0035] Step 1: Preprocess the questions and text paragraphs in the dataset. The preprocessing includes classifying the questions according to the interrogative words to obtain the classified questions, so as to determine the expected answer type of the questions; then mark the entities of the text paragraphs according to the expected answer type obtained by the question classification to obtain the marked text paragraphs;

[0036] The Step 1 dataset is derived from the Stanford Question Answering Dataset (SQuAD), a reading comprehension dataset consisting of questions asked by crowdworkers on a set of Wikipedia articles, where the answer to each question is a piece of text from the corresponding article, and some questions may not be answerable. SQuAD 1.1 contains more than 100,000 question-answer pairs for more than 500 articles, and SQuAD 2.0 combines more than 100,000 questions in SQuAD 1.1 and adds more than 50,000 unanswerable questions, which are designed by crowdworkers in an adversarial way and look similar to answerable questions. Figure 4 This is a flow chart of the open domain question answering system of the present invention.

[0037] The specific steps of Step 1 are:

[0038] Step 1.1: Detect the types of interrogative words in the question sentences (interrogative pronouns "who, what, which, whose" and interrogative adverbs "when, where, how, why") and classify the question sentences to obtain the classified questions, so as to determine whether the expected answer type of the question sentence is a noun or an adverbial;

[0039] Step 1.2: Based on the expected answer type of the question, use Spacy NER to detect the relevant entity types in the text paragraph, highlight the entities in the text paragraph that contain the expected answer type of the question, reduce the influence of irrelevant entities in the answer sentence, and obtain the marked text paragraph, such as Figure 2 Shown is a diagram of the entity detection framework of the present invention;

[0040] Step 1.3: Finally, the processed question and text paragraph are input into the deep transfer model BERT to calculate the semantic similarity.

[0041] Step 2: Send the classified questions and marked text paragraphs in Step 1 as input to the deep transfer model. In the deep transfer model BERT, an encoder is used to process them and unify the text length. The BERT model converts the input questions and text paragraphs into vector form. Each word in the question and text paragraph is obtained by adding the word vector, segment vector and position vector to generate the question feature vector and the candidate text segment feature vector.

[0042] The Step 2 includes:

[0043] The classified questions and marked text paragraphs in Step 1 are sent as input to the deep transfer model BERT. The deep transfer model BERT is used to unify the text length of the input. The model converts the input questions and text paragraphs into vector form and generates an attention masking matrix attention_mask, thereby obtaining the question feature vector and the candidate text segment feature vector;

[0044] When unifying the text length, short sentences are increased to a fixed length by padding, and long sentences are truncated to a fixed length. The matrix value of the padded part of the text is 0, and the matrix value of the original part of the text is 1. In this way, when calculating the self-attention mechanism, the deep transfer model BERT will only focus on the input part, and reduce the attention weight of the padded part to almost 0.

[0045] In Step 2, the classified questions and marked text paragraphs in Step 1 are sent as input to the deep transfer model. The encoder is used in the deep transfer model BERT to process them and unify the text length. The model converts the input questions and text paragraphs into vector form and generates question feature vectors and candidate text segment feature vectors.

[0046] The BERT model converts the input questions and text paragraphs into vector form. Each word in the question and text paragraph is obtained by adding the word vector, segment vector and position vector. The word vector is the vector representation of each word in the question and text paragraph. The segment vector is to help the model distinguish the input questions and text paragraphs. The position vector records the position information of each word in the model.

[0047] As shown in formula (1), after obtaining the input of question and text paragraph, the BERT model encodes the question q and text paragraph t:

[0048] x i =p i +e wi +s i (1)

[0049] where x i Yes X q or X t The i-th word w in i The vector representation of X q or X t are the vector representation sequences of question q and text paragraph t, respectively, i It is the word w i The position vector, e wi It is the word w i The word vector of s i It is a segment vector representation that enables the model to distinguish between questions and text paragraphs. After obtaining the matrix representation of questions and text paragraphs, they are input into the BERT model, combined with the attention masking matrix processed with the expected answer type information, and encoded using the self-attention mechanism.

[0050] Step 3: Through the self-attention mechanism of the deep transfer model BERT, focus on the marked text paragraphs, calculate the semantic similarity between the marked text paragraphs and the question, and obtain the best answer text;

[0051] In Step 3, the self-attention mechanism of the deep transfer model BERT includes a multi-head attention module, a forward neural network, and a layer normalization module; after obtaining the vector representation sequence X of the question q and the text paragraph t q and X tFinally, they are concatenated to obtain the input sequence X of the deep transfer model BERT, and X is input into the multi-head attention module to obtain the context information of each word in the input question and text paragraph. Finally, the output of the final self-attention structure is obtained through a feedforward neural network and layer normalization.

[0052] Step 4: Calculate the final probability of the answer and the best answer text in the similarity layer through the fully connected layer and sigmoid normalization.

[0053] In the process of attention calculation, because Step 2 masks the unmarked keywords in the answer sentence, the deep transfer model BERT pays more attention to the marked keywords, enhancing the model's feature extraction capability, and making the question and text paragraph have a deeper semantic fusion in the calculation stage; after BERT encoding calculation, the [CLS] tag output by the model integrates the semantic information of each word in the input question and text paragraph. Using this tag, this method uses a fully connected layer and a sigmoid function to measure the similarity between the question and the text paragraph. As shown in Formula (2), the calculation formula of the sigmoid function is that the closer the final output value is to 1, the higher the probability that it is the correct answer.

[0054]

[0055] In order to illustrate the effect of the present invention, the Stanford question-answering dataset in the experiment of the present invention uses EM and F1 as evaluation indicators. Among them, EM is the exact matching result, that is, the answer given by the model is exactly the same as the standard answer; F1 is the fuzzy matching result, which can be understood as the machine answering part of the content correctly, which is calculated based on the answer given by the model and the standard answer. Among them, the higher the EM value and F1 value, the better the performance of the model.

[0056] In order to verify the effectiveness of the answer extraction method of the question-answering system based on deep model transfer learning, a simple deep transfer model (pre-trained-fine-tuned model BERT-base-uncase-english) was used to perform the answer extraction task on the Stanford question-answering dataset. The experimental evaluation indicators EM and F1 reached 0.808 and 0.883. The example flowchart of the present invention using only the deep transfer model for the answer extraction task is as follows: Figure 3 shown.

[0057] Based on the existing experimental results using a simple deep transfer model, it can be seen that the present invention combines the deep transfer model, question classification processing and labeled text segments to make better use of text information, and fully utilizes the combination of transfer learning and deep learning. It can more accurately obtain the correct answer and improve the accuracy of the answer extraction task.

[0058] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A question-answering system answer extraction method based on deep model transfer learning, Features: First, the data set is preprocessed. When the text paragraphs and questions enter the deep transfer model to calculate the vector representation, the questions in the data set are classified by the question classification method of question words to obtain the expected answer type corresponding to the question. According to the classification information of the question, the text paragraphs are processed and the entities corresponding to the expected answer type in the question are marked. Secondly, the processed questions and text paragraphs are input into the deep transfer model. The self-attention mechanism of the deep transfer model is used to focus on the marked text paragraphs and calculate the semantic similarity between the text paragraphs and the questions. Finally, the final probability is calculated in the similarity layer through the fully connected layer and sigmoid normalization. The specific steps of the method are as follows: Step 1: Preprocess the questions and text paragraphs in the data set. Preprocessing includes classifying the questions according to question words to obtain the classified questions, so as to determine the expected answer type of the questions. Then, the entities of the text paragraph are marked according to the expected answer type obtained by question classification to obtain a marked text paragraph; Step 2: The classified questions and marked text paragraphs in Step 1 are sent as input to the deep transfer model. The encoder is used in the deep transfer model BERT to process them, unify the text length, and generate question feature vectors and candidate text segment feature vectors. Step 3: Through the self-attention mechanism of the deep transfer model BERT, focus on the marked text paragraphs, calculate the semantic similarity between the marked text paragraphs and the question, and obtain the best answer text; Step 4: Calculate the final probability of the answer and the best answer text through the fully connected layer and sigmoid normalization in the similarity layer; The Step 2 includes: The classified question and marked text paragraphs in Step 1 are sent as input to the deep transfer model BERT. The deep transfer model BERT is used to unify the text length of the input and generate an attention masking matrix attention_mask to obtain the question feature vector and the candidate text segment feature vector; When unifying the text length, short sentences are padded to a fixed length, and long sentences are truncated to a fixed length. The matrix value of the padded part of the text is 0, and the matrix value of the original part of the text is 1. In this way, when calculating the self-attention mechanism, the deep transfer model BERT will only focus on the input part, and reduce the attention weight of the padded part to almost 0; In Step 3, the self-attention mechanism of the deep transfer model BERT includes a multi-head attention module, a forward neural network, and a layer normalization module; after obtaining the vector representation sequence X of the question q and the text paragraph t q and X t Finally, they are concatenated to obtain the input sequence X of the deep transfer model BERT, and X is input into the multi-head attention module to obtain the context information of each word in the input question and text paragraph. Finally, the output of the final self-attention structure is obtained through a feedforward neural network and layer normalization.

2. According to claim 1, the method for extracting answers to a question-answering system based on deep model transfer learning, Features: The specific steps of Step 1 are: Step 1.1: Detect the question word type in the question sentence and classify the question sentence to obtain the classified question sentence, so as to determine the expected answer type of the question sentence; Step 1.2: Based on the expected answer type of the question, use Spacy NER to detect the relevant entity types in the text paragraph, highlight the entities in the text paragraph that contain the expected answer type of the question, reduce the influence of irrelevant entities in the answer sentence, and obtain the marked text paragraph; Step 1.3: Finally, the processed question and text paragraph are input into the deep transfer model BERT to calculate the semantic similarity.