Candidate Passage Generation Based on Text Classification and Multi-hop Question Answering Method

Through the combination of text classification and intermediate hop inference, the single hop problem generator is trained using the ready-made data set, which solves the problems of high manual annotation requirements and label noise in the multi-hop question answers, and achieves a more accurate and robust multi-hop inference process.

CN115878794BActive Publication Date: 2025-07-18ZHEJIANG ZHELIXIN CREDIT INVESTIGATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211229355.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2025-07-18
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

The prior art has problems in the multi-hop question answering the problem that manual annotation needs are high and pseudo-supervision may introduce label noise, and the existing methods fail to effectively utilize the supporting facts of each reasoning step, resulting in inaccurate answers.

Method used

The candidate paragraph generation method based on text classification is adopted, and through prompt learning and intermediate hop inference, the single hop problem generator is trained using the ready-made single hop problem dataset to generate sub-problems and perform unsupervised decomposition, and multi-hop inference is performed in combination with the unified reader model.

Benefits of technology

It improves the accuracy and robustness of multi-hop questions, reduces the need for manual labeling, avoids label noise, and enhances the prediction performance of single-hop question-and-answer model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115878794B_ABST
    Figure CN115878794B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating candidate paragraphs based on text classification and multi-hop question answering, belonging to the technical field of natural language processing. The present invention classifies paragraph texts into candidate paragraphs for the original question based on prompt language, and by providing an intermediate hop reasoner, each reasoning step generates a more accurate question decomposition based on the current supporting facts; by providing a single-hop question generator, an existing single-hop question dataset is used to train a single-hop question generator to directly generate sub-questions in an unsupervised manner, without the need for manual annotation after question decomposition, and avoiding the risk of label noise that may be introduced by pseudo-supervision; in addition, the single-hop question dataset used to train the single-hop question generator is also used as one of the samples for training the single-hop question answering model, making the data used by the single-hop question answering model and the single-hop question generator more consistent, which is beneficial to improving the prediction performance of the single-hop question answering model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a candidate paragraph generation and multi-hop question answering method based on text classification. Background Art

[0002] Multi-hop questions refer to those questions that require multi-hop reasoning on the knowledge graph to answer. For example, if you want to answer the question "Who are the directors of the movies starring Jackie Chan", you need a multi-hop reasoning path formed by multiple triples <Jackie Chan, starring, New Police Story>, <New Police Story, director, Chen Musheng> to answer it.

[0003] The multi-hop problem is a hot task in the field of natural language processing in recent years. It requires aggregating information from multiple documents and performing multi-hop reasoning to infer the answer. The current methods are mainly divided into two categories. The first category uses a one-step reader to capture the interaction between the question and the relevant context to predict the answer and supporting sentences (that is, the interaction between the question and the relevant context is captured by the pre-trained reader model for the input question, and the answer is directly output). The prediction accuracy of this type of method is not high. The second category simulates an explainable multi-step reasoning process, decomposes the multi-hop problem into multiple simple single-hop problems and solves them, but the existing methods of decomposing the problem have the following two problems:

[0004] 1. Problem decomposition is highly dependent on manual annotation or automatically constructed pseudo-supervision. The former requires a lot of time for manual annotation, while the latter may introduce label noise.

[0005] 2. Single-hop question generation is based only on the original question without considering the supporting facts involved in each jump reasoning step, which often leads to misguided decomposition and inaccurate interpretation, thereby predicting inaccurate answer to the question.

[0006] In addition, answering multi-hop questions requires aggregating information from multiple documents (candidate paragraphs). The relevance of the aggregated documents to answering multi-hop questions is an important prerequisite for ensuring the accuracy of answering multi-hop questions. Therefore, how to quickly and accurately screen out candidate paragraphs from numerous paragraphs has become a technical problem that needs to be urgently solved in answering multi-hop questions. Summary of the invention

[0007] The present invention provides a candidate paragraph generation and multi-hop question answering method based on text classification. First, the idea of prompt learning is used to quickly and accurately extract candidate paragraphs required for answering multi-hop questions from numerous paragraphs. Then, by providing an intermediate jump reasoner, each reasoning step is based on the current supporting facts, resulting in a more accurate problem decomposition, thereby making the entire multi-hop reasoning process more accurate and more robust. By providing a single-hop question generator, a single-hop question generator is trained using an existing single-hop question data set, and sub-questions are directly generated in an unsupervised manner. It is no longer necessary to manually label the problem after decomposition, and the risk of label noise introduced by pseudo-supervision is avoided. In addition, the single-hop question data set used to train the single-hop question generator is also used as one of the samples for training a single-hop question answering model, so that the single-hop question answering model and the data used by the single-hop question generator are more consistent, which is conducive to improving the prediction performance of the single-hop question answering model.

[0008] To achieve this object, the present invention adopts the following technical solutions:

[0009] A method for generating candidate paragraphs and answering multi-hop questions based on text classification is provided, and the steps include:

[0010] S1, extract the original question Keywords in and tag them ;

[0011] S2, for a given paragraph text , using template functions Will Convert to language model Input , In the original text of the paragraph A hint language for classification tasks is added, which contains the mask position that needs to be predicted and filled in with the label;

[0012] S3, the language model Predict the label that fills the mask position ;

[0013] S4, label converter The label Mapped to a set of tag words in a pre-built tag system The corresponding label words in The paragraph text obtained as prediction Type;

[0014] S5, judging the label word With the label Is it consistent?

[0015] If so, use the paragraph text as a candidate paragraph for answering the original question and add it to the candidate paragraph set.

[0016] If not, filter out the paragraph text ;

[0017] S6. Input the original question into a pre-trained paragraph ranking model to calculate the probability scores representing the relevance of each candidate paragraph in the candidate paragraph set to answering the original question . Then select the candidate paragraphs with the top scores and the jump paragraphs linked to the candidate paragraph ranked first as the relevant context for answering the original question , denoted as ;

[0018] S7. Input the original question , the relevant context and the sub-question-answer pair obtained from the previous intermediate jump into a unified reader model that is iteratively updated and trained with the input-output data of each jump for intermediate jump answer reasoning, and output the sub-question-answer pair corresponding to the current intermediate jump and the single-hop supporting sentence ;

[0019] S8. Use the sub-question-answer pair output from the previous jump of the final jump , the original question , the relevant context and the preset answer type as the input for the unified reader model for final jump answer reasoning, and output the multi-hop question answer corresponding to the original question and the multi-hop supporting sentence .

[0020] Preferably, the method steps for training the language model in training step S2 include:

[0021] A1. For each used as a training sample, calculate the probability score of each tag word in the tag word set being filled in the masked position , and the calculation method is expressed by the following formula (1):

[0022]

[0023] A2. Calculate the probability distribution through the softmax function , , and the calculation method is expressed by the following formula (2):

[0024]

[0025] In formulas (1)-(2), represents the label of the said label word ;

[0026] represents the label set of the text classification task.

[0027] A3. According to and , and using the constructed loss function, calculate the model prediction loss, and the constructed loss function is expressed by the following formula (3):

[0028]

[0029] In formula (3), represents the fine-tuning coefficient;

[0030] represents the distribution predicted by the model and the gap between the true distribution;

[0031] represents the score predicted by the model and the gap between the true score;

[0032] A4. Judge whether the termination condition of the model iterative training is reached,

[0033] If so, terminate the iteration and output the said language model ;

[0034] If not, adjust the model parameters and return to step A1 to continue the iterative training.

[0035] Preferably, the said language model is a fusion language model formed by fusing a number of language sub-models , and the method for training the fusion language model includes the steps:

[0036] B1. Define a set of template functions , and the set of template functions contains a number of different said template functions ;

[0037] B2. For each of the training samples, , through the corresponding language sub-model , calculate each label word in the set of label words and fill in the probability score of the masked position . , The calculation method is expressed by the following formula (4):

[0038]

[0039] B3. For each of the associated template functions , are fused to obtain , which is calculated by the following formula (5):

[0040]

[0041] In formula (5), represents the number of the template functions in the set of template functions ;

[0042] represents the weight of the template function in the calculation of ;

[0043] B4. Calculate the probability distribution through the softmax function, which is calculated by the following formula (6):

[0044]

[0045] In formulas (5) and (6), represents the label in the label set that has a mapping relationship with the label word ;

[0046] represents the label set of the text classification task;

[0047] B5. According to and , and using the constructed loss function, calculate the model prediction loss. The constructed loss function is expressed by the following formula (7):

[0048]

[0049] In formula (7), represents the fine-tuning coefficient;

[0050] Represents the distribution predicted by the model The gap with the true distribution;

[0051] Represents the score predicted by the model The gap with the true score;

[0052] B6. Determine whether the termination condition for the iterative training of the model is reached

[0053] If so, terminate the iteration and output the fused language model;

[0054] If not, adjust the model parameters and return to step B2 to continue the iterative training.

[0055] Preferably, the language model or the language sub-model is a BERT language model.

[0056] Preferably, the fine-tuning coefficient .

[0057] Preferably, in step S6 .

[0058] Preferably, the unified reader model In each intermediate or final jump, identify the single-hop support sentence of the current jump through the following method steps :

[0059] C1. Combine the input original question , the relevant context and the sub-question-answer pair formed by the previous jump into a connection sequence expressed by the following expression (9):

[0060]

[0061] In the above expression, represents the connection sequence representation input to the single-hop support sentence recognizer in the jump;

[0062] represents the jump;

[0063] represents the delimiter of a certain candidate paragraph in the relevant context selected in step S1;

[0064] Represents the multi-hop question of the original input;

[0065] Represents the sub-question generated by the

[0066] Represents the answer obtained by solving the sub-question generated by the jump;

[0067] Represents the th sentence in the

[0068] th text paragraph in the candidate paragraph;

[0069] Represents the number of sentences in the

[0070] C2, based on the special markings of each sentence representation, construct a binary classifier to predict the probability that each sentence is a supporting fact for the current jump, and sentences with a probability value greater than are used as single-hop supporting sentences for the current jump, constituting ; ;

[0071] C3, by minimizing the binary cross-entropy loss function, optimize the unified reader model used for all jumps The binary cross-entropy loss function is expressed by the following formula (8):

[0072]

[0073] In formula (8), represents the binary cross-entropy loss function adopted by the unified reader model used in the jump; ;

[0074] represents whether the sentence is a label for the supporting fact of the jump;

[0075] represents the total number of sentences in the relevant context.

[0076] Preferably, .

[0077] Preferably, the sub - problems of the current jump are generated through the following method steps:

[0078] D1, extract the overlapping words of the single - jump support sentences identified in the current jump and the original problem ;

[0079] D2, add each of the extracted overlapping words to the single - jump support sentence ;

[0080] D3, use the single - jump support sentences added with the overlapping words as the input of a pre - trained single - jump problem generator, and the single - jump problem generator generates the sub - problems of the current jump decomposition according to the input . .

[0081] Preferably, use the single - jump support sentences identified in the current jump and the single - jump sub - problems generated in the current jump as the input of a pre - trained single - jump Q&A model, predict and output the single - jump answer corresponding to the single - jump sub - problem , and the samples for training the single - jump problem model are the single - jump sub - problems generated for each intermediate jump and the single - jump problem dataset used when training the single - jump problem generator.

[0082] The present invention has the following beneficial effects:

[0083] 1. By adding hint language for the classification task in the paragraph text, the hint language contains mask positions where labels need to be predicted and filled, converting the paragraph text classification problem into a classification prediction problem similar to cloze test, simplifying the process of paragraph text classification prediction, and being able to more accurately analyze the paragraph text from the perspective of the matching relationship between the type of the paragraph text and the labels of the keywords in the original problem , and mining deeper information, improving the accuracy of paragraph text classification.

[0084] 2. By providing an intermediate - jump reasoner, each reasoning step is based on the current supporting facts, resulting in a more accurate problem decomposition, thus making the entire multi - jump reasoning process more accurate and robust.

[0085] 3. By providing a single-hop question generator, an existing single-hop question dataset is utilized to train a single-hop question generator, which directly generates sub-questions in an unsupervised manner, eliminating the need for manual annotation of question decomposition and avoiding the risk of label noise introduced by pseudo-supervision.

[0086] 4. The single-hop question dataset used to train the single-hop question generator is used as one of the samples for training the single-hop question answering model, making the data used by the single-hop question answering model more consistent with that of the single-hop question generator, which is beneficial to improving the prediction performance of the single-hop question answering model. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0088] Figure 1 is a flowchart of the implementation steps of a multi-hop question answering method based on text classification provided by an embodiment of the present invention;

[0089] Figure 2 is a comparative example diagram of the effects of the existing method and the method provided by the present application for decomposing multi-hop questions into multiple simple single-hop questions and solving them;

[0090] Figure 3 is a logical reasoning diagram of a multi-hop question answering method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and through specific embodiments.

[0092] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams, rather than physical diagrams, and should not be construed as a limitation of this patent; for better illustration of the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.

[0093] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if terms such as "upper", "lower", "left", "right", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0094] In the description of the present invention, unless otherwise clearly specified and limited, if terms such as "connection" are used to indicate the connection relationship between components, this term should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection, or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0095] The method for generating candidate paragraphs based on text classification and multi-hop question answering provided by the embodiments of the present invention, as Figure 1 , includes the steps:

[0096] S1, extract the keywords in the original question and label them. For example, for the original question "In which township does person a offer courses?",

[0097] extract its keywords such as "person a" and / or "course", and label the keywords "person a" and "course" with, for example, "literature"; It should be noted here that there are many existing methods for extracting keywords from the original question

[0098] and labeling the keywords, so no specific description will be given. S2, for the given paragraph text use the template function to convert it into the input of the language model , add the prompt language of the classification task to the original paragraph text , and the prompt language contains the masked positions where the labels need to be predicted and filled;

[0099] S3, the language model Predict the label for the masked position ;

[0100] S4, label converter Map the label to the set of label words in a pre-constructed label system corresponding label word as the predicted paragraph text type;

[0101] It should be noted here that in this embodiment, the technical core of classifying the input paragraph text is to adopt the idea of prompt learning. Prompt learning can simplify the classification process and improve the classification efficiency, and has higher classification superiority for small-scale data sets. Specifically, in order to give full play to the powerful question-answering and reading comprehension capabilities of the paragraph text classifier, the input paragraph text is processed according to a specific pattern, and task prompt language is added to it to make it more suitable for the question-answering form of the language model. The principle of paragraph classification based on the prompt learning-based paragraph text classifier is as follows:

[0102] Let be a pre-trained language model (preferably a BERT language model), is the set of label words in a pre-constructed label system, and the masked word is used to fill the content of the masked position covered in the input of the language model and let be the set of labels for the text classification task (paragraph text classification task). After tokenizing each paragraph text, the word sequence input to the language model is obtained, and then a custom template function is used to convert into the input of the language model In the classification task prompt language is added, and the prompt language includes the masked position where the label needs to be predicted and filled. After conversion, the paragraph text type prediction problem can be converted into a cloze problem, that is, the language model takes the cloze problem in the form as the input, and the most suitable word predicted to fill the masked position is used as the classification prediction result of the paragraph text expressed.

[0103] It should be emphasized that this application is based on the idea of prompt learning and makes better use of the language model The Q&A and reading comprehension capabilities. At the same time, since the classification problem is converted into a cloze problem, the prediction process is simpler, improving the classification efficiency of the policy text classifier. Further, in this embodiment, a mapping from the label set of the text classification task to the label word set in the pre-constructed label system is defined as the label converter . For example, for the label , this label converter maps it to the label word

[0104] For each template function and label converter , this embodiment classifies the paragraph text through the following steps:

[0105] Given an input paragraph text (preferably the original paragraph text word sequence), use the template function to convert into the input of the language model . The language model will predict the most suitable label for the masked position in . Then, use the label converter to map this label to the label word in the policy document element system , and use it as the classification of the paragraph text . Preferably, this embodiment uses a pre-trained Chinese BERT model as the language model . Its prediction method for the masked position follows the pre-training task of the BERT model, that is, it uses the output corresponding to the masked position in to predict the label of the masked position (the prediction method is the same as the Masked Language Model pre-training task of the BERT model and will not be elaborated).

[0106] For example, regarding the template function , assume that is defined as " . Generally speaking, this is a text paragraph about _____.", where "_____" represents the masked position, thus adding a prompt language for the classification task to the original paragraph text . For example, it is "In which township does person a offer courses?", for this paragraph text , after adding the above prompt language, the language model The classification task is to predict the label of the mask position "_____" in "In which town does person a offer courses? Overall, this is a text paragraph about _____." After predicting the label behind the mask position, the predicted label Mapped to a set of tag words in the tag system The corresponding label words in As the predicted paragraph text Type.

[0107] The following is the training language model for this embodiment The method is described:

[0108] Language Model The BERT model is preferably used. There are many existing training methods for the BERT model, which can be applied to this application to train the language model. , the difference is that this embodiment is used to train the language model The sample is a template function The converted And the label converter The converted label word set The corresponding label words in , and the loss function for evaluating model performance improved by the present application to improve classification accuracy.

[0109] Training the language model In this application, the sample data set is randomly divided into a training set and a validation set in a ratio of 7:3. The training process is as follows:

[0110] For each paragraph text, a sequence of only one mask position is generated , for the tag word set in the tag system Each tag word in The probability of filling in the mask position calculates a score (due to the label In the tag word collection There is a label word with a mapping relationship in , so the predicted label Filling in the probability score of the mask position is equivalent to predicting the corresponding label word The probability score of filling in the mask position), which is determined by the language model Prediction represents the predicted probability that the label word can be filled in the mask position. More specifically, for a sequence , this application calculates the label set of the text classification task Tags in The method for filling in the probability score at this mask position is expressed by the following formula (1):

[0111]

[0112] In formula (1), represents the probability score of filling the label at the mask position. Since the label has a mapping relationship with the set of label words in the label system the corresponding label word in it, so is equivalent to representing the probability score of filling the label word at the mask position;

[0113] , for example, the label of the label word "military" can be mapped to , and the label of the label word "humanities" can be mapped to . By establishing such a mapping relationship, the task is changed from assigning a meaningless label to an input sentence to selecting the word most likely to fill the mask position.

[0114] After calculating the scores of all label words in filling the same mask position, a probability distribution is obtained through the softmax function. The specific calculation method is expressed by the following formula (2):

[0115]

[0116] In formula (2), represents the set of labels for the text classification task;

[0117] Then, according to and , and using the constructed loss function, the model prediction loss is calculated. The constructed loss function is expressed by the following formula (3):

[0118]

[0119] In formula (3), represents the fine-tuning coefficient (preferably 0.0001);

[0120] represents the gap between the distribution predicted by the model and the true one-hot vector distribution;

[0121] represents the gap between the score predicted by the model and the true score;

[0122] Finally, determine whether the termination condition for model iterative training is reached.

[0123] If so, terminate the iteration and output the language model. ;

[0124] If not, adjust the model parameters and continue with the iterative training.

[0125] To further improve the model training effect and thus enhance the classification performance of the language model, preferably, the language model is a fused language model formed by fusing a number of language sub-models . The method for training the fused language model is as follows:

[0126] First, define a set of template functions . The set of template functions contains a number of different template functions . For example, the template function is " . What is this text passage related to? _____", and another example, the template function is "What is this text passage related to? It is related to _____", and so on. For different template functions , in this embodiment, the following method is used to train the fused language model:

[0127] For each used as a training sample, calculate the probability score of each label word in the label word set filled in the mask position through the corresponding language sub-model . The calculation method is expressed by the following formula (4):

[0128]

[0129] In formula (4),

[0129] ;

[0130] Then, fuse associated with each template function to obtain , which is calculated by the following formula (5):

[0131]

[0132] In formula (5), represents the number of template functions in the set of template functions ;

[0133] Indicates a template function The weight occupied during the calculation ;

[0134] Then, calculate the probability distribution through the softmax function , and the calculation method is expressed by the following formula (6):

[0135]

[0136] In formula (2), Indicates the label set of the text classification task. Finally, according to and , and using the constructed loss function, calculate the model prediction loss, and the constructed loss function is expressed by the following formula (8):

[0137]

[0138] In formula (8), Indicates the fine-tuning coefficient (preferably 0.0001);

[0139] Indicates the distribution predicted by the model The gap between the true distribution;

[0140] Indicates the score predicted by the model The gap between the true score.

[0141] After predicting the type of the paragraph text , as shown in Figure 1 , transfer to the step:

[0142] S5, determine whether the label word is consistent with the label ,

[0143] If so, add the paragraph text as a candidate paragraph for answering the original question to the candidate paragraph set,

[0144] If not, filter out the paragraph text ;

[0145] After obtaining the candidate paragraph set of the original question , enter the multi-hop question answering session, that is, transfer to Figure 1 the steps shown in

[0146] S6, the original question Input into a pre-trained passage ranking model to calculate the probability scores representing the relevance of each candidate passage to answering the original question and then select the top ( preferably equal to 3, when adding the jump passage linked to the top-ranked candidate passage, that is, the relevant context of the original question ), the candidate passages and the jump passage linked to the top-ranked candidate passage, which are relevant to answering the original question There are 4 candidate passages. Since in each inference step of the intermediate jump in this embodiment, it is based on the current supporting facts, a more accurate question decomposition is generated. Therefore, compared with the existing single-hop question decomposition method in the background art that is only based on the original question without considering the current supporting facts for each decomposition step, the value of can be smaller. Through repeated comparison of experimental data, when it has almost no impact on the accuracy of answering multi-hop questions, but due to the decrease in the value of , the overall speed of answering multi-hop questions is significantly improved. In addition, this application adds the jump passage linked to the top-ranked candidate passage to the relevant context of the original question ;

[0147] S7. Input the original question and the relevant context into a pre-trained unified reader model (also known as the intermediate jump reasoner) to perform intermediate jump answer reasoning, and output the sub-question-answer pair corresponding to each intermediate jump and the single-hop supporting sentence ;

[0148] S8. Use the sub-question-answer pair output by the previous hop of the final hop, the original question , the relevant context and the answer type as the input of the unified reader model to perform final hop answer reasoning, and output the multi-hop question answer corresponding to the original question and the multi-hop supporting sentence .

[0149] The following is combined with Figure 2 , Figure 3A detailed description of the specific implementation method for multi-hop question answering is as follows:

[0150] As Figure 2 shown, for example, for the multi-hop question "In which township does person a offer courses?" (i.e., the original question ), according to the multi-hop question decomposition method described in the background art, based only on the original question without considering the supporting facts involved in each jump reasoning step, this multi-hop question may be decomposed into Sub-Q1: Where does person a offer courses and Sub-Q2: In which township is person a. However, through the method provided in this application, this multi-hop question is decomposed into Step1-Q: In which manor does person a offer courses and Step2-Q: Which township does the Bhaktivedanta Manor belong to, as well as identifying single-hop support sentences Step1-S and Step2-S from the candidate paragraphs as the basis for generating Step1-Q and Step2-Q. Obviously, the generation of Step1-Q and Step2-Q is more likely to infer the correct answer because there is evidence to rely on (supported by Step1-S and Step2-S respectively).

[0151] In this embodiment, given an original question and a context containing multiple candidate paragraphs, the goal is to identify the context relevant to answering this original question , predict the final answer , and explain the answer with supporting sentences .

[0152] To reduce the interference of too many candidate paragraphs on question answering during the multi-hop reasoning process and improve the question answering efficiency, in this embodiment, first, the candidate paragraphs most relevant to answering the original question are screened out from all candidate paragraphs as the relevant context of the question, denoted as . The specific screening method for the relevant context is as follows: Given multiple candidate paragraphs as training samples for the paragraph ranking model, a paragraph ranking model is trained. The paragraph ranking model consists of a RoBERTa encoder and a binary classification layer. This model takes each original question and each candidate paragraph as input, and the sigmoid function in the binary classification layer outputs the probability score of each candidate paragraph being relevant to the original question . Using the correct question-related paragraphs in the training data as supervision to optimize a cross-entropy loss function can train the paragraph ranking model. Then, a two-hop selection strategy is adopted. For the first hop, in the set containing paragraphs relevant to the original question Select the candidate paragraph with the highest score from the candidate paragraphs with the same phrase, then jump to the jump paragraph linked by the wiki hyperlink embedded in the candidate paragraph with the highest score, and finally sort the jump paragraph and the probability scores in descending order before ( preferably equal to 3) of the candidate paragraphs as the context for answering the original question and denote it as 。

[0153] It should be emphasized here that step S6 of the multi-hop question answering method provided in this application, that is, finding the context of the original question is very important. In the subsequent intermediate jump reasoning and final jump reasoning, identifying the single-hop support sentences that serve as the basis for generating sub-questions in each hop, generating sub-questions for each hop, predicting the answers corresponding to the sub-questions for each hop, and outputting the answer corresponding to the final original question must all be based on this context obtained in step S6 。 Since this application incorporates the jump paragraph linked by the candidate paragraph with the highest score into the context corresponding to the original question , in the sub-question generation and answer reasoning for each hop in the intermediate jump, the influence of the second-hop paragraph (i.e., the jump paragraph) linked by the candidate paragraph with the highest score in the first hop on sub-question generation and sub-question answering is considered, making it less likely for sub-question generation and sub-question answering to deviate from the original question itself, and selecting the candidate paragraphs ranked among the top as the relevant context for this original question , considering the comprehensive influence of different candidate paragraphs on the accuracy of sub-question generation and sub-question answering, and since a limited number of relevant contexts are selected , ensuring the efficiency of sub-question generation and sub-question answering, and thus ensuring the efficiency of multi-hop question answering. Here, it should be further noted that since the specific training process of the paragraph ranking model is not within the scope of the claims of this application, the specific training process of the paragraph ranking model will not be described in detail here.

[0154] After screening out the relevant context of the original question

[0155] , the multi-hop question answering method provided in this embodiment enters the intermediate jump reasoning process. Intermediate jump reasoning means performing multi-hop reasoning step by step based on the screened relevant context 。 In this embodiment, a unified reader model is used 。 , step by step. In this embodiment, a unified reader model is adopted (i.e., the intermediate hop reasoner or the final hop reasoner) to identify the single-hop supporting sentences for each intermediate hop , and then generate and answer corresponding single-hop sub-questions according to the identified single-hop supporting sentences , and pass the original question , relevant context , and the sub-question-answer pair obtained from the current intermediate hop to the unified reader model for the question answering reasoning of the next hop.

[0156] The unified reader model adopted in this embodiment includes 3 models, namely, a single-hop supporting sentence recognizer, a single-hop question generator, and a single-hop question answering model.

[0157] The single-hop supporting sentence recognizer takes the original question , relevant context , and the sub-question-answer pair formed by the previous hop as input (when the previous hop is the first hop, since no sub-question-answer pair is generated, there is only the original question , so when the second hop is an intermediate hop, the input to the single-hop supporting sentence recognizer is the original question and relevant context ). It attempts to find a single-hop supporting sentence from the relevant context that can be used as the basis for generating the sub-question of the current hop and answering the generated sub-question . Specifically, the concatenation sequence of the original question , relevant context , and the sub-question-answer pair of the previous hop input to the single-hop supporting sentence recognizer is expressed by the following expression (9):

[0158]

[0159] In expression (9), represents the concatenation sequence representation input to the single-hop supporting sentence recognizer in the th hop;

[0160] represents the th hop;

[0161] represents the delimiter in the relevant context selected in step S1 for a certain candidate paragraph, and what follows represents a paragraph, for example that constitutes a certain candidate paragraph in the relevant context ;

[0162] Represents the multi-hop problem of the original input;

[0163] Represents the sub-problem generated by the

[0164] Represents the answer obtained by solving the sub-problem generated by the th hop;

[0165] Represents the th sentence in the

[0166] Represents the number of text paragraphs in the candidate paragraph;

[0167] Represents the th text paragraph in the candidate paragraph. Then, based on the representation of each sentence with special tokens a binary classifier is constructed to predict the probability that each sentence is a supporting fact for the current hop . Predicting the probability that each sentence is a supporting fact for the current hop can adopt existing supporting fact prediction methods, so the specific calculation method for is not elaborated here;

[0168] Finally, by minimizing the binary cross-entropy loss function, the unified reader model used for the th hop is optimized. The binary cross-entropy loss function is expressed by the following formula (10):

[0169]

[0170] In formula (10), represents the binary cross-entropy loss function adopted by the unified reader model used in the th hop; represents the label indicating whether the sentence

[0171] is a supporting fact for the th hop; represents the total number of sentences in the relevant context

[0172] represents the relevant context .

[0173] Identify the current Single-hop support sentence for jumping After that, enter the generation process of the sub-jumping problem. This application does not use manual annotation or pseudo-supervised methods to train the single-hop question generation model, but directly uses a ready-made single-hop question corpus to pre-train a single-hop question generator to generate the sub-question of the current hop according to the single-hop support sentence identified in the current hop and the original question to generate the sub-question of the current hop . Specifically, first extract the single-hop support sentence identified in the current hop and the original question overlapping words, and then add the overlapping words to the single-hop support sentence (for example, concatenated in front of the original single-hop support sentence ), and then use the single-hop support sentence with the overlapping words added as the input of the pre-trained single-hop question generator (the input form is expressed as [CLS] [SEP] [SEP], for example Figure 3 in is "In which township does person a offer courses?", is "Person a is a leader in the b field, offers courses at the Bhaktivedanta Manor, and lectures on his own comments on domestic and international events.", then is "Person a's courses", and the single-hop question generator generates the sub-question decomposed by the current hop

[0174] It should be noted here that adding the overlapping words to the single-hop support sentence is beneficial to guiding the generation of sub-questions to be more in line with the reasoning goal of the original question . Another thing to note is that since the specific training method of the single-hop question generator is not within the scope of the claims of this application, the specific training process will not be described

[0175] Generate the sub-question of the current hop After that, this application uses the single-hop support sentence and the generated single-hop sub-question as the input of the pre-trained single-hop question answering model, and predicts and outputs the single-hop answer corresponding to the single-hop sub-question . It should be noted here that in order to improve the single-hop question answering model's prediction of the single-hop answer The accuracy rate. When training the single-hop Q&A model, one of the samples used is the single-hop question dataset also used when training the single-hop question generator. Since the same single-hop question dataset is used for both training the single-hop Q&A model and the single-hop question generator, it ensures the data consistency of some training samples, reduces the noise error caused by introducing inconsistent samples, and the prediction accuracy is higher.

[0176] It should be noted here that for the single-hop supporting sentence and the single-hop sub-question The training of the single-hop Q&A model with these as samples can be obtained based on existing training methods. And since the specific training process of the single-hop Q&A model is not within the scope of the claims of this application, the training process of the single-hop Q&A model will not be elaborated specifically.

[0177] After completing several intermediate hops, it enters the process of generating the answer to the multi-hop question and identifying the multi-hop supporting sentence in the last hop (the final hop ). Specifically, as Figure 3 shown, in the final hop, the sub-question-answer pair of the previous hop (i.e., the last hop of the intermediate hops) is used to build a bridge between the intermediate hop and the final hop, and then the same unified reader model used in the intermediate hop reasoning process is used to predict the final answer to the original question , and at the same time provide the multi-hop supporting sentence as the basis for answering the original question . As Figure 3 shown, the connection sequence input to the unified reader model in the final hop is:

[0178]

[0179] Comparing expression (9) and expression (11), it can be seen that in the final hop, two additional tags yes or no are inserted into the connection sequence input to the unified reader model before the relevant context for answer prediction. In this embodiment, there are 2 types of answer types corresponding to the original question , which are: yes, no. Yes means the answer type of the original question is yes; no means the answer type of the original question is no. For example, for the original question "Is the first Chinese athlete to win an Olympic gold medal Xu XX?", its answer type is "yes".

[0180] To complete the last-hop reasoning, first use a binary classifier to identify the relevant context Whether each sentence in is a supporting fact for the entire multi-hop question (i.e., the original question ), calculate the loss of identifying the supporting sentence through the loss function and then make a prediction of the final answer segment. The prediction method is as follows: Add a linear layer with a softmax function on all context representations (Softmax is a function for calculating probabilities, which can calculate the probability that each character is the start position or the end position of the answer on the representations of all characters in the relevant context (i.e., the relevant context ), to obtain the probability that each (i.e., the th character in the relevant context ) is the start position of the answer or the probability of being the end position , and denote the maximum probabilities of being the start position of the answer and being the end position of the answer as and respectively, and then obtain the content between the positions where and

[0181] are located as the answer to the multi-hop question of the final prediction output For the prediction loss of the start position and the end position of the multi-hop question answer in the relevant context, it is calculated by the following formula (12):

[0182]

[0183] To improve the training speed and model performance of the unified reader model , the present invention also specifically constructs a joint loss function for the unified reader model . The constructed joint loss function is expressed by the following formula (13):

[0184]

[0185] In formula (13), represents the joint loss function;

[0186] represents the binary cross-entropy loss function adopted by the intermediate hop inference in the intermediate hop of the th hop;

[0187] represents the binary cross-entropy loss function adopted by the final hop inference in the final th hop;

[0188] represents inferring the original question The corresponding multi-hop question answer The total number of hops required;

[0189] 、 respectively represent 、 the weighted hyperparameters when participating in constructing the joint loss function;

[0190] represents the prediction loss of the starting and ending positions of the multi-hop question answer corresponding to the original question by the final-hop reasoner in the relevant context ;

[0191] or is expressed by the following formula (14):

[0192]

[0193] In formula (14), represents the binary cross-entropy loss function adopted by the unified reader model used in the th hop, when it represents that the current hop is the intermediate hop of the th hop, when it represents that the current hop is the final hop of the th hop;

[0194] represents whether the th sentence in the th paragraph of the relevant context is the label of the supporting fact for the th hop; ;

[0195] represents the total number of sentences in the relevant context ;

[0196] is expressed by the following formula (15):

[0197]

[0198] In formula (15), 、 respectively represent the maximum probabilities of the answer start position and answer end position of the label content extracted from the relevant context as the multi-hop question answer to the original question ;

[0199] Method for training a unified reader model with a joint loss function is as follows:

[0200] Using the joint loss function as the loss function adopted when training the unified reader model and using the sub-question answer pairs obtained at each intermediate hop ( representing the sub-questions decomposed at each intermediate hop, being the sub-questions predicted for each intermediate hop corresponding answers), the original question the relevant context and the preset answer type as joint training samples, and jointly training to obtain a unified reader model ;

[0201] Then, input the original question and the relevant context into the unified reading model to perform intermediate hop and final hop answer reasoning, and finally output the multi-hop question answer corresponding to the original question and the multi-hop supporting sentences .

[0202] To verify the performance of the unified reader model trained by the joint optimization method of the present application , the present application evaluated the model performance using HotPotQA as the question-answering data set. The evaluation process requires answering questions and predicting supporting facts to explain the reasoning at the same time. It includes two benchmark settings: Distractor (finding answers given 10 paragraphs) and fullwiki (not given paragraphs, need to retrieve relevant paragraphs in wiki to find answers). The present application focuses on the Distractor setting to mainly test the multi-hop reasoning ability while ignoring the information retrieval part. The data set consists of 90447, 7405, and 7405 data points in the training set, development set, and test set respectively. Each instance has 10 candidate paragraphs, and only two paragraphs contain the necessary sentences to support the question. In terms of automatic evaluation, exact match (EM) and F1 of answer prediction, support fact prediction, and their combination are used as metrics. In addition, to train the single-hop question generator and the single-hop question-answering model, SQuAD is used as the single-hop question corpus.

[0203] In an embodiment, ELECTRA large is used as the main model for the step-by-step reasoning method and the single-hop question answering model, and BART-large is used to train the single-hop question generator. All these models are implemented using Huggingface. The training batch size used is 48, and fine-tuning is performed for 10 epochs. Adam is used as the optimizer with a learning rate of 3e-5. This application uses a linear learning rate with a 10% warm-up ratio. The hyperparameters for balancing the loss weights are selected as = 10 and = 5.

[0204] This application conducts a performance comparison between the unified reader model trained by the joint training method and the current state-of-the-art multi-hop question answering reasoning models (including models based on question decomposition and models based on one-step readers), and the comparison results are shown in Table 1 below. Compared with the previous models based on question decomposition (DecompRC and ONUS in Table 1) and models based on one-step readers (TAP2~HGN in Table 1), it can be seen from Table 1 that the unified reader model (StepReasoner) proposed in this application has significantly improved in answer prediction, supporting sentence prediction, and joint scores.

[0205]

[0206] Table 1

[0207] Meanwhile, in this scenario example, an ablation experiment is conducted on the joint training method of the model proposed in this application, and the experimental results are shown in Table 2 below. In Table 2, w / o represents without. In the method of w / o joint training, joint optimization is not used, and the pipeline reasoning model is directly used. In the methods of w / o bias.supp and w / o bias.ques, two components for reducing exposure bias are not used, which are used to reduce the inconsistency of single-hop supporting sentences and single-hop sub-questions between training and testing respectively.

[0208]

[0209] Table 2

[0210] It can be seen from Table 2 that overall, using all three components simultaneously can achieve better results. Joint optimization of the unified reader model for all hops can improve the tolerance to intermediate errors and improve the reasoning performance. After not using any measures to mitigate exposure bias, the effect also drops significantly, indicating that these two measures for reducing the training-test differences of single-hop supporting sentences and single-hop questions both have better generalization ability.

[0211] This application also compares the robustness of the unified reader model trained by using existing pre-trained models with existing methods and the unified reader model trained by the joint training method provided by this application The comparison results are shown in Table 3 below. In Table 3, the models trained by existing methods include BERT-base uncased, ELECTRA-large, and ALBERT-xxlarge-v2. It can be seen that these existing pre-trained models are used as the initial models, and the performance of the models trained by the joint training method provided by this application (denoted as "StepReasoner-BERT", "StepReasoner-ELECTRA", and "StepReasoner-ALBERT" in Table 3) has been improved, especially in terms of the EM score. This indicates that the unified reader model trained by the joint training method proposed in this application is more robust and effective for training based on various pre-trained models.

[0212]

[0213] Table 3

[0214] The unified reader model trained by the joint training method of this application For the comparison of the inference effects of different inference types in multi-hop reasoning, please refer to Table 4 below. Table 4 includes four inference categories: "Bridge", "Implicit-Bridge", "Comparison", and "Intersection" ("Bridge": bridge problem, which requires inferring an explicit intermediate bridge entity first and then finding the answer to the question; "Implicit-Bridge": implicit bridge problem, which requires inferring an implicit intermediate bridge entity first and then finding the answer to the question; "Comparison": comparison problem, which requires comparing the attributes of two entities; "Intersection": intersection problem, which requires finding an answer that satisfies multiple attributes / constraints simultaneously). It can be seen that the multi-hop question-answering reasoning method provided by this application is effective for different inference types, especially for "Implicit-Bridge" and "Intersection", because it is easier to obtain wrong answers to these two types of questions by directly identifying entities that satisfy a query attribute from a single piece of evidence and ignoring multi-hop reasoning involving other evidence, thus obtaining a quick solution. This observation also verifies the effectiveness of the single-hop questions that gradually generate interpretable multi-hop reasoning based on intermediate single-hop supporting sentences provided by this application.

[0215]

[0216] Table 4

[0217] To prove the effectiveness of generating single-hop questions based on the identified single-hop supporting sentences, several different single-hop question generation methods are incorporated into the step-by-step reasoning framework, and the Q&A results are compared on ELECTRA. For the comparison data of the Q&A results, please refer to Table 5 below. It can be seen that the method based on Supp has the best performance, generating more accurate and informative sub-questions based on single-hop supporting sentences, which is more effective than the single-hop questions generated by other strategies.

[0218]

[0219] Table 5

[0220] It should be noted that the above specific implementation manners are merely preferred embodiments of the present invention and the applied technical principles. Those skilled in the art should understand that various modifications, equivalent replacements, changes, etc. can be made to the present invention. However, as long as these transformations do not deviate from the spirit of the present invention, they should be within the protection scope of the present invention. In addition, some terms used in the specification and claims of this application are not restrictive, but are only for the convenience of description.

Claims

1. A method for generating candidate paragraphs and multi-hop question answering based on text classification, characterized in that the steps Including: S1, Extract the keywords in the original question and tag them and tag them ; S2. For the given passage text , use the template function to convert it into the input of the language model , and add the prompt language for the classification task to the original passage text . The prompt language contains the masked positions where the labels need to be predicted and filled in; S3, the language model predicts the label to be filled in the masked position ; S4, Label Converter Map the said label to the set of label words in a pre-constructed label system and the corresponding label word as the predicted type of the said passage text ; S5, determine whether the label word is the same as the label or not. If so, then use the paragraph text as a candidate paragraph for answering the original question and add it to the candidate paragraph set If not, filter out the paragraph text ; S6. Input the original question into a pre-trained passage ranking model to calculate the probability scores representing the relevance of each candidate passage in the candidate passage set to answering the original question . Then select the candidate passages with the top scores and the jump passages linked to the candidate passage ranked first as the relevant context for answering the original question , denoted as ; S7. Input the original problem , the relevant context , and the sub-question-answer pair obtained from the previous intermediate hop into a unified reader model that is iteratively updated with the input-output data of each hop as training samples to perform intermediate-hop answer reasoning and output the sub-question-answer pair corresponding to the current intermediate hop and single-hop supporting sentences ; S8, the sub-question-answer pair output by the previous hop of the final hop , the original question , the relevant context and a preset answer type for the unified reader model to perform answer inference for the final hop with the input, and output the multi-hop question answer corresponding to the original question and multi-hop supporting sentences . .

2. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 1, wherein The language model in training step S2 The method steps include: A1. For each of the training samples , calculate each label word in the set of label words and the probability score filled in the mask position . The calculation method is expressed by the following formula (1): A2. Calculate the probability distribution through the softmax function , , and the calculation method is expressed by the following formula (2): In Formulas (1)-(2), represents the label of the label word ; Represents the set of labels for the text classification task; A3, according to and , and using the constructed loss function, calculate the model prediction loss, and the constructed loss function is expressed by the following formula (3): In formula (3), represents the fine-tuning coefficient; Indicates the distribution predicted by the model The gap with the true distribution; The score predicted by the model The gap with the true score; A4. Determine whether the termination condition of the model iterative training is reached. If so, terminate the iteration and output the language model ; If not, adjust the model parameters and return to step A1 to continue the iterative training.

3. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 2, characterized in that The language model is a fusion language model formed by fusing a number of language sub-models The method for training the fusion language model includes the steps of: B1, define a set of template functions , the set of template functions includes several different said template functions ; B2, for each of the training samples , through the corresponding language sub-model , calculate each label word in the set of label words the probability score of filling the masked position , , The calculation method is expressed by the following formula (4): B3, for associating each of the said template functions of are fused to obtain , which is calculated by the following formula (5): In formula (5), represents the number of the template functions in the set of the template functions; representing the said template function at the time of calculation the weight occupied B4. Calculate the probability distribution through the softmax function , which is calculated through the following formula (6): In Formulas (5) and (6), represents the label set in which the label has a mapping relationship with the label word Represents the set of labels for the text classification task; B5, according to and , and using the constructed loss function, calculate the model prediction loss, and the constructed loss function is expressed by the following formula (7): In formula (7), represents the fine-tuning coefficient; Represents the distribution predicted by the model The gap between the true distribution; The score predicted by the model The gap with the true score; B6. Determine whether the termination condition of the model iterative training is reached. If so, terminate the iteration and output the fused language model. If not, adjust the model parameters and return to step B2 to continue the iterative training.

4. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 3, characterized in that, The language model or the language sub-model is the BERT language model.

5. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 2 or 3, characterized in that, Fine-tuning coefficient .

6. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 1, characterized in that, In step S6, .

7. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 1, wherein The unified reader model In each intermediate hop or final hop, the single-hop support sentence of the current hop is identified through the following method steps : C1, the original problem of the input , the relevant context and the sub-problem-answer pair formed by the previous hop are formed into a connection sequence expressed by the following expression (9): In the above expression, represents the connection sequence representation input to the single-hop support sentence recognizer in the Indicates the Jump; Indicates the relevant context in which a certain candidate paragraph is selected in step S1 Separator in Indicates the multi-hop problem of the original input; Indicates the sub-problems generated by the jump; Indicates the answer obtained by solving the sub-question generated by the jump; Indicates the th paragraph in the candidate paragraph and the th sentence; Indicates the number of text paragraphs in the candidate paragraph; Indicates the number of sentences in the text paragraphs in the candidate paragraph; C2, based on each sentence with special tokens representation, construct a binary classifier to predict the probability that each sentence is the supporting fact for the current hop, and take sentences with probability values greater than as single-hop supporting sentences for the current hop, forming ; ; C3, the unified reader model used for all hops is optimized by minimizing the binary cross-entropy loss function The binary cross-entropy loss function is expressed by the following formula (8): In formula (8), represents the binary cross-entropy loss function adopted by the unified reader model used in optimizing the Indicates a sentence Whether it is the tag that jumps to support the fact; Indicates the total number of sentences in the relevant context 8. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 7, characterized in that, 。 9. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 1 or 7, characterized in that, Generate the current sub-problem to jump: D1, extracted from the current single-hop support sentences recognized in the and the original question overlapping words; D2, adding each of the extracted overlapping words to the single-hop supporting sentence ; D3, with each of the single-hop support sentences added with the overlapping words as the input of a pre-trained single-hop question generator, and the single-hop question generator generates sub-questions for the current hop decomposition according to the input .

10. The method for generating candidate paragraphs based on text classification and multi-hop question answering according to claim 9, characterized in that, Using the single-hop support sentence recognized in the current hop and the single-hop sub-question generated in the current hop as the input of a pre-trained single-hop Q&A model, predict and output the single-hop answer corresponding to the single-hop sub-question. The samples for training the single-hop Q&A model are the single-hop sub-questions generated for each intermediate hop and the single-hop question dataset used when training the single-hop question generator. ​

Citation Information

Patent Citations

  • Construction method of question-answering system based on document set multi-hop reasoning

    CN111538819A

  • Multi-hop question answering method based on multi-hop reasoning joint optimization

    CN114780707A