A data augmentation and training method for multi-hop question answering retrieval models
By adopting a positive example denoising strategy based on propositional clauses in the multi-hop question-and-answer search model, the problem of noise problems and negative sampling strategies is solved, and the information capture accuracy and multi-hop reasoning ability of the model are improved.
Patent Information
- Application Number
- CN202411728003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-28
AI Technical Summary
There is a noise problem in the multi-hop question-and-answer retrieval model, which is difficult to filter out irrelevant information. The existing in-batch negative sampling strategy is simple and random, making it difficult to effectively learn complex logical relationships.
The positive example denoising strategy based on proposition clauses is adopted, and the proposition clauses are extracted from the document through large language models and prompt engineering technology, and the semantic similarity is calculated and sorted using BERT Score. The TopK proposition clauses are retained for splicing to form the denoised positive example text, and input them into the multi-hop question and answer pre-trained language model for training.
It significantly reduces the interference of the problem-independent information in document paragraphs, improves the accuracy and efficiency of the model to capture relevant information, and enhances the accuracy and robustness of the model in multi-hop inference tasks.
Smart Images

Figure CN119669755B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to a training method for a multi-hop question answering model, specifically to a data augmentation and training method for a multi-hop question answering retrieval model. Background Art
[0002] During the training process of the multi-hop question answering model, the core idea is to use known supporting facts to predict new supporting facts, and then train an efficient sentence encoder. This method effectively encodes the document content by converting the entire document into an embedding vector of a fixed dimension. However, when dealing with documents of variable length, one of the main challenges faced by the model is how to filter out irrelevant information, namely the so-called "noise". Since the amount of information contained in the document is huge, not all information is directly related to the training objective, so the noise problem has become a key factor affecting the model performance.
[0003] In addition, although the in-batch negative sampling strategy adopted by the current MDR (multi-hop dense retrieval) model can reduce the cost of negative sample sampling to a certain extent, due to its relatively random and simple selection mechanism, the model is difficult to effectively learn complex logical relationships, further limiting the expressiveness of the model.
[0004] To sum up, the problems existing in the prior art are as follows:
[0005] 1. The MDR algorithm for the multi-hop question answering retrieval model encounters the noise problem. One of the main challenges is how to filter out irrelevant information, that is, the "noise" information unrelated to the multi-hop question.
[0006] 2. Although the in-batch negative sampling strategy adopted by the current model can reduce the cost of negative sample sampling to a certain extent, due to its relatively random and simple selection mechanism, the model is difficult to effectively learn complex logical relationships, further limiting the expressiveness of the model. Summary of the Invention
[0007] To solve the above problems existing in the prior art, the present invention provides a data augmentation and training method for a multi-hop question answering retrieval model.
[0008] The object of the present invention is achieved by the following technical solutions:
[0009] A data augmentation and training method for a multi-hop question answering retrieval model includes the following steps:
[0010] Step 1: Obtain a multi-hop question-answering dataset, which consists of complex multi-hop questions and their corresponding document sets. These document sets include first-hop retrieval documents, second-hop retrieval documents, and other relevant documents;
[0011] Step 2: Denoise the positive examples in the first-hop documents and second-hop documents in the document set to obtain the denoised documents as new positive examples for model training, and use the remaining parts of the documents as supplementary negative examples for training. The specific steps are as follows:
[0012] Step 21: Use a large language model and prompt engineering techniques to extract propositional clauses from the first-hop documents and second-hop documents in the document set to obtain the corresponding set of propositional clause candidates;
[0013] Step 22: According to the multi-hop question, calculate the semantic similarity scores for each propositional clause in the candidate set using BERT Score and sort them in descending order to obtain the sorted set of propositional clause candidates;
[0014] Step 23: Set the retention percentage α, retain the top K propositional clauses, and splice the top K propositional clauses in the original word order to obtain the denoised positive example text;
[0015] Step 3: Input the data obtained in Step 2 into a multi-hop question-answering pre-trained language model for training. The specific steps are as follows:
[0016] Step 31: Based on the text embedding model of Bert, convert the first-hop documents, second-hop documents, and negative example documents into corresponding text embedding vectors;
[0017] Step 32: Design a multi-hop question-answering pre-trained language model
[0018] Step 321: Select the MDR model as the baseline model;
[0019] Step 322: Introduce a positive example denoising strategy based on propositional clauses. Propositional clauses are defined as atomic expressions in the text, and each proposition contains an independent fact or information point. The process of positive example denoising is as follows: First, use a proposition extraction model to identify and extract multiple proposition sets from a single text paragraph. Then, for each proposition in the set, calculate its semantic relevance score with the given question q and the relevant supporting documents;
[0020] Step 323: Introduce a hyperparameter a to control the ratio between the number of retained propositions and the total number of propositions. According to the hyperparameter a, select the propositions within this ratio range to be retained. These propositions will be re - spliced in the order they appear in the original text to maintain text coherence. In addition, to avoid losing propositions containing answer information, the propositions related to the answer are also included in the retained set. The remaining proposition set = P\ is used to generate additional negative example samples to further enrich the training data set;
[0021] Step 33: Input the data obtained in Step 31 into the multi - hop question - answering pre - trained language model designed in Step 32 for training.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] The present invention proposes a positive - example denoising strategy based on propositional clauses. Under this strategy, propositional clauses are defined as atomic expressions in a document that can independently express a single fact or information point and are presented in a concise and self - contained form. By using propositional clauses as an intermediate step, the interference of irrelevant information in the document paragraphs can be significantly reduced, thereby improving the capture accuracy and efficiency of the model for relevant information. This strategy enhances the accuracy and robustness of the model in multi - hop reasoning tasks by reducing redundant information. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of data augmentation;
[0025] Figure 2 is a schematic diagram of the training and usage method for a multi - hop question - answering model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The technical solutions of the present invention will be further described below in conjunction with the drawings, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.
[0027] The present invention provides a data augmentation and training method for a multi - hop question - answering retrieval model, as Figure 2 shown, the method includes the following steps:
[0028] Step 1: Obtain a multi - hop question - answering data set, which consists of complex multi - hop questions and their corresponding document sets. These document sets include first - hop retrieval documents, second - hop retrieval documents, and other relevant documents.
[0029] The multi-hop question answering dataset is a dataset specifically designed for training multi-hop question answering models. It contains a large number of complex multi-hop questions and corresponding document collections. Each question requires multiple steps (or "hops") of reasoning to find the answer, demanding that the model can learn and understand the semantic associations between texts in different hops. The present invention uses the publicly available data resources of HotpotQA and MuSiQue for verification. Among them, HotpotQA is a multi-hop reasoning question answering dataset developed by researchers at Stanford University in the United States. This dataset aims to promote the understanding and solving capabilities of complex questions in the field of natural language processing, especially those questions that require integrating information from multiple sources to answer. HotpotQA requires the model to not only understand a single document but also effectively integrate information from different documents to complete multi-step reasoning. MuSiQue (Multihop Questions via Single-hop Question Composition) is a multi-hop reasoning question answering dataset introduced by the Knowledge Engineering Laboratory of Tsinghua University. The design purpose of this dataset is to overcome the problems that can be answered by shortcuts existing in existing datasets. MuSiQue adopts a bottom-up question composition method, enhancing the logical relevance between questions by ensuring that the answer to each sub-question is necessary for answering subsequent questions.
[0030] Step 2: Denoise the positive examples of the first-hop documents and the second-hop documents in the document collection to obtain the denoised documents as new positive examples for model training, and use the remaining parts of the documents as supplementary negative examples for training. The specific steps are as follows:
[0031] Step 21: Use large language models and prompt engineering techniques to extract propositional clauses from the first-hop documents and the second-hop documents in the document collection to obtain the corresponding candidate set of propositional clauses. For example, the original text document: Before the restoration work was carried out from 1990 to 2001, the Leaning Tower of Pisa tilted 5.5 degrees, but now the Leaning Tower of Pisa tilts about 3.99 degrees. This means that the top of the Leaning Tower of Pisa is horizontally offset from the center by 3.9 meters (12 feet 10 inches). Then the extracted set of propositional clauses is: 1. Before the restoration work was carried out from 1990 to 2001, the Leaning Tower of Pisa tilted 5.5 degrees. 2. The Leaning Tower of Pisa now tilts about 3.99 degrees. 3. The top of the Leaning Tower of Pisa is horizontally offset from the center by 3.9 meters (12 feet 10 inches).
[0032] Step 22: According to the multi-hop question, calculate the semantic similarity scores for each propositional clause in the candidate set using BERT Score and sort them in descending order to obtain the sorted candidate set of propositional clauses.
[0033] BERT Score is a semantic similarity calculation method based on the BERT (Bidirectional Encoder Representations from Transformers) model. It utilizes the powerful language understanding ability of BERT to convert text into semantic vector representations and evaluates the semantic relevance between texts by calculating the similarity between these vectors. During the calculation process, BERT Score takes into account lexical matching (Precision and Recall) in the text and the cosine similarity between embedding vectors (a variant of the F1 Score). The calculation steps of BERT Score are as follows:
[0034] Step 221, calculate word embeddings:
[0035] Perform text word embedding on each propositional clause text N in the multi-hop question text C and the set of propositional clauses corresponding to the positive example document, respectively obtaining vector sequences emb_C and emb_N.
[0036] Step 222, calculate the cosine similarity matrix:
[0037] Calculate the cosine similarity between each word embedding in the multi-hop question text and the propositional clause text to obtain a similarity matrix S, and its expression is shown in formula (1):
[0038]
[0039] Among them, S ij represents the result of calculating the cosine similarity between the i-th word embedding in the multi-hop question text and the j-th word embedding in the propositional clause text; cos_similarity(emb_C i , emb_N J ) represents calculating the cosine similarity between two word embedding vectors emb_C i and emb_N J ; emb_C i represents the word embedding vector of the i-th word in the multi-hop question text; emb_N J represents the word embedding vector of the j-th word in the propositional clause text; In, emb_C i * emb_N j represents the dot product of two word embedding vectors, ||emb_C i || and ||emb_N j || respectively represent the norm of the word embedding vector of the i-th word in the multi-hop question text and the norm of the word embedding vector of the j-th word in the propositional clause text. By calculating the dot product divided by the product of the norms of the two vectors, the cosine value of the angle between the two vectors, that is, their cosine similarity, is obtained.
[0040] Step 223: Calculate precision, recall, and the harmonic mean F1:
[0041] For each word in the multi-hop question text, find the word with the highest similarity in the propositional clause text, and calculate the average similarity as precision, whose expression is shown in formula (2):
[0042]
[0043] where P represents precision, that is, the precision rate, |C| in represents the number of words in the multi-hop question text. c ∈ C represents traversing each word c in the multi-hop question text, where C is the set of all words in the multi-hop question text, r ∈ N represents traversing each word r in the propositional clause text, where N is the set of all words in the propositional clause text, and s(c, r) represents the similarity between the word c in the multi-hop question text and the word r in the propositional clause text, max r∈R s(c, r) represents finding the value with the highest similarity to a specific word c in the multi-hop question text among all words in the propositional clause text, Σ c∈ C max r∈R s(c, r) represents finding the word with the highest similarity in the propositional clause text for each word in the multi-hop question text and accumulating these maximum similarity values.
[0044] For each word in the propositional clause text, find the word with the highest similarity in the multi-hop question text, and calculate the average similarity as recall, whose expression is shown in formula (3):
[0045]
[0046] where R represents recall, that is, the recall rate, |N| in represents the number of words in the propositional clause text. c ∈ N represents traversing each word c in the propositional clause text, where N is the set of all words in the propositional clause text, r ∈ C represents traversing each word r in the multi-hop question text, where C is the set of all words in the multi-hop question text. s(c, r) represents the similarity between the word c in the propositional clause text and the word r in the multi-hop question text. max r∈C s(c, r) represents finding the value with the highest similarity to a specific word c in the propositional clause text among all words in the multi-hop question text, Σ c∈R max r∈ Cs(c,r) means finding the word with the highest similarity in the multi-hop question text for each word in the propositional clause text, and accumulating these maximum similarity values.
[0047] The final BERTScore is the harmonic mean of P and R, and its expression is shown in formula (4):
[0048]
[0049] Among them, F1 represents the final BERTScore. P represents the precision calculated previously, which is the average of the maximum similarities found for each word in the multi-hop question text in the propositional clause text, indicating the matching degree of the multi-hop question text and the propositional clause text in terms of precision. R represents recall, which is the average of the maximum similarities found for each word in the propositional clause text in the multi-hop question text, indicating the degree to which the information in the propositional clause text is recalled in the multi-hop question text. is the formula for calculating the harmonic mean of two numbers. Here, the final BERTScore is obtained by calculating the harmonic mean of precision and recall, comprehensively considering the performance of both precision and recall, making the evaluation result more comprehensive and balanced.
[0050] After calculating the semantic similarity scores between all candidate clauses and the multi-hop question, these scores are sorted in descending order. In the sorted set, the clauses with higher scores have a stronger semantic correlation with the question, so they are more likely to contain the answer or relevant information to the question. The top few clauses with the highest scores can be selected as the input for the subsequent reasoning steps according to actual needs.
[0051] Step 23: Set the retention percentage α, retain the top K propositional clauses, and splice the top K propositional clauses in the original word order to obtain the denoised positive example text. The specific steps are as follows:
[0052] Step 231: Define the retention percentage α: The retention percentage α is a value between 0 and 1, used to determine how many proportions of clauses should be retained in the sorted candidate set of propositional clauses. For example, if α = 0.2, it means that the top 20% of the sorted clauses will be retained.
[0053] Step 232, Calculate TopK: The value of TopK can be calculated by multiplying the retention percentage α by the total number of candidate clauses. For example, if the total number of candidate clauses is 100 and α = 0.2, then TopK = 20, which means the top 20 clauses after sorting will be retained. Note: In practical applications, if the calculated value of TopK is not an integer, it is usually necessary to round or truncate. Additionally, if the total number of candidate clauses is small, it may be necessary to adjust the value of α according to the actual situation to ensure that enough clauses are retained.
[0054] Step 233, Concatenate in the original order to obtain the denoised positive example text, which specifically includes the following content:
[0055] 1) Select the TopK propositional clauses: According to the calculated value of TopK, select the top TopK clauses from the sorted candidate set of propositional clauses. These clauses are sorted in descending order of the semantic similarity score with the multi-hop question, so the propositional clause with the highest score has the strongest relevance to the question text.
[0056] 2) Concatenate in the original order: After selecting the TopK clauses, they need to be concatenated in the order they appeared in the original document (i.e., the original order). This step is to maintain the logical coherence and information integrity of the text. Although the order of the clauses may be disrupted during the sorting process, concatenating in the original order can restore their logical relationship.
[0057] 3) Obtain the denoised positive example text: After the processing in step 2) above, a denoised positive example text composed of the TopK propositional clauses concatenated in the original order is obtained. This text removes the information that is irrelevant or redundant to the question and retains the clauses that are most likely to contain the answer or relevant information. It will be used as the input for the subsequent reasoning and answer generation steps to help the model understand and answer questions more accurately.
[0058] Step 3, Input the data obtained in Step 2 into the multi-hop question answering pre-trained language model for training, that is, input the data into the multi-hop question answering pre-trained language model designed by the present invention for training. The specific steps are as follows:
[0059] Step 31, The text embedding model based on Bert is used to convert the question into a question embedding vector, and convert the first-hop document and the second-hop document into corresponding text embedding vectors. These vectors capture the semantic information of the text and can be represented in the semantic vector space, making similar sentences closer in the vector space. Similarly, the negative example documents are also converted into corresponding text embedding vectors for contrast learning with the positive examples during the training process, so as to optimize the model and enhance the model's retrieval ability. The specific steps are as follows:
[0060] Step 311: Convert the multi-hop question text into a question embedding vector. When dealing with multi-hop question answering tasks, first, the user's question needs to be converted into a question embedding vector. This is achieved by inputting the question text into a BERT-based text embedding model. BERT is a pre-trained deep bidirectional encoder that can well capture the semantic information of the text through the language representations learned on large-scale text data. The question embedding vector captures the semantic information of the question, enabling the model to understand the intention and key content of the question.
[0061] Step 312: Convert the first-hop document and the second-hop document into corresponding text embedding vectors. In multi-hop question answering, the first-hop document usually refers to the text that directly answers the question, while the second-hop document is the information further required based on the currently obtained first-hop document and the question text. These texts are also converted into corresponding text embedding vectors so that the model can understand and process them. These text embedding vectors also capture the semantic information of the text, enabling the model to identify the associations and differences between texts.
[0062] These text embedding vectors are represented in a semantic vector space, where each dimension represents a specific semantic feature of the text. Similar sentences (i.e., semantically similar or related sentences) will have a closer distance, while different sentences will have a farther distance. This representation method enables the model to more easily identify and understand the semantic relationships between texts.
[0063] Step 313: Convert the negative example documents into corresponding embedding vectors. To optimize the model's ability to identify the correct information flow, negative example documents are usually introduced. Negative example documents refer to texts that are irrelevant or incorrect to the question or the correct answer. These negative example documents are also converted into corresponding embedding vectors for contrastive learning with positive examples (i.e., the question, the first-hop document, and the second-hop document) during the training process.
[0064] Step 32: Design a multi-hop question answering pre-trained language model
[0065] The model idea designed by the present invention is to train a sentence encoder by using previous supporting facts to predict subsequent supporting facts. This method realizes the efficient encoding of document content by converting the text of the entire document into an embedding vector of a fixed dimension. However, in the face of documents of variable length, the MDR model will inevitably encounter noise problems during processing because a large amount of information contained in the document may not be directly related to the goal of semantic association with the training document. For this reason, the MDR model is selected as the baseline model of the present invention. The in-batch strategy adopted by the MDR model, although it plays a role in reducing the cost of negative sampling, its negative sample selection mechanism is too random and simple, which limits the ability of the model to fully learn and capture complex logical relationships. To solve these problems, the present invention introduces a positive example denoising strategy based on propositional clauses. A propositional clause is defined as an atomic expression in the text, and each proposition contains an independent fact or information point and is presented in a concise and self-contained natural language format. This strategy can effectively avoid too much semantic and sentence noise information contained in the paragraph. By using propositions as an intermediate step, the present invention can maximize the reduction of information unrelated to the problem in the text segment.
[0066] The process of positive example denoising is as Figure 1 shown, and the denoising residues of other samples in the batch are combined as negative examples. Specifically, first, a proposition extraction model is used to identify and extract multiple proposition sets from a single text paragraph. Then, for each proposition in the set, its semantic relevance score with the given question q and the relevant supporting documents is calculated. The score can be calculated using the ordinary Bert Score or the Bert-based Rank model can also be selected to calculate the score.
[0067] To effectively screen out important propositions, the present invention introduces a hyperparameter a, which is used to control the ratio between the number of retained propositions and the total number of propositions. According to the hyperparameter a, the present invention selects the propositions within this ratio range to be retained, and these propositions will be re-concatenated in the order in which they appear in the original text to maintain the coherence of the text. In addition, to avoid losing propositions containing answer information, the present invention also includes those propositions related to the answer in the retention set. The remaining proposition set = P\ is used to generate additional negative example samples to further enrich the training data set, where P is the proposition set obtained from a single document by the proposition extraction model. The formation formulas are shown in Formulas (5) and (6):
[0068] Pa + ←augment(P + ,q,K ɑ ) (5)
[0069] Pa - ←P + \Pa+ (6)
[0070] Among them, P + is the original positive example document, K ɑ is the number of selected sentences, q is the given question of the training set sample, and Pa + is the denoised positive example document. Pa - represents the part after removing the positive examples and is used as the negative examples for training other samples. Through the above method, the present invention not only improves the quality of the data and the robustness of the model, but also enhances the model's ability to distinguish relevant texts from irrelevant texts by introducing positive and negative example samples.
[0071] Step 33: Input the data obtained in Step 31 into the multi-hop question-answering pre-trained language model designed in Step 32 for training.
[0072] During the model training process, the model will learn how to distinguish positive and negative example documents. Through contrastive learning, the model can gradually optimize its ability to identify the correct information flow, that is, it can more accurately judge which texts are necessary to answer the question and which texts are irrelevant or incorrect. This ability is crucial for improving the accuracy and efficiency of the multi-hop question-answering system.
[0073] During the model training process, the relationship between the question embedding vector and the relevant text embedding vector is optimized by calculating different contrastive losses. The contrastive losses include:
[0074] a. For the question embedding vector, two sets of positive and negative examples are used to calculate the first two contrastive losses. In these two sets, the first-hop document embedding vector is used as the positive example, while the second-hop document embedding vector and the in-batch negative example embedding vector are used as the negative examples, thereby obtaining contrastive loss 1.
[0075] b. Combine the question embedding vector with the first-hop document embedding vector, use the second-hop document embedding vector as the positive example, and the in-batch negative example embedding vector as the negative example, thereby obtaining contrastive loss 2.
[0076] c. Add a new contrastive loss 3, which is specifically optimized for the relationship between the question embedding vector and the positive example embedding vectors of two consecutive steps (i.e., the first hop and the second hop). In this case, the positive example is jointly composed of the first-hop document embedding vector and the second-hop document embedding vector, and the negative example is the newly added negative example embedding vector after denoising. Such a design is to further improve the performance of the model in complex multi-hop reasoning scenarios.
[0077] The present invention uses the BM25 algorithm to mine difficult-to-process negative examples from the Wikipedia knowledge base and combines the denoising residuals of other samples within the batch as negative examples. This comprehensive strategy not only improves the model's ability to identify positive examples but also enhances the model's training effect through carefully selected negative examples, thereby improving the overall performance of the model. The training loss functions are shown in Formulas (7), (8), and (9):
[0078]
[0079] L total = L1 + λL2 (9)
[0080] Among them, L1 represents the loss function using denoised positive samples, which is similar to the loss of the original MDR but uses denoised positive samples; L2 represents an additional loss function that includes negative samples formed from discarded positive propositions and previously unused hard negative samples; L total represents the overall loss function; represents the feature mapping function that converts the input into its feature representation; T represents the temperature hyperparameter used to scale the dot product result and affect the behavior of the softmax function; λ represents the hyperparameter that balances the influence of the two loss functions L1 and L2; x i represents the query, that is, the question; represents the denoised positive sample; x k represents the negative samples within the batch, including hard negative samples from BM25, and B is the total number of batches; x n represents the negative samples formed from discarded positive propositions and previously unused hard negative samples, and M is the total number of negative samples.
[0081] Embodiment:
[0082] Suppose the HotpotQA dataset is used. One specific multi-hop question and its corresponding document set are as follows:
[0083] Question Q: What is the nationality of the director of the movie "The Godfather"?
[0084] Document set D:
[0085] The first-hop document d_1: Francis Ford Coppola is an American film director, producer, and screenwriter. His most famous works are directing the "Godfather" trilogy. Coppola was born on April 7, 1939, in Detroit, Michigan, and he obtained a drama degree from Hofstra University.
[0086] Second-hop document d_2: "The Godfather" is a 1972 American crime film directed by Francis Ford Coppola. The film is based on Mario Puzo's 1969 novel of the same name. "The Godfather" is considered one of the greatest films in the history of cinema and has won numerous Academy Awards, including the Best Picture award.
[0087] Positive example denoising:
[0088] Extract propositional clauses:
[0089] d_1:
[0090] 1. Francis Ford Coppola is an American film director.
[0091] 2. His most famous works are directing the "Godfather" trilogy.
[0092] 3. Coppola was born on April 7, 1939, in Detroit, Michigan.
[0093] 4. He obtained a drama degree from Hofstra University.
[0094] d_2:
[0095] 1. "The Godfather" is a 1972 American crime film directed by Francis Ford Coppola.
[0096] 2. The film is based on Mario Puzo's 1969 novel of the same name.
[0097] 3. "The Godfather" is considered one of the greatest films in the history of cinema.
[0098] 4. The film has won numerous Academy Awards, including the Best Picture award.
[0099] Calculate semantic similarity scores:
[0100] d_1:
[0101] 1. Francis Ford Coppola is an American film director. Score: 0.85.
[0102] 2. His most famous works are directing the "Godfather" trilogy. Score: 0.90.
[0103] 3. Coppola was born on April 7, 1939, in Detroit, Michigan. Score: 0.70.
[0104] 4. He obtained a drama degree from Hofstra University. Score: 0.65.
[0105] d_2:
[0106] 1. *The Godfather* is a 1972 American crime film directed by Francis Ford Coppola. Score: 0.88.
[0107] 2. The film is adapted from Mario Puzo's 1969 novel of the same name. Score: 0.75.
[0108] 3. *The Godfather* is considered one of the greatest films in the history of cinema. Score: 0.70.
[0109] 4. The film won several Academy Awards, including the Best Picture award. Score: 0.65.
[0110] Retain the top K propositional clauses (assuming \(\alpha = 0.5\), retain the first 2):
[0111] The denoised positive example text \(Pa + ) = [d_1* = "Francis Ford Coppola is an American film director. His most famous works are directing the *Godfather* trilogy" d_2* = "*The Godfather* is a 1972 American crime film directed by Francis Ford Coppola. The film is adapted from Mario Puzo's 1969 novel of the same name."]
[0112] Additional negative examples \(Pa -) = ["Coppola was born on April 7, 1939, in Detroit, Michigan. He received a degree in drama from Hofstra University.", "*The Godfather* is considered one of the greatest films in the history of cinema. The film won several Academy Awards, including the Best Picture award."]
[0113] Negative example generation:
[0114] Use the BM25 algorithm to mine hard negative examples from the Wikipedia knowledge base, such as:
[0115] "Francis Ford Coppola is a versatile filmmaker whose works cover a variety of genres."
[0116] "He not only directed the *Godfather* series but also classic films such as *Apocalypse Now*."
[0117] "There are also many members of Coppola's family working in the film industry, including his son Roman Coppola."
[0118] "He is also a famous film producer and has produced many successful films."
[0119] loss1 is similar to the loss of the original MDR but uses denoised positive samples. A few negative examples mined by the BM25 algorithm are used as negative samples.
[0120] The positive examples of loss2 are the concatenation of two denoised positive samples. For the negative examples of the BM25 algorithm, in addition to the negative examples used in loss1, the denoising residuals (\(Pa-\)) of other samples within the batch are used as the negative examples of loss2.
[0121] Inference process:
[0122] During the process of using the model for inference, steps such as the denoising step during training are not required. Instead, the trained model is directly used to embed the questions and candidate documents, and the similarity is calculated for document retrieval.
Claims
1. A data augmentation and training method for a multi-hop question-answering retrieval model, characterized in that The method comprises the following steps: Step 1: Obtain a multi-hop question answering dataset, which consists of complex multi-hop questions and their corresponding document sets, including first-hop retrieval documents, second-hop retrieval documents, and other related documents; Step 2: Perform positive example denoising on the first-hop document and the second-hop document in the document collection, and use the denoised documents as new positive examples for model training. The remaining parts of the documents are used as supplementary negative examples for training. The specific steps are as follows: Step 21: extract proposition clauses from the first-hop document and the second-hop document in the document set using a large language model and prompt engineering technology to obtain a corresponding proposition clause candidate set; Step 22: According to the multi-hop problem, the semantic similarity score of each proposition clause in the candidate set is calculated using BERT Score and sorted in descending order to obtain a sorted proposition clause candidate set; Step 23, set the retention percentage α, retain the TopK proposition clauses, and concatenate the TopK proposition clauses in the original order to obtain the denoised positive example text; Step 3: Input the data obtained in step 2 into the multi-hop question-answering pre-trained language model for training. The specific steps are as follows: Step 31: Based on the Bert text embedding model, the first-hop document, the second-hop document, and the negative example document are converted into corresponding text embedding vectors; Step 32: Design a multi-hop question-answering pre-trained language model Step 321, selecting the MDR model as the baseline model; Step 322: introduce a positive example denoising strategy based on proposition clauses. Proposition clauses are defined as atomic expressions in the text. Each proposition contains an independent fact or information point. The process of positive example denoising is as follows: first, a proposition extraction model is used to identify and extract multiple proposition sets from a single text paragraph. Then, for each proposition in the set, the semantic relevance score of the proposition with the given question q and the related supporting documents is calculated; Step 323: Introduce a hyperparameter to control the ratio between the number of retained propositions and the total number of propositions. According to the hyperparameter a, propositions within this ratio are selected and retained. These propositions will be reassembled according to their order in the original text, and propositions related to the answer will also be included in the retained set. The remaining proposition set = P\ is used to generate additional negative samples, where P is the proposition set obtained after a single document is processed by the proposition extraction model. Step 33: input the data obtained in step 31 into the multi-hop question-answering pre-trained language model designed in step 32 for training.
2. The data enhancement and training method for a multi-hop question-answering retrieval model according to claim 1, characterized in that The specific steps of step 22 are as follows: Step 221, calculate word embedding: For multi-hop question text Each proposition clause text of the proposition clause set corresponding to the positive document Perform text word embedding to obtain vector sequences emb_C and emb_ ; Step 222, calculate the cosine similarity matrix: Calculate the cosine similarity between each word embedding in the multi-hop question text and the proposition clause text to obtain the similarity matrix S, which is expressed as shown in the following formula: in, Indicates the calculation of the multi-hop problem text The embedding of the word is related to the proposition clause text. The cosine similarity results between the embeddings of the words; Indicates the calculation of two word embedding vectors and The cosine similarity of Represents the first The word embedding vector of each word; Represents the proposition clause in the text The word embedding vector of each word; middle, represents the dot product of two word embedding vectors, and Respectively represent the first The modulus of the word embedding vector of the word and the first The modulus of the word embedding vector of the word; Step 223, calculate precision, recall and harmonic mean F1: For each word in the multi-hop question text, find the word with the highest similarity in the proposition clause text and calculate the average similarity as precision, which is expressed as follows: Among them, P represents precision, that is, the accuracy rate. represents the number of words in the multi-hop question text, represents traversing each word c in the multi-hop question text, It means traversing each word r in the proposition clause text, represents the similarity between word c in the multi-hop question text and word r in the proposition clause text, It means that for a specific word c in the multi-hop question text, find the value with the highest similarity among all the words in the proposition clause text. It means that for each word in the multi-hop question text, find the word with the highest similarity in the proposition clause text, and add up these maximum similarity values; For each word in the proposition clause text, find the word with the highest similarity in the multi-hop question text and calculate the average similarity as recall, which is expressed as follows: Among them, R represents recall, that is, the recall rate. Indicates the number of words in the proposition clause text; The final BERTScore is the harmonic mean of P and R, and its expression is shown in the following formula: Among them, F1 represents the final BERTScore; After calculating the semantic similarity scores between all candidate clauses and the multi-hop question, these scores are sorted in descending order.
3. The data enhancement and training method for a multi-hop question-answering retrieval model according to claim 1, characterized in that The specific steps of step 23 are as follows: Step 231, define the retention percentage α: the retention percentage α is a value between 0 and 1, which is used to determine the proportion of clauses that should be retained in the sorted candidate set of propositional clauses; Step 232, calculate TopK: the value of TopK is calculated by multiplying the retention percentage α by the total number of candidate clauses; Step 233: splice in the original word order to obtain the denoised positive example text.
4. The data enhancement and training method for a multi-hop question-answering retrieval model according to claim 3, characterized in that The step 233 specifically includes the following contents: 1) Select the TopK proposition clauses: According to the calculated TopK values, select the topK clauses from the sorted proposition clause candidate set; 2) Splicing in original order: After selecting the top K clauses, they need to be spliced in the order in which they appear in the original document; 3) Obtaining the denoised positive example text: After the processing in the above step 2), a denoised positive example text consisting of the TopK proposition clauses in the original order is obtained.
5. The data enhancement and training method for a multi-hop question-answering retrieval model according to claim 1, characterized in that The specific steps of step 31 are as follows: Step 311, converting the multi-hop question text into a question embedding vector; Step 312: The first-hop document and the second-hop document are converted into corresponding text embedding vectors; Step 313: The negative example documents are converted into corresponding embedding vectors.
6. The data enhancement and training method for a multi-hop question-answering retrieval model according to claim 1, characterized in that In step 33, during the model training process, the relationship between the question embedding vector and the related text embedding vector is optimized by calculating the following contrast loss: in, Represents the loss function using denoised positive samples; represents an additional loss function that includes negative samples formed from discarded positive propositions and previously unused hard negative samples; represents the overall loss function; Represents the feature mapping function, which converts the input into its feature representation; Represents the temperature hyperparameter, which is used to scale the dot product result and affect the behavior of the softmax function; Represents hyperparameters, balance and The impact of the two loss functions; It represents a query, i.e. a question; represents the denoised positive sample; represents negative samples in the batch, including hard negative samples from BM25, is the total number of batches; represents negative samples formed from discarded positive propositions and previously unused hard negative samples, is the total number of negative samples.