A model training method, system and information retrieval method based on attribution text
By building an attribution text-based training method in a large language model, using relevant document sets to generate samples and perform noise enhancement, the problem of high manual labeling costs is solved, high-quality attribution text is efficiently generated, and the output accuracy and user satisfaction of the large language model are improved.
Patent Information
- Application Number
- CN202510998368.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-21
AI Technical Summary
When existing large language models generate attribution text, the cost of manually annotated high-quality open source attribution text training data is high, making it difficult to improve output accuracy and unable to meet user information retrieval needs.
By extracting relevant document sets from the original dataset, using a large language model to generate summaries with citation relationships, constructing samples, filtering low-quality samples through the F1 score, adding a noise enhancement strategy for training, and optimizing the large language model to improve its accuracy in generating attribution text.
High-quality attribution text training samples can be generated without manual labeling, which improves the training efficiency and output quality of large language models, reduces costs, provides more reliable answers that meet user needs, and improves information processing efficiency.
Smart Images

Figure CN120509454B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model training method, system and information retrieval method based on attribution text. Background Art
[0002] In recent years, with the increasing convenience of natural language interaction, more and more users are turning to large language models (LLMs) to meet their information retrieval needs. However, despite the rich knowledge acquired during pre-training, LLMs' output can sometimes deviate from user instructions and produce hallucinations, significantly limiting their ability to meet information retrieval needs. Furthermore, due to the lack of clear attribution, verifying the correctness of LLM-generated content is difficult.
[0003] To better meet the needs of information retrieval, attribution text generation has recently attracted significant attention from both academia and industry. The goal is to improve the reliability of generated content by providing citations for each assertion. Despite the importance of attribution text generation, existing open-source attribution systems still have significant room for improvement. Currently, building attribution systems typically involves a three-step process: first, external search engines are used to retrieve relevant passages; second, these passages are incorporated into carefully designed templates along with the user query; and finally, these templates serve as input to LLMs, guiding them to generate responses with appropriate citations.
[0004] Because existing LLMs are designed to follow user instructions, this approach can cause the model to generate responses based on irrelevant retrieval passages, leading to incorrect attribution. Furthermore, even when relevant passages are provided as input, relying solely on the task generalization capabilities of LLMs still struggles to follow user instructions and generate correctly cited responses that match the question. The high cost of manually annotated, high-quality, open-source attribution text training data makes it difficult to improve the output accuracy of large language models used to generate attribution text, making it impossible to meet user information retrieval needs. Summary of the Invention
[0005] To this end, the technical problem to be solved by the present invention is to overcome the problem in the prior art that the cost of manually annotated high-quality open source attribution text training data is high, which makes it difficult to improve the output accuracy of the large language model used to generate attribution text.
[0006] To solve the above technical problems, the present invention provides a model training method based on attribution text, comprising:
[0007] Extract relevant document sets from the original dataset;
[0008] Using a large language model, a summary with citation relationships is generated based on all documents in the relevant document set. The summary is used as the answer, and the citation relationship is used as the reference of the answer. The large language model is then used to generate questions for the answers. The corresponding questions, answers, citations, and related document sets constitute a sample.
[0009] After generating multiple samples, calculate the F1 score of each sample, remove samples with F1 scores lower than the filtering threshold, and obtain filtered samples;
[0010] For each filtered sample, we randomly select irrelevant documents from the sample's related document set from the original dataset, add them to the sample's related document set and shuffle their order. We also change the reference of the sample's answer so that it links to the correct document in the related document set, thus obtaining a noise-enhanced sample.
[0011] A training set is constructed using noise-enhanced samples. The relevant document set and question of each sample in the training set are used as input, and the answers and citations are used as labels. The large language model is fine-tuned through supervision to obtain a fully trained large language model.
[0012] Preferably, unbiased answers and misquotes are constructed for each sample in the training set, and preference optimization is performed on the trained large language model, including:
[0013] The answers and citations of the samples in the training set are the preferred answers and correct citations of the samples;
[0014] Perform random reference addition, random reference deletion, and random reference change on the correct reference to obtain the incorrect reference;
[0015] Retain irrelevant documents from the relevant document set in the sample, input the irrelevant documents and sample questions into the large language model, and obtain unfavorable answers and misquotes for the sample;
[0016] The trained large language model is optimized based on the sample's preferred answers and correct citations, as well as the non-preferred answers and incorrect citations.
[0017] Preferably, if the original data set is a structured data set, extracting the relevant document set from the original data set includes:
[0018] Randomly select a paragraph from the original dataset as a document and construct a related document set containing only one document;
[0019] Multiple entity triples are randomly selected from the original data set, and the paragraphs containing the two entities in the entity triples are obtained as documents in the related document set to construct a related document set containing multiple documents.
[0020] Preferably, if the original data set is an unstructured data set, extracting a relevant document set from the original data set includes:
[0021] Classify the documents in the original data set and retain the categories with the number of documents not less than a first preset number;
[0022] For each retained category, calculate the similarity between all documents in the category and arrange them in descending order; retain the second preset number of document pairs with the highest similarity;
[0023] For each document pair, the key information and the most overlapping strings of each document are extracted as new documents, and the related document set is constructed with the new documents.
[0024] Preferably, when generating multiple samples based on a related document set, the process includes:
[0025] Using a large language model, multiple summaries with citation relationships are generated based on all documents in the current relevant document set. Each summary is regarded as an answer, and the citation relationship of each summary is used as the reference of its answer. Then, the large language model is used to generate a question for each answer. The corresponding question, answer and reference constitute a question-answer pair. Each question-answer pair and the current relevant document set constitute a sample, and multiple samples are obtained.
[0026] Preferably, a large language model is used to generate a summary with a citation relationship based on all documents in the relevant document set, and the summary is used as the answer, and the citation relationship is used as the reference of the answer; then the large language model is used to generate a question for the answer, and the corresponding question, answer, reference and related document set constitute a sample. The large language model adopts the gpt-4-1106-preview model or the Doubao-1.5-Pro-32k model.
[0027] Preferably, after generating multiple samples, the F1 score of each sample is calculated using the reference quality standard proposed by ALCE, including:
[0028] The sample's responses are divided into multiple statements. Each statement is used as a hypothesis, and each reference in the statement is used as a premise. The statement and its reference are input into the natural language inference model for judgment. If there is an implication relationship between the statement and the reference, the reference is marked as a valid reference, otherwise it is an invalid reference.
[0029] Calculate the ratio of the number of valid citations of each statement to the total number of citations of the statement, and take the average of the ratios as the citation recall rate;
[0030] The ratio of the total number of valid citations to the total number of citations is used as the citation precision rate;
[0031] The F1 score of a sample is calculated based on the citation recall and citation precision.
[0032] Preferably, the natural language inference model is a t5-xxl-true-nli-mixture model.
[0033] The present invention also provides a model training system based on attribution text, comprising:
[0034] Related document acquisition module, used to extract relevant document sets from the original data set;
[0035] The sample generation module is used to use the large language model to generate a summary with citation relationships based on all documents in the relevant document set, using the summary as the answer and the citation relationship as the reference of the answer; the large language model is then used to generate a question for the answer, and the corresponding question, answer, reference and related document set constitute a sample;
[0036] The sample filtering module is used to generate multiple samples, calculate the F1 score of each sample, and remove samples with F1 scores lower than the filtering threshold to obtain filtered samples;
[0037] The noise enhancement module is used to randomly select irrelevant documents from the original dataset for each filtered sample, add them to the sample's related document set and shuffle their order, and simultaneously change the reference of the sample's answer to link it to the correct document in the related document set to obtain a noise-enhanced sample;
[0038] The supervised fine-tuning module is used to construct a training set using noise-enhanced samples. It takes the relevant document set and questions of each sample in the training set as input, and uses answers and citations as labels to perform supervised fine-tuning on the large language model to obtain a fully trained large language model.
[0039] The present invention also provides an information retrieval method, comprising:
[0040] The retriever retrieves relevant documents based on the question input by the user to form a relevant document set. The question and the relevant document set are input into the large language model trained by the above-mentioned attribution text-based model training method to obtain the answer and its corresponding reference in the relevant document set.
[0041] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0042] The present invention discloses a model training method based on attribution text. First, the documents related to each other in the original data set are clustered to obtain a related document set. Then, answers, citations, and questions are generated based on the related document set. Samples are constructed with the corresponding related document set, questions, answers, and citations. Samples with low citation quality are then eliminated based on the citation quality standard to improve the quality of the training set. Finally, irrelevant documents are added to the sample for noise enhancement. The training set constructed in this way can help the large language model to increase its attention to relevant documents during training and avoid the large language model from generating answers based on irrelevant documents. The present invention does not require manual annotation and can automatically generate high-quality attribution text training samples, thereby improving the training efficiency of the large language model used to generate attribution text, reducing the cost of manual annotation, and improving the quality of the answers and citations output by the model to provide reliable answers that better meet user needs and improve the efficiency of information processing in various industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0044] Figure 1 It is a flow chart of a model training method based on attribution text of the present invention;
[0045] Figure 2 It is a structural diagram of a model training method based on attribution text of the present invention;
[0046] Figure 3 It is a sample example diagram constructed by the present invention;
[0047] Figure 4 This is the result graph of the effect of filtering threshold on sample quality;
[0048] Figure 5 This is the result graph of different natural language inference models filtering samples, where Figure 5 (a) in the figure is the filtering result of the t5-xxl-true-nli-mixture model. Figure 5 (b) in the figure is the filtering result diagram of the ModernBERT-base-nli model. DETAILED DESCRIPTION
[0049] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0050] Reference Figure 1 As shown, the first embodiment of the present invention provides a model training method based on attribution text, including:
[0051] S1: Extract relevant document sets from the original dataset;
[0052] S2: Using a large language model, generate a summary with a citation relationship based on all documents in the relevant document set, use the summary as the answer, and the citation relationship as the reference of the answer; then use the large language model to generate a question for the answer, and the corresponding question, answer, citation and related document set constitute a sample ,in For a set of related documents, For the problem, To answer, is the reference, i is the sample index;
[0053] S3: After generating multiple samples, calculate the F1 score of each sample and remove samples with F1 scores lower than the filtering threshold to obtain filtered samples;
[0054] S4: For each filtered sample, randomly select irrelevant documents from the original dataset, add them to the sample's related document set and shuffle their order. Simultaneously, change the reference of the sample's answer so that it links to the correct document in the related document set, thus obtaining a noise-enhanced sample.
[0055] S5: Construct a training set using noise-enhanced samples. Take the relevant document set and question for each sample in the training set as input, and use the answer and citation as labels to perform supervised fine-tuning (SFT) on the large language model to obtain a fully trained large language model.
[0056] Figure 2 This is a structural diagram of a model training method based on attribution text in the present invention.
[0057] Preferably, for each sample in the training set Constructing a non-preferred answer and misquotes In the preference optimization (PO) phase, the trained large language model is optimized for preference, including:
[0058] The answers and citations of the samples in the training set are the preferred answers and correct citations of the samples;
[0059] Perform random reference addition, random reference deletion, and random reference change on the correct reference to obtain the incorrect reference;
[0060] To prevent the language model from focusing on irrelevant documents, we perform irrelevant document focusing, which means retaining irrelevant documents from the relevant document set in the sample. We input questions about irrelevant documents and samples into the large language model to obtain unfavorable answers and incorrect citations for the sample.
[0061] The trained large language model is optimized based on the sample's preferred answers and correct citations, as well as the non-preferred answers and incorrect citations.
[0062] Specifically, random citation addition is to prevent the language model from containing too many incorrect citations, and some citations are randomly added to construct incorrect citations; random citation deletion is to delete some golden citations to construct incorrect citations, aiming to avoid the omission of key citations; random citation change is to avoid citing irrelevant documents, and some key citations are replaced with random citations as incorrect citations.
[0063] Through the above four strategies of random citation addition, random citation deletion, random citation change and irrelevant document focus, different types of "inferior" responses and corresponding incorrect citations are obtained for use in training and preference answers. Comparative non-preferential response and misquotes .
[0064] In the inference phase, the set of relevant documents By external search engine based on the query After the search is performed, n represents the total number of documents in the relevant document set. Each document consists of one or more sentences and is 100 words long. Given a query and the corresponding retrieved set of related documents , attributed text generation aims to generate an answer and citations , m is the total number of citations. Each reference in is an index to one of the retrieved documents. Figure 3 For a sample example.
[0065] In the second embodiment provided by the present invention, a structured data set is used as the original data set. Then, in S1, the relevant document set is extracted from the original data set, including:
[0066] Randomly select a paragraph from the original dataset as a document and construct a related document set containing only one document;
[0067] Multiple entity triples are randomly selected from the original data set, and the paragraphs containing the two entities in the entity triples are obtained as documents in the related document set to construct a related document set containing multiple documents.
[0068] Specifically, Example 2 uses the structured dataset Wiki-Graphs as the original dataset Wiki-Graphs consist of many entity triples Composition, representing entity With entity Have a relationship .entity With entity Link to paragraph sets individually and .
[0069] If you need to extract the relevant document set The number of relevant documents in n=1, then from the original dataset Randomly select a paragraph and construct a related document set , i represents the sample index.
[0070] If you need to extract the relevant document set The number of relevant documents n≥2, we can first select a target entity and then search for the relationship with the entity in Wiki-Graphs. All adjacent entities of the target entity are obtained to obtain multiple entity triplets, and further obtain the corresponding paragraphs of the target entity and the adjacent entities in Wiki-Graphs. In this way, a set of Wikipedia paragraphs with direct semantic associations with the target entity in the graph can be obtained, and then the paragraphs containing more than two entities are selected as documents in the related document set to construct a related document set containing multiple documents. .
[0071] In Example 2, a large language model is used in S2 to generate a summary with a citation relationship based on all documents in the relevant document set, and the summary is used as the answer, and the citation relationship is used as the reference of the answer; then the large language model is used to generate a question for the answer, and the corresponding question, answer, reference and related document set constitute a sample. The large language model adopts the gpt-4-1106-preview model.
[0072] Specifically, the relevant document set was selected Then, if the number of documents in the relevant document set is , recognizing that generating questions for a given answer is more direct than answering the question, we use the gpt-4-1106-preview model and the template with the prompt "generate an answer and cite it correctly" to generate a question for the document A summary with a citation relationship is used as the answer, and the citation relationship is used as the reference of the answer. Subsequently, the gpt-4-1106-preview model is used to generate a question based on the answer to obtain a question-answer pair. The question-answer pair and the related document set are directly used to form a sample. .
[0073] If the number of documents in the relevant document set In order to improve the speed of sample data construction, the gpt-4-1106-preview model is used according to the relevant document set Generate a summary with citations for all documents in the , the reference relationship is the reference of the answer ; Then use the gpt-4-1106-preview model to generate a question for the answer ,The corresponding questions, answers, quotes and their related document sets constitute a sample.
[0074] Specifically, the model generates a question-answer pair for a set of related documents in the format of Question:…, Answer:…, Reference:…, and then uses regular expressions to extract the content after fixed keywords (Question / Answer / Reference), extracting the corresponding question, answer, and reference into a question-answer pair. In the form of, and then with the relevant document set to form a sample .
[0075] A question-answer pair can be generated based on a set of related documents to construct a sample, or multiple question-answer pairs can be generated based on a set of related documents to construct multiple samples accordingly.
[0076] Specifically, multiple question-answer pairs are generated based on a relevant document set, and multiple samples are constructed accordingly, including: using a large language model to generate multiple summaries with reference relationships based on all documents in the current relevant document set, taking each summary as an answer, and the reference relationship of each summary as the reference of its answer; then using the large language model to generate a question for each answer, the corresponding question, answer and reference constitute a question-answer pair, each question-answer pair and the current relevant document set constitute a sample, and multiple samples are obtained.
[0077] In Example 2, 31,823 entity triples were selected to obtain 31,823 related document sets, half of which contained single documents and the other half contained multiple documents. This example generates a question-answer pair based on a related document set and constructs a sample, obtaining 31,823 samples.
[0078] In order to filter the data, it is necessary to remove samples with low citation quality. In S3, Example 2 uses the citation quality standard proposed by ALCE (Automatic LLMs' Citation Evaluation) to evaluate each sample. The quality of the citations is calculated by calculating the F1 score for each sample, including:
[0079] The sample answers Divide the data into multiple statements; treat each statement as a hypothesis and each reference in the statement as a premise. Input the statement and its references into the Natural Language Inference (NLI) model to determine whether each reference fully supports or does not fully support each statement.
[0080] The natural language inference model will output entailment, neutrality, or contradiction for each statement-reference relationship. Entailment means that the reference effectively supports the statement, otherwise it is considered unsupportive. If the relationship between the statement and the reference is entailment, the reference is marked as valid, otherwise it is invalid.
[0081] Calculate the ratio of the number of valid citations of each statement to the total number of citations of the statement, and take the average of the ratios as the citation recall rate;
[0082] The ratio of the total number of valid citations to the total number of citations is used as the citation precision rate;
[0083] The F1 score of a sample is calculated based on the citation recall and citation precision.
[0084] Specifically, the citation recall is used to evaluate whether the cited paragraph fully supports the reply content. The formula is:
[0085] ;
[0086] in, is the citation recall rate, The number of statements divided for the response.
[0087] Specifically, citation precision is used to identify irrelevant citations. When a citation fails to confirm a statement, but the remaining citations can still support the statement, the citation is considered irrelevant to the statement. The formula is:
[0088] ;
[0089] in, is the citation accuracy.
[0090] The F1 score of the sample is calculated based on the citation recall rate and citation precision rate. The formula is:
[0091] ;
[0092] in, is the F1 score of the sample.
[0093] Preferably, the natural language inference model used in Example 2 is the t5-xxl-true-nli-mixture model.
[0094] In Example 2, the filtering threshold is set to 0.9 for data filtering. After filtering, Example 2 retains 13,225 samples.
[0095] Due to the limitations of the retriever, the retrieved documents inevitably contain some irrelevant information. In order to solve the above challenges, S4 of the present invention introduces a noise document enhancement strategy: for each filtered sample, Randomly select irrelevant documents and add them to the relevant document set of the sample and disrupt the order, and simultaneously change the reference of the sample answer Link it with the correct document in the relevant document set to obtain the noise-enhanced sample.
[0096] Given that entity-based document clustering methods are limited to data with structured information, the present invention also utilizes data sources lacking structured information for data generation to demonstrate the flexible scalability of the method of the present invention.
[0097] In the third embodiment provided by the present invention, an unstructured data set is used as the original data set. Then, in S1, a relevant document set is extracted from the original data set, including:
[0098] Classify the documents in the original data set and retain the categories with the number of documents not less than a first preset number;
[0099] For each retained category, calculate the similarity between all documents in the category and arrange them in descending order; retain the second preset number of document pairs with the highest similarity;
[0100] For each document pair, the key information and the most overlapping strings of each document are extracted as new documents, and the related document set is constructed with the new documents.
[0101] Example 3 uses the ArXiver dataset as the original dataset, which contains 63,357 ArXiver papers published from January 2023 to October 2023. To ensure that each category under the dataset can identify a sufficient number of relevant documents, this embodiment selects 34 categories containing no less than 500 documents. For each retained category, in order to obtain documents that are related to each other within the category, this embodiment uses the GTE-ModernBERT-Base model to calculate the similarity between each pair of documents in the same category, and arranges them in order from high to low. To ensure the diversity of the data, the second preset number is set to 200, that is, for categories with more than 200 pairs of document pairs, 200 pairs are retained from the document pairs with the highest similarity. For the original dataset, a total of 5171 mutually related document pairs, that is, 5171 related document sets, were obtained.
[0102] Because the documents are very long, this example uses regular expressions to extract key information from each document pair. This information is the declarative fragment that semantically carries the core content of the document, such as the claims, conclusions, or theorems in the paper. The overlapping parts of the two documents are then calculated, and the string with the most overlap is retained from each document pair as the new document, from which the related document set is constructed. The claims refer to the core ideas put forward by the author in the paper.
[0103] In S2 of Example 3, a large language model is used to generate a summary with a citation relationship based on all documents in the relevant document set, and the summary is used as the answer, and the citation relationship is used as the reference of the answer; then the large language model is used to generate questions for the answers, and the corresponding questions, answers, citations and their related document sets constitute a sample. Considering the balance between cost and performance, the large language model adopts the Doubao-1.5-Pro-32k model. This embodiment generates multiple question-answer pairs based on a relevant document set, and correspondingly constructs multiple samples. A total of 19,924 question-answer pairs were generated, and finally 19,924 samples were constructed.
[0104] In S3 of Example 3, the citation quality standard proposed by ALCE is also used to evaluate the citation quality of each sample, and the F1 score of each sample is calculated. The natural language inference model adopts the t5-xxl-true-nli-mixture model, and samples with a citation F1 score lower than 0.9 are deleted.
[0105] In summary, the model training method based on attribution text described in the present invention first clusters the interrelated documents in the original data set to obtain a related document set, then generates answers, citations and questions based on the related document set, and constructs samples with the corresponding related document set, questions, answers and citations; then, samples with low citation quality are eliminated according to the citation quality standard to improve the quality of the training set; finally, irrelevant documents are added to the sample for noise enhancement, and the training set constructed in this way can help the large language model to increase its attention to relevant documents during training, and avoid the large language model from generating answers based on irrelevant documents. The present invention does not require manual annotation, and can automatically generate high-quality attribution text training samples, thereby improving the training efficiency of the large language model used to generate attribution text, reducing the cost of manual annotation, and improving the quality of its output answers and citations, so as to provide reliable answers that better meet user needs and improve the efficiency of information processing in various industries.
[0106] In order to verify the generalization ability of the method of the present invention, the fourth embodiment provided by the present invention selected two types of LLMs as backbone models: LLaMA2, LLaMA3 and Mistral. For the LLaMA2 series, two models of different scales were selected: LLaMA2-7B and LLaMA2-13B. For the Mistral series, Mistral-7B was selected. In most of the main experiments and ablation experiments, ChatGPT (gpt-3.5-turbo-0301) with a 4K context window was used. The results of ChatGPT-16K (gpt-3.5-turbo-16k-0613) and GPT-4 (gpt-4-0613; 8K context window) were also used. The method of the present invention is denoted as .
[0107] This example evaluates LLaMA and its chat version against open-source models. Furthermore, the proposed method is compared with Hagrid, Self-RAG, and Calm. Hagrid is a 3214-sample SFT dataset for attribution text generation; Self-RAG aims to enable large language models to self-reflectively decide whether to retrieve a document; and Calm enables smaller language models to validate the output of larger models.
[0108] This example evaluates the proposed method and all baseline methods on the short text question answering dataset PopQA and two long text question answering datasets ALCE-ASQA and ALCE-ELI5, as well as FanOutQA. Given that the data in this example comes from Wikipedia paragraphs and the ELI5 data is collected from Reddit, the performance of the ELI5 model can alleviate concerns about performance improvements that may be caused by data contamination. In addition, this example also introduces the ArXiver dataset as a data source, which also helps to improve performance and solve this problem. A comprehensive analysis was further conducted by calculating the BLEU scores between all questions in the training set and the questions in the ASQA and PopQA datasets.
[0109] The results show that no question pair has a BLEU score exceeding 0.9, which indicates that there is no significant text overlap between the datasets, thus effectively reducing the risk of data contamination.
[0110] Table 1. Evaluation results of the proposed method and all baseline methods
[0111]
[0112] Table 1 shows the method of the present invention and all baseline methods. With the help of our method, the open source backbone model can achieve excellent citation performance, significantly surpassing the powerful closed-source model. In addition, our method can bring significant performance improvements to all backbone models on all datasets, which fully demonstrates that Ablation experiments on the supervised fine-tuning and PO stages also demonstrate the effectiveness of our method for each stage. The PO stage significantly improves the quality of prudence. Furthermore, integrating data from diverse sources significantly improves performance, effectively validating the generalization capability of our method.
[0113] In order to further analyze the method of the present invention, this embodiment also studies the influence of the filtering threshold on the sample quality. Figure 4 shown. Figure 4 The results show that only higher filtering thresholds can bring benefits, indicating that data quality is more important than data quantity.
[0114] This example also evaluates the impact of different natural language inference models in S4 on sample filtering results. Figure 5 As shown, Figure 5 (a) in the figure is the filtering result of the t5-xxl-true-nli-mixture model. Figure 5(b) in the figure is the filtering result diagram of the ModernBERT-base-nli model. It can be seen that under the same filtering threshold, the performance is improved as the amount of SFT data increases. When the amount of SFT data exceeds 5000, the performance improvement is no longer significant. At the same time, this embodiment found that when the amount of Direct Preference Optimization (DPO) data is less than 9000, there is a clear linear relationship, that is, as the DPO data increases, the performance improves, which fully verifies the effectiveness of the method of the present invention. When the amount of DPO data exceeds 9000, the improvement is not significant. In addition, the ModernBERT-base-nli model may retain more samples while reducing performance. Therefore, the present invention ultimately chooses to use t5-xxl-true-nli-mixture for data filtering.
[0115] In addition, the present invention has also found that using documents of the same category but different subcategories as noise documents can achieve the best performance. If the correlation of the noise documents is too high, the performance may be degraded.
[0116] In this example, 100 randomly selected examples from the resulting training set were manually evaluated. Following Self-RAG, evaluation was conducted along two dimensions: relevance (the appropriateness of the output and its consistency with the topic of the question) and support (the sufficiency of the evidence provided to validate the answer). The results showed that 84% of the examples were classified as relevant and 78% as supportive. These results confirm the high quality of the training set constructed by the present invention. This example was manually evaluated by three annotators, and the Spearman correlation coefficient was used to calculate inter-annotator agreement, resulting in a value of 0.72.
[0117] The method of the present invention can automatically generate high-quality training samples for attributed text generation without manual annotation. Using this method, the Mistral-7B model achieved a citation recall of 84.0 and a precision of 87.0 on ASQA, significantly outperforming GPT-4, which achieved a citation recall of 73.0 and a precision of 76.5. The experimental results of this example also demonstrate the effectiveness of the method of the present invention in reducing the impact of noisy documents and avoiding irrelevant citations.
[0118] Based on the above-mentioned attribution text-based model training method, the present invention also provides an attribution text-based model training system, comprising:
[0119] Related document acquisition module, used to extract relevant document sets from the original data set;
[0120] The sample generation module is used to use the large language model to generate a summary with citation relationships based on all documents in the relevant document set, using the summary as the answer and the citation relationship as the reference of the answer; the large language model is then used to generate a question for the answer, and the corresponding question, answer, reference and related document set constitute a sample;
[0121] The sample filtering module is used to generate multiple samples, calculate the F1 score of each sample, and remove samples with F1 scores lower than the filtering threshold to obtain filtered samples;
[0122] The noise enhancement module is used to randomly select irrelevant documents from the original dataset for each filtered sample, add them to the sample's related document set and shuffle their order, and simultaneously change the reference of the sample's answer to link it to the correct document in the related document set to obtain a noise-enhanced sample;
[0123] The supervised fine-tuning module is used to construct a training set using noise-enhanced samples. It takes the relevant document set and questions of each sample in the training set as input, and uses answers and citations as labels to perform supervised fine-tuning on the large language model to obtain a fully trained large language model.
[0124] The present invention also provides an information retrieval method, comprising:
[0125] The retriever retrieves relevant documents based on the question input by the user to form a relevant document set. The question and the relevant document set are input into the large language model trained by the above-mentioned attribution text-based model training method to obtain the answer and its corresponding reference in the relevant document set.
[0126] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0128] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0130] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A model training method based on attribution text, characterized in that: include: Extract relevant document sets from the original dataset; Using the first language model, one or more summaries with citation relationships are generated based on all documents in a related document set. Each summary is treated as an answer, and the citation relationship of each summary is used as the reference of its answer. Then, using the second language model, a question is generated for each answer. The corresponding question, answer, and reference constitute a question-answer pair. Each question-answer pair and the current related document set constitute a sample. After generating multiple samples, calculate the F1 score of each sample, remove samples with F1 scores lower than the filtering threshold, and obtain filtered samples; For each filtered sample, we randomly select irrelevant documents from the sample's related document set from the original dataset, add them to the sample's related document set and shuffle their order. We also change the reference of the sample's answer so that it links to the correct document in the related document set, thus obtaining a noise-enhanced sample. A training set is constructed using noise-enhanced samples. The relevant document set and questions for each sample in the training set are used as input, and the answers and citations are used as labels. The pre-trained large language model is fine-tuned through supervision to obtain a fully trained large language model.
2. The model training method based on attribution text according to claim 1, characterized in that: Construct unfavorable answers and misquotes for each sample in the training set, and perform preference optimization on the trained large language model, including: The answers and citations of the samples in the training set are the preferred answers and correct citations of the samples; Perform random reference addition, random reference deletion, and random reference change on the correct reference to obtain the incorrect reference; Retain irrelevant documents from the relevant document set in the sample, input the irrelevant documents and sample questions into the large language model, and obtain unfavorable answers and misquotes for the sample; The trained large language model is optimized based on the sample's preferred answers and correct citations, as well as the non-preferred answers and incorrect citations.
3. The model training method based on attribution text according to claim 1, characterized in that: If the original dataset is a structured dataset, extract the relevant document set from the original dataset, including: Randomly select a paragraph from the original dataset as a document and construct a related document set containing only one document; Multiple entity triples are randomly selected from the original data set, and the paragraphs containing the two entities in the entity triples are obtained as documents in the related document set to construct a related document set containing multiple documents.
4. The model training method based on attribution text according to claim 1, characterized in that: If the original dataset is unstructured, extract relevant document sets from the original dataset, including: Classify the documents in the original data set and retain the categories with the number of documents not less than a first preset number; For each retained category, calculate the similarity between all documents in the category and arrange them in descending order; retain the second preset number of document pairs with the highest similarity; For each document pair, the key information and the most overlapping strings of each document are extracted as new documents, and the related document set is constructed with the new documents.
5. The model training method based on attribution text according to claim 1, characterized in that: A large language model is used to generate a summary with citation relationships based on all documents in the relevant document set. The summary is used as the answer, and the citation relationship is used as the reference of the answer. The large language model is then used to generate questions for the answers. The corresponding questions, answers, citations, and their related document sets constitute a sample. The large language model uses the gpt-4-1106-preview model or the Doubao-1.5-Pro-32k model.
6. The model training method based on attribution text according to claim 1, characterized in that: After generating multiple samples, the F1 score of each sample is calculated using the reference quality criteria proposed by ALCE, including: The sample's responses are divided into multiple statements. Each statement is used as a hypothesis, and each reference in the statement is used as a premise. The statement and its reference are input into the natural language inference model for judgment. If there is an implication relationship between the statement and the reference, the reference is marked as a valid reference, otherwise it is an invalid reference. Calculate the ratio of the number of valid citations of each statement to the total number of citations of the statement, and take the average of the ratios as the citation recall rate; The ratio of the total number of valid citations to the total number of citations is used as the citation precision rate; The F1 score of a sample is calculated based on the citation recall and citation precision.
7. The model training method based on attribution text according to claim 6, characterized in that: The natural language inference model is the t5-xxl-true-nli-mixture model.
8. A model training system based on attribution text, characterized in that: include: Related document acquisition module, used to extract relevant document sets from the original data set; The sample generation module is used to use the first language model to generate one or more summaries with citation relationships based on all documents in a related document set, treating each summary as an answer and the citation relationship of each summary as the reference of its answer. The second language model is then used to generate a question for each answer. The corresponding question, answer, and reference constitute a question-answer pair, and each question-answer pair and the current related document set constitute a sample. The sample filtering module is used to generate multiple samples, calculate the F1 score of each sample, and remove samples with F1 scores lower than the filtering threshold to obtain filtered samples; The noise enhancement module is used to randomly select irrelevant documents from the original dataset for each filtered sample, add them to the sample's related document set and shuffle their order, and simultaneously change the reference of the sample's answer to link it to the correct document in the related document set to obtain a noise-enhanced sample; The supervised fine-tuning module is used to construct a training set using noise-enhanced samples. It takes the relevant document set and questions of each sample in the training set as input, and uses answers and citations as labels to perform supervised fine-tuning on the pre-trained large language model to obtain a fully trained large language model.
9. An information retrieval method, characterized in that: include: A retriever retrieves relevant documents according to a question input by a user to form a relevant document set, and inputs the question and the relevant document set into a large language model trained by a model training method based on attribution text as described in any one of claims 1 to 7 to obtain an answer and its corresponding reference in the relevant document set.
Citation Information
Patent Citations
Systems and methods for parameter ensembling for reducing hallucination in abstractive summarization
US20230376677A1
LLM fine-tuning for code generation
US20250094138A1