A method, system, device and medium for fact checking of large language models based on retrieval enhancement

By introducing search enhancement methods in large language models, documents are retrieved and filtered from trusted knowledge bases, and fact units are extracted and optimized, the problem of insufficient accuracy when generating facts in large language models is solved, and efficient and reliable content generation is achieved.

CN119204025BActive Publication Date: 2025-05-13SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411707628.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-13
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Large language models may lack accuracy or authenticity when generating specific facts, and there is "illusion" phenomenon. Traditional information retrieval technology has a long response time, making it difficult to meet the needs of users' immediate feedback.

Method used

Through a search-enhanced method, document search and filtering is performed from a trusted knowledge base, fact units in the original output of the large language model are extracted and semantic enhancement is performed, labels are set for each fact unit based on the searched document set, correction and optimization are performed, and the original output is finally revised using the optimized fact unit.

Benefits of technology

It significantly improves the accuracy and reliability of content generated by large language models, reduces the occurrence of "illusion" phenomena, realizes the entire process from the original output to the final revised output, and improves processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204025B_ABST
    Figure CN119204025B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, device and medium for fact verification of a large language model based on retrieval enhancement, belonging to the technical field of natural language processing, and the method steps are as follows: based on an input task, a document is retrieved from a trusted knowledge base, and a document is screened according to a document evaluation result, and the retrieved document is deleted and compressed according to the relevance of the user input task to obtain a document set; a fact unit is extracted from the original output of the large language model and semantically enhanced; a label is set for the fact unit based on the document set, and a score is given according to the similarity between the fact unit and the document set; the fact unit is corrected, processed and optimized according to the label; the output of the large language model is revised based on the user input task using the fact unit, and the consistency is checked before output. The present invention improves the accuracy and reliability of the content generated by the large language model by introducing the document retrieval and fact verification mechanism of the trusted knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and specifically relates to a method, system, device and medium for performing fact verification of a large language model based on retrieval enhancement. Background Art

[0002] LLM is the abbreviation of Large Language Model.

[0003] NLP is the abbreviation of Natural Language Processing.

[0004] RAG, short for Retrieval-Augmented Generation, is an artificial intelligence technology that combines information retrieval technology with language generation models.

[0005] With the development of artificial intelligence technology, especially the continuous improvement of deep learning algorithms, LLM has become an important function in the field of NLP. Large language models can generate natural and fluent texts through training with a large amount of text data, and have achieved remarkable results in many tasks such as machine translation, text summarization, and question-answering systems. However, although LLM can generate high-quality text, the content generated by the model may lack accuracy or authenticity when specific facts are required. That is, the model may generate information that does not appear in the training data or content that is inconsistent with the facts. This scene is called "hallucination", which is a manifestation of the factual inconsistency of large language models.

[0006] Relevant personnel have explored various ways to improve the accuracy of LLM-generated content. Among them, RAG is an effective means. By integrating the relevant knowledge of the retrieval into the context of LLM, the model can refer to accurate information to generate answers. The application of RAG technology has improved the accuracy of the generated content to a certain extent, but even if the correct context information is retrieved, LLM may still generate incorrect content. In addition, in practical applications, how to integrate the retrieval system with the language model to achieve rapid response while ensuring high accuracy is a challenge. Usually, traditional information retrieval technology often takes a long time to return results and cannot meet the user's needs for instant feedback. Summary of the invention

[0007] In a first aspect, an embodiment of the present application provides a method for performing fact verification of a large language model based on retrieval enhancement, comprising the following steps:

[0008] S1. Retrieve documents from a trusted knowledge base based on the task input by the user, and screen them according to the document evaluation results during the retrieval process, and delete and compress the retrieved documents according to the relevance of the task input by the user after the retrieval is completed to obtain a document set;

[0009] S2. Extract fact units from the raw output of the large language model and perform semantic enhancement;

[0010] S3. Setting a label for each fact unit based on the retrieved document set, and scoring each fact unit according to the similarity between the fact unit and the document set;

[0011] S4. Correct and process the fact unit according to the label, and then optimize the fact unit;

[0012] S5. Use the optimized fact unit to revise the output of the large language model based on the user input task, and output it after the consistency check of the revised result passes.

[0013] Furthermore, the specific steps of step S1 are as follows:

[0014] S11. Analyze the user input task and extract key entities and keywords;

[0015] S12. Retrieve relevant documents from a dynamically updated trusted knowledge base based on key entities and keywords;

[0016] S13. Evaluate the relevance, authority, and update frequency indicators of the retrieved documents, and select a first document set that meets the requirements based on the evaluation results;

[0017] S14. Reorder the documents in the initial document set, and select a second document set whose key entity and keyword relevance is higher than a threshold, then delete redundant information and truncate or merge low-priority content from the documents in the second document set to complete document simplification and compression, and obtain a third document set.

[0018] Furthermore, the specific steps of step S13 are as follows:

[0019] S131. Construct the user input task Q and each retrieved document The semantic similarity of , using cosine similarity and sentence vector embedding to calculate the relevance score through the following formula :

[0020]

[0021] Among them, Embed( ) and Embed( ) represent the vector embedding representation of task input and document respectively;

[0022] S132. Get each retrieved document Sources and citations , determine the trust score of the source , combined with the preset source weights and citation weight The authoritative score Auth is calculated by the following formula: ):

[0023] Auth( )= + ;

[0024] S133. Retrieved documents The update frequency score Rec is calculated based on the time difference ΔT between the release time and the current time using the following formula: ):

[0025]

[0026] Among them, τ is an adjustment parameter used to control the time decay rate;

[0027] S134. For each retrieved document Score by relevance 、Authoritative score Auth( ) and update frequency score Combined with their respective weight coefficients, the comprehensive score is calculated using the following formula :

[0028]

[0029] Among them, α is the relevance score The weight coefficient of , β is the authoritative score Auth( ), γ is the update frequency score The weight coefficient of

[0030] S135. For each retrieved document Based on comprehensive rating The relationship with the set threshold T determines the documents used for fact verification and generates the first document set R:

[0031] ={ ∣ ≥T}.

[0032] Furthermore, the specific steps of step S2 are as follows:

[0033] S21. Decompose the original output of the large language model into independent factual statements without pronouns, perform preprocessing, sentence segmentation, and fact extraction to obtain fact units, and remove duplicates from the fact units;

[0034] S22. Perform dependency syntax analysis, coreference resolution, and logical relationship identification on fact units to complete semantic enhancement and conduct consistency checks.

[0035] Furthermore, the specific steps of step S3 are as follows:

[0036] S31. Setting a label for each fact unit based on the retrieved third document set, wherein the label includes a true label, a false label, and an unmentioned label;

[0037] S32. Calculate the matching degree of each fact unit with the document in the third document set by a semantic similarity algorithm;

[0038] S33. Calculate the support degree of the documents in the third document set for each fact unit according to the matching degree, and generate a credibility score for each fact unit.

[0039] Furthermore, the specific steps of step S4 are as follows:

[0040] S41. Correct the fact unit with the false label and replace it with the content consistent with the document in the third document set;

[0041] S42. Select to retain or delete the fact units with unmentioned labels in order of credibility scores from low to high;

[0042] S43. Learn user preferences by referring to the user's query patterns and feedback data, and adjust the content optimization strategy based on the learning results;

[0043] S44. Based on the corrected fact unit, optimization is performed in combination with the user input task and the third document set to obtain an optimized sentence set, thereby completing the fact content optimization.

[0044] Furthermore, the specific steps of step S5 are as follows:

[0045] S51. Revise the original output of the large language model based on the optimized sentence set and user input tasks and use the content optimization strategy to complete the fact revision;

[0046] S52. Perform consistency check on the revised output and output it after passing the check.

[0047] In a second aspect, an embodiment of the present application further provides a system for performing fact verification of a large language model based on retrieval enhancement, comprising:

[0048] A trusted document retrieval module is used to retrieve documents from a trusted knowledge base based on a task input by a user, and to screen the documents according to the document evaluation results during the retrieval process, and to delete and compress the retrieved documents according to their relevance to the task input by the user after the retrieval is completed, thereby obtaining a document set;

[0049] The fact unit extraction module is used to extract fact units from the original output of the large language model and perform semantic enhancement;

[0050] The fact verification module is used to set a label for each fact unit based on the retrieved document set and score each fact unit according to the similarity between the fact unit and the document set;

[0051] The fact correction and content optimization module is used to correct and process fact units according to labels, and then optimize the fact units;

[0052] The fact revision module is used to revise the output of the large language model based on the user input task using the optimized fact unit, and output the revised result after the consistency check passes.

[0053] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for fact-checking a large language model based on retrieval enhancement as described in the first aspect are implemented.

[0054] In a fourth aspect, an embodiment of the present application further provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for fact-checking a large language model based on retrieval enhancement as described in the first aspect.

[0055] It can be seen from the above technical solutions that the present invention has the following advantages:

[0056] In the method, system, device and medium for fact verification of a large language model based on retrieval enhancement provided by the present application, by retrieving and screening documents from a trusted knowledge base, the high authority and accuracy of the documents are ensured, and a reliable source of information is provided for subsequent fact verification. Furthermore, the original output of the large language model is decomposed into independent fact units, and semantic enhancement is performed, which improves the model's ability to understand and express facts. Through fact verification, a label is set for each fact unit and the similarity is quantified, providing clear guidance and quantitative basis for fact correction and content optimization. The fact units are corrected and optimized according to the labels to ensure the accuracy, readability and fluency of the generated content. Finally, the original output is revised using the optimized fact units, which improves the reliability of the model and enhances the user's satisfaction and trust in the content generated by the large language model. The present invention can be flexibly deployed and applied in various environments and platforms, providing a strong guarantee for the accuracy and credibility of large language models in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0058] Figure 1 The figure is a flow chart of an embodiment of a method for performing fact verification of a large language model based on retrieval enhancement according to the present invention.

[0059] Figure 2 The flowchart of another embodiment of the method for performing fact verification of a large language model based on retrieval enhancement of the present invention is shown.

[0060] Figure 3 A schematic diagram of a system for performing fact verification of a large language model based on retrieval enhancement according to the present invention. DETAILED DESCRIPTION

[0061] In the specific steps of the method for performing large-scale language model fact verification based on retrieval enhancement, which will be described in detail below, various embodiments of the present disclosure will be described more comprehensively. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.

[0062] For example, with the rapid development of artificial intelligence technology, especially the continuous improvement of deep learning algorithms, large language models (LLMs) have become the core force in the field of natural language processing (NLP). These models are trained with massive amounts of text data to create natural and fluent text content, and have demonstrated excellent performance in many application scenarios such as machine translation, text summarization, and question-answering systems. However, although LLMs are able to generate high-quality text, their content may lack accuracy or authenticity when specific facts need to be presented. This phenomenon of introducing information that has never appeared in the training data or is inconsistent with the facts during the generation process is called "hallucination", which reveals the shortcomings of large language models in terms of factuality.

[0063] In order to improve the accuracy of LLM generated content, researchers have explored a variety of strategies, among which the retrieval-augmented generation model (RAG) is an effective method. RAG integrates the retrieved relevant knowledge into the context of LLM, enabling the model to generate answers based on accurate information, thereby improving the accuracy of the generated content to a certain extent. However, even if the correct contextual information is retrieved, LLM may still generate incorrect content based on this information. In addition, in actual deployment, how to efficiently integrate the retrieval system with the language model to achieve fast response and ensure high accuracy has become a difficult problem that needs to be solved urgently. Traditional information retrieval technology often has a long response time and is difficult to meet users' needs for instant feedback.

[0064] In response to the above problems, this embodiment provides a method for fact verification of large language models based on retrieval enhancement. By introducing the document retrieval and fact verification mechanism of a trusted knowledge base, the accuracy and reliability of the content generated by the large language model are significantly improved, and the occurrence of the "hallucination" phenomenon is reduced. At the same time, through document retrieval, fact unit extraction, fact verification, fact correction and content optimization, the whole process from original output to final revised output is automated and intelligently processed, thereby improving processing efficiency and accuracy.

[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] See also Figure 1 The figure is a flowchart of a method for performing fact verification of a large language model based on retrieval enhancement in a specific embodiment, the method comprising the following steps:

[0067] S1. Retrieve documents from a trusted knowledge base based on the task input by the user, and screen them according to the document evaluation results during the retrieval process, and delete and compress the retrieved documents according to the relevance of the task input by the user after the retrieval is completed to obtain a document set;

[0068] Exemplarily, the trusted knowledge base may be a specific database, an authoritative news source, an encyclopedia, or professional literature;

[0069] It should be noted that by performing trusted document retrieval in a trusted knowledge base, it is ensured that the retrieved documents have high authority and accuracy, providing a reliable source of information for subsequent fact verification. In the retrieval process, screening is performed based on the document evaluation results, reducing the interference of irrelevant or low-quality documents, improving retrieval efficiency and accuracy, and deleting and compressing the retrieved documents based on their relevance to the user input task, resulting in a more streamlined and relevant document set, providing a clearer and more focused information basis for subsequent fact verification.

[0070] S2. Extract fact units from the raw output of the large language model and perform semantic enhancement;

[0071] It should be noted that decomposing the original output of the large language model into independent fact units helps to more clearly identify and verify the accuracy of each fact, and semantically enhances the fact units, which improves the model's understanding and expression capabilities of the fact units, and provides a more accurate and in-depth semantic basis for subsequent fact verification;

[0072] S3. Setting a label for each fact unit based on the retrieved document set, and scoring each fact unit according to the similarity between the fact unit and the document set;

[0073] It should be noted that setting a label for each fact unit helps to quickly identify the status of the fact unit, provides clear guidance for subsequent processing, and provides a quantitative basis for fact correction and content optimization through similarity scoring;

[0074] S4. Correct and process the fact unit according to the label, and then optimize the fact unit;

[0075] It should be noted that the fact units are corrected according to the labels to ensure the accuracy of the facts in the generated content, and the fact units are optimized to improve the readability and fluency of the generated content and enhance the user experience;

[0076] S5. Revise the output of the large language model based on the user input task using the optimized fact unit, and output the revised result after the consistency check passes;

[0077] It should be noted that the use of optimized fact units to revise the original output ensures the accuracy of the final output and improves the reliability of the model. The consistency check ensures the consistency of the revised output with the original input task and avoids information loss or misunderstanding.

[0078] The method for fact-checking a large preloaded model based on retrieval enhancement in this embodiment improves the accuracy of the content generated by the large language model and enhances the reliability of the model by introducing document retrieval of a trusted knowledge base.

[0079] Further, as a refinement and expansion of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, as shown below Figure 2 As shown, another method for fact checking of a large language model based on retrieval enhancement is provided, and the method comprises the following steps:

[0080] S1. Based on the user input task, documents are retrieved from the trusted knowledge base, and the documents are screened according to the document evaluation results during the retrieval process. After the retrieval is completed, the documents are deleted and compressed according to the relevance of the retrieved documents to the user input task to obtain a document set. The specific steps of step S1 are as follows:

[0081] S11. Analyze the user input task and extract key entities and keywords;

[0082] S12. Retrieve relevant documents from a dynamically updated trusted knowledge base based on key entities and keywords;

[0083] It should be noted that relevant documents are retrieved from the trusted knowledge base based on key entities and keywords, and a dynamic knowledge base update function is added, so that real-time updated content can be automatically obtained from the knowledge base to ensure the timeliness of the verification content;

[0084] S13. Evaluate the relevance, authority, and update frequency indicators of the retrieved documents, and select a first document set that meets the requirements based on the evaluation results;

[0085] It should be noted that the use of lightweight evaluators can quickly evaluate the relevance, authority, and update frequency of retrieved documents;

[0086] S14. Reorder the documents in the initial document set, and select a second document set whose key entity and keyword relevance is higher than a threshold, and then delete redundant information and truncate or merge low-priority content from the documents in the second document set to complete document simplification and compression to obtain a third document set;

[0087] It should be noted that, based on the first document set after evaluation, the matching degree with the user input task is improved, and the content is compressed to fit the context window length of LLM. Specifically,

[0088] Re-ranking and screening: Screen and re-rank the initial search results R to improve the relevance to the task input X; during the screening process, retain the highly relevant documents containing key entities and keywords, remove the content irrelevant to the task, and obtain the document set after preliminary screening :

[0089]

[0090] Compression: Compress the content of the filtered document set, delete redundant information such as "references" and "footnotes" to ensure that the document content is compact and the core information is clear. By truncating or aggregating low-priority content, a streamlined document R' is obtained:

[0091] ;

[0092] Final Result That is, the document set that has been filtered, reordered, and compressed is used for fact verification and generation enhancement in subsequent steps to ensure that it meets both task requirements and the context window limitations of the model;

[0093] S2. Extract fact units from the original output of the large language model and perform semantic enhancement; the specific steps of step S2 are as follows:

[0094] S21. Decompose the original output of the large language model into independent factual statements without pronouns, perform preprocessing, sentence segmentation, and fact extraction to obtain fact units, and remove duplicates from the fact units;

[0095] Preprocessing is to clean the output generated by LLM to remove extra spaces and special characters;

[0096] Sentence segmentation is to split the output generated by LLM into separate sentences;

[0097] Fact extraction is to extract the most basic fact units from each sentence through natural language processing technology, including but not limited to key information such as names of people, places, organizations, time, quantity, etc.;

[0098] Deduplication is to remove duplicate fact units to ensure that each fact unit is unique. Specifically,

[0099] Assume that the output of the LLM task is , the extracted basic fact unit is S:

[0100] ) = { }

[0101] where n is the number of statements, is the i-th fact unit;

[0102] The output of LLM elements is classified into independent, pronoun-free factual statements to achieve fine-grained segmentation.

[0103] S22. Perform dependency syntax analysis, coreference resolution, and logical relationship identification on fact units, complete semantic enhancement, and perform consistency checks;

[0104] Specifically, dependency syntactic analysis is used to deeply analyze the internal structure of sentences and identify the dependency relationship between words to ensure that the core meaning expressed in the sentence can be accurately extracted; co-reference problems that may exist in the text are solved through co-reference resolution to confirm whether noun phrases in different sentences or within sentences refer to the same entity, ensuring the consistency and fluency of text expression; logical relationship recognition technology is used to capture the logical relationship between sentences, such as cause and effect, transition, etc., to form an overall grasp of the logic of the text; after completing these analyses, consistency checks are performed on the extracted fact units to ensure the logical integrity of these units and that they do not conflict with the overall semantics of the text;

[0105] S3. Set a label for each fact unit based on the retrieved document set, and score each fact unit according to the similarity between the document set and the fact unit; the specific steps of step S3 are as follows:

[0106] S31. Setting a label for each fact unit based on the retrieved third document set, wherein the label includes a true label, a false label, and an unmentioned label;

[0107] Specifically, based on the retrieved trusted document content, each extracted fact unit is verified one by one, and the units that meet the requirements are marked as "true", the units with contradictions are marked as "false", and the units that cannot be verified in the retrieval are marked as "not mentioned". Only the fact units that meet the verification are retained to ensure consistency with the retrieved document;

[0108] S32. Calculate the matching degree of each fact unit with the document in the third document set by a semantic similarity algorithm;

[0109] Specifically, the matching degree between each fact unit and the search document is calculated through the semantic similarity algorithm, and the accuracy of the verification is further judged to ensure that each fact unit has undergone a high standard of semantic matching verification;

[0110] S33. Calculate the support degree of the document in the third document set for each fact unit according to the matching degree, and generate a credibility score for each fact unit;

[0111] Specifically, a credibility score is generated for each fact unit, and the degree of support for each unit is calculated based on the trusted documents to ensure high credibility and data support for the output results;

[0112] S4. Correct and process the fact unit according to the label, and then optimize the fact unit; the specific steps of step S4 are as follows:

[0113] S41. Correct the fact unit with the false label and replace it with the content consistent with the document in the third document set;

[0114] Specifically, fact units marked as “false” are corrected and replaced with content consistent with the retrieved document to maintain the authenticity of the content;

[0115] S42. Select to retain or delete the fact units with unmentioned labels in order of credibility scores from low to high;

[0116] Specifically, for the “not mentioned” factual units, if they can remain consistent with the context, they will be appropriately retained; otherwise, they will be deleted, with simplicity and coherence as the core, ensuring that the text is concise and easy to understand;

[0117] S43. Learn user preferences by referring to the user's query patterns and feedback data, and adjust the content optimization strategy based on the learning results;

[0118] Specifically, by designing a user preference learning module, referring to the user's query patterns and feedback data, the content optimization strategy is gradually adjusted to make the content more in line with user needs;

[0119] S44. Based on the corrected fact unit, the optimization is performed in combination with the user input task and the third document set to obtain an optimized statement set, thereby completing the fact content optimization;

[0120] Specifically, on the basis of correction, the generated content is further adjusted to improve the language fluency and naturalness while meeting the requirements of the target task. The task input Q is maintained throughout the correction process to prevent deviation from the initial task goal.

[0121] During correction, the task input Q is included to avoid divergence from the task input. The corrected statement C is:

[0122] ;

[0123] S5. Use the optimized fact unit to revise the output of the large language model based on the user input task, and output it after the consistency check of the revised result passes; the specific steps of step S5 are as follows:

[0124] S51. Revise the original output of the large language model based on the optimized sentence set and user input tasks and use the content optimization strategy to complete the fact revision;

[0125] Specifically, use the corrected sentence set C to output the original task After optimization, the task input Q is retained during the revision process to ensure consistency with the task input. The final revised output O is expressed as follows:

[0126] ;

[0127] S52. Perform consistency check on the revised output and output it after passing the check;

[0128] It should be noted that the content of consistency checking is to check whether the corrected fact unit completely matches the requirements of the original task input, and whether there are any logical errors, omissions or duplicate information in the output.

[0129] In an embodiment of the present invention, based on step S13, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.

[0130] The specific steps of step S13 are as follows:

[0131] S131. Construct the user input task Q and each retrieved document The semantic similarity of , using cosine similarity and sentence vector embedding to calculate the relevance score through the following formula :

[0132]

[0133] Among them, Embed( ) and Embed( ) represent the vector embedding representation of task input and document respectively;

[0134] It should be noted that documents with a relevance score above the set threshold are considered relevant;

[0135] S132. Get each retrieved document Sources and citations , determine the trust score of the source , combined with the preset source weights and citation weight The authoritative score Auth is calculated by the following formula: ):

[0136] Auth( )= + ;

[0137] It should be noted that documents with lower authority will be filtered out;

[0138] S133. Retrieved documents The update frequency score is calculated based on the time difference ΔT between the release time and the current time using the following formula :

[0139]

[0140] Among them, τ is an adjustment parameter used to control the time decay rate;

[0141] It should be noted that the closer the value of the adjustment parameter τ is to 1, the newer the document is;

[0142] S134. For each retrieved document Score by relevance 、Authoritative score Auth( ) and update frequency score Combined with their respective weight coefficients, the comprehensive score is calculated using the following formula :

[0143]

[0144] Among them, α is the relevance score The weight coefficient of , β is the authoritative score Auth( ), γ is the weight coefficient of the update frequency score Rec( )’s weight coefficient;

[0145] It should be noted that the values ​​of α, β, and γ can be adjusted according to task requirements;

[0146] S135. For each retrieved document Based on comprehensive rating The relationship with the set threshold T determines the documents used for fact verification and generates the first document set R:

[0147] ={ ∣ )≥T}.

[0148] For example, a large language model (LLM) needs to generate a report on "global climate change", which needs to accurately reflect the scientific facts and data of current global climate change. In order to ensure the accuracy of the report, the method of the present invention is used to perform fact verification.

[0149] Specifically,

[0150] 1. Document retrieval and screening

[0151] Tasks that obtain user input - need to generate a report on global climate change, extract key entities and keywords such as "global climate change", "greenhouse gas emissions", and "global temperature rise", retrieve relevant documents from dynamically updated trusted knowledge bases such as IPCC reports, NASA data, NOAA data, etc., use cosine similarity and sentence vector embedding to calculate the semantic similarity between each document and the task input, screen out documents with high relevance, calculate the authority score of each document based on the trust score and citation status of the document source, calculate the update frequency score of each document based on the time difference between the document's release time and the current time, calculate the comprehensive score of each document based on the relevance, authority, and update frequency scores, and screen out documents with a comprehensive score higher than the set threshold to form a first document set, reorder the documents in the first document set, screen out a second document set with key entity and keyword relevance higher than the threshold, and then delete redundant information and truncate or merge low-priority content from the documents in the second document set to complete document simplification and compression to obtain a third document set.

[0152] 2. Fact Unit Extraction and Semantic Enhancement

[0153] Decompose the original output of LLM into independent factual statements without pronouns, and perform preprocessing, sentence segmentation, and fact extraction to obtain fact units, such as fact unit A: Since the 19th century, the global average temperature has risen by about 1 degree Celsius; fact unit B: The frequency and intensity of extreme weather events such as heat waves, droughts, and floods have increased in the past few decades; fact unit C: The emission of greenhouse gases such as carbon dioxide is one of the main causes of global warming; fact unit D: Glaciers in some areas are melting rapidly, causing sea level rise; then remove duplicates from the fact units, perform dependency syntax analysis, coreference resolution, and logical relationship identification on the fact units, complete semantic enhancement, and perform consistency checks;

[0154] 3. Fact Verification and Labeling

[0155] Based on the third document set, a label is set for each fact unit, including true labels, false labels and unmentioned labels, and the matching degree between each fact unit and the documents in the third document set is calculated by a semantic similarity algorithm. The support degree of the documents in the third document set for each fact unit is calculated according to the matching degree, and a credibility score is generated for each fact unit. For example, for fact unit A, if the fact unit comes from an internationally authoritative climate research institution and has been verified many times, then its credibility score may be very high, such as 90 points; for fact unit B, if it is based on the results of multiple independent studies and the consistency between these studies is very high, then its credibility score will also be relatively high; for fact unit C, if it is based on a large amount of scientific evidence and model predictions, then its credibility will also be very high; for fact unit D, if it comes from field observations and long-term monitoring data, then its credibility will also be relatively high.

[0156] IV. Correction and Optimization of Fact Units

[0157] The fact units with false labels are corrected and replaced with content consistent with the documents in the third document set. The fact units with unmentioned labels are selected for retention or deletion in order of credibility score from low to high. User preference learning is performed with reference to user query patterns and feedback data. The content optimization strategy is adjusted according to the learning results. The corrected fact units are optimized in combination with user input tasks and the third document set to obtain an optimized statement set. For example, when processing fact unit A, it is necessary to further verify its data source and calculation method; when processing fact unit B, it is necessary to integrate and compare the results of multiple studies; when processing fact unit C, it is necessary to pay attention to the latest scientific progress and research results; when processing fact unit D, it is necessary to pay attention to the latest observation data and trend analysis of glacier melting.

[0158] 5. Output Revision and Consistency Check

[0159] Based on the optimized statement set and user input tasks and using content optimization strategies, the original output of LLM is revised, and the revised output is checked for consistency to ensure that the content is consistent with the information in the third document set. After the verification is passed, the output is produced to form the final "Global Climate Change" report.

[0160] Ultimately, the generated "Global Climate Change" report successfully avoided the "illusion" problem that may be caused by LLM, ensuring that the facts in the report are accurate and reliable. At the same time, the report was optimized based on the user's query patterns and feedback data, improving user satisfaction and trust.

[0161] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0162] like Figure 3 As shown, the following is an embodiment of the system for performing fact verification of a large language model based on retrieval enhancement provided by an embodiment of the present disclosure. The system and the method for performing fact verification of a large language model based on retrieval enhancement in the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the system for performing fact verification of a large language model based on retrieval enhancement, reference can be made to the embodiment of the method for performing fact verification of a large language model based on retrieval enhancement in the above-mentioned embodiments.

[0163] The system includes:

[0164] A trusted document retrieval module is used to retrieve documents from a trusted knowledge base based on a task input by a user, and to screen the documents according to the document evaluation results during the retrieval process, and to delete and compress the retrieved documents according to their relevance to the task input by the user after the retrieval is completed, thereby obtaining a document set;

[0165] The fact unit extraction module is used to extract fact units from the original output of the large language model and perform semantic enhancement;

[0166] The fact verification module is used to set a label for each fact unit based on the retrieved document set and score each fact unit according to the similarity between the fact unit and the document set;

[0167] The fact correction and content optimization module is used to correct and process fact units according to labels, and then optimize the fact units;

[0168] The fact revision module is used to revise the output of the large language model based on the user input task using the optimized fact unit, and output the revised result after the consistency check passes.

[0169] The system for fact verification of a large language model based on retrieval enhancement provided in this embodiment ensures the high authority and accuracy of the document by retrieving and screening the document from a trusted knowledge base, and provides a reliable source of information for subsequent fact verification. Furthermore, the original output of the large language model is decomposed into independent fact units, and semantic enhancement is performed, which improves the model's ability to understand and express facts. Through fact verification, a label is set for each fact unit and the similarity is quantified, providing clear guidance and quantitative basis for fact correction and content optimization. The fact units are corrected and optimized according to the labels to ensure the accuracy, readability and fluency of the generated content. Finally, the original output is revised using the optimized fact units, which improves the reliability of the model and enhances the user's satisfaction and trust in the content generated by the large language model. The present invention can be flexibly deployed and applied in various environments and platforms, providing a strong guarantee for the accuracy and credibility of large language models in practical applications.

[0170] The method for fact-checking a large language model based on retrieval enhancement provided in the embodiment of the present application can be applied to electronic devices. It will be appreciated by those skilled in the art that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or less components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0171] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, buttons, a camera, a display, and a SIM card interface, etc.

[0172] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0173] The processor may include one or more processing units, for example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated into one or more processors.

[0174] The processor can be the nerve center and command center of the electronic device. The controller can generate an operation control signal according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0175] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. The memory may store instructions or data that the processor has just used or is cyclically used. If the processor needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor, and thus improves system efficiency.

[0176] The above-mentioned electronic device implements the method of the present application for fact verification of large language models based on retrieval enhancement, which significantly improves the accuracy and reliability of the content generated by large language models by introducing the document retrieval and fact verification mechanism of the trusted knowledge base, and reduces the occurrence of "hallucination" phenomena. At the same time, through document retrieval, fact unit extraction, fact verification, fact correction and content optimization, the whole process from original output to final revised output is automated and intelligently processed, thereby improving processing efficiency and accuracy.

[0177] The electronic device provided in this embodiment stores a method for fact verification of a large language model based on retrieval enhancement. By introducing a document retrieval and fact verification mechanism of a trusted knowledge base, the accuracy and reliability of the content generated by the large language model are significantly improved, and the occurrence of the "hallucination" phenomenon is reduced. At the same time, through document retrieval, fact unit extraction, fact verification, fact correction and content optimization, the whole process from original output to final revised output is automated and intelligently processed, thereby improving processing efficiency and accuracy.

[0178] The storage medium provided in the present application stores a program product that can implement a method for fact-checking a large language model based on retrieval enhancement.

[0179] The method for fact verification of a large language model based on retrieval enhancement includes: retrieving documents from a trusted knowledge base based on a user input task, screening according to document evaluation results during the retrieval process, and deleting and compressing the retrieved documents according to the relevance to the user input task after the retrieval is completed to obtain a document set; extracting fact units from the original output of the large language model and performing semantic enhancement; setting labels for each fact unit based on the retrieved document set, and scoring each fact unit according to the similarity between the fact unit and the document set; correcting and processing the fact units according to the labels, and then optimizing the fact units; using the optimized fact units to revise the output of the large language model based on the user input task, and outputting the revised results after the consistency check passes.

[0180] In some possible implementations, the method for performing large-scale language model fact verification based on retrieval enhancement of the present disclosure may be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above “Exemplary Method” section of this specification according to various exemplary implementations of the present disclosure.

[0181] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0182] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for fact checking of a large language model based on retrieval enhancement, characterized in that: The steps include: S1. Based on the user input task, documents are retrieved from the trusted knowledge base, and the documents are screened according to the document evaluation results during the retrieval process. After the retrieval is completed, the documents are deleted and compressed according to the relevance of the retrieved documents to the user input task to obtain a document set. The specific steps of step S1 are as follows: S11. Analyze the user input task and extract key entities and keywords; S12. Retrieve relevant documents from a dynamically updated trusted knowledge base based on key entities and keywords; S13. Evaluate the relevance, authority and update frequency indicators of the retrieved documents, and select the first document set that meets the requirements based on the evaluation results; obtain the source and citation status of each retrieved document, determine the trust score of the source, and then calculate the authority score based on the preset source weight and citation weight; calculate the update frequency score of the retrieved documents based on the time difference between the release time and the current time; S14. Reorder the documents in the first document set, select the second document set whose key entity and keyword relevance is higher than the threshold, delete redundant information and truncate or merge low-priority content from the documents in the second document set, complete document simplification and compression, and obtain a third document set; S2. Extract fact units from the raw output of the large language model and perform semantic enhancement; S3. Setting a label for each fact unit based on the retrieved document set, and scoring each fact unit according to the similarity between the fact unit and the document set; S4. Correct and process the fact unit according to the label, and then optimize the fact unit; S5. Use the optimized fact unit to revise the output of the large language model based on the user input task, and output it after the consistency check of the revised result passes.

2. The method for fact checking of a large language model based on retrieval enhancement according to claim 1, characterized in that: The specific steps of step S13 are as follows: S131. Construct the user input task Q and each retrieved document The semantic similarity of , using cosine similarity and sentence vector embedding to calculate the relevance score through the following formula : Among them, Embed( ) and Embed( ) represent the vector embedding representation of task input and document respectively; S132. Get each retrieved document Sources and citations , determine the trust score of the source , combined with the preset source weights and citation weight The authoritative score Auth is calculated by the following formula: ): Auth( )= + ; S133. Retrieved documents The update frequency score is calculated based on the time difference ΔT between the release time and the current time using the following formula : Among them, τ is an adjustment parameter used to control the time decay rate; S134. For each retrieved document Score by relevance 、Authoritative score Auth( ) and update frequency score Combined with their respective weight coefficients, the comprehensive score is calculated using the following formula : Among them, α is the relevance score The weight coefficient, β is the authoritative score Auth( ), γ is the update frequency score The weight coefficient of S135. For each retrieved document Based on comprehensive rating The relationship with the set threshold T determines the documents used for fact verification and generates the first document set R: ={ ∣ ≥T}。 3. The method for fact checking of a large language model based on retrieval enhancement according to claim 1, characterized in that: The specific steps of step S2 are as follows: S21. Decompose the original output of the large language model into independent factual statements without pronouns, perform preprocessing, sentence segmentation, and fact extraction to obtain fact units, and remove duplicates from the fact units; S22. Perform dependency syntax analysis, coreference resolution, and logical relationship identification on fact units to complete semantic enhancement and conduct consistency checks.

4. The method for fact checking of a large language model based on retrieval enhancement as claimed in claim 3, characterized in that: The specific steps of step S3 are as follows: S31. Setting a label for each fact unit based on the retrieved third document set, wherein the label includes a true label, a false label, and an unmentioned label; S32. Calculate the matching degree of each fact unit with the document in the third document set by a semantic similarity algorithm; S33. Calculate the support degree of the documents in the third document set for each fact unit according to the matching degree, and generate a credibility score for each fact unit.

5. The method for fact checking of a large language model based on retrieval enhancement as claimed in claim 4, characterized in that: The specific steps of step S4 are as follows: S41. Correct the fact unit with the false label and replace it with the content consistent with the document in the third document set; S42. Select to retain or delete the fact units with unmentioned labels in order of credibility scores from low to high; S43. Learn user preferences by referring to the user's query patterns and feedback data, and adjust the content optimization strategy based on the learning results; S44. Based on the corrected fact unit, optimization is performed in combination with the user input task and the third document set to obtain an optimized sentence set, thereby completing the fact content optimization.

6. The method for fact checking of a large language model based on retrieval enhancement according to claim 5, characterized in that: The specific steps of step S5 are as follows: S51. Revise the original output of the large language model based on the optimized sentence set and user input tasks and use the content optimization strategy to complete the fact revision; S52. Perform consistency check on the revised output and output it after passing the check.

7. A system for fact checking of large language models based on retrieval enhancement, characterized in that: include: The trusted document retrieval module retrieves documents from the trusted knowledge base based on the user input task, and screens them according to the document evaluation results during the retrieval process. After the retrieval is completed, the retrieved documents are deleted and compressed according to the relevance of the user input task to obtain a document set. The specific process of obtaining the document set is as follows: Analyze user input tasks and extract key entities and keywords; Retrieve relevant documents from a dynamically updated trusted knowledge base based on key entities and keywords; Evaluate the relevance, authority and update frequency indicators of the retrieved documents, and select the first document set that meets the requirements based on the evaluation results; obtain the source and citation status of each retrieved document, determine the source trust score, and then calculate the authority score based on the preset source weight and citation weight; calculate the update frequency score of the retrieved documents based on the time difference between the release time and the current time; Reorder the documents in the first document set, select the second document set whose key entity and keyword relevance is higher than a threshold, delete redundant information and truncate or merge low-priority content from the documents in the second document set, complete document simplification and compression, and obtain a third document set; The fact unit extraction module extracts fact units from the original output of the large language model and performs semantic enhancement; The fact verification module sets a label for each fact unit based on the retrieved document set and scores each fact unit according to its similarity with the document set; The fact correction and content optimization module corrects and processes fact units according to labels, and then optimizes the fact units; The fact revision module uses optimized fact units to revise the output of the large language model based on the user input task, and outputs the revised result after the consistency check passes.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for performing fact verification of a large language model based on retrieval enhancement as described in any one of claims 1 to 6 are implemented.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for performing fact verification of a large language model based on retrieval enhancement as claimed in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Pseudo-correlation feedback model information retrieval method and system based on BERT

    CN110442777A

  • Retrieval-Augmented Generation Solution System

    KR102732291B1