Intelligent dialogue method, system and device based on big language model illusion relief and medium

The large language model generates multiple rounds of inference results and uses small language models for logical semantic detection, which solves the problem of large language model hallucination in intelligent dialogue, and achieves the reliability and accuracy of the results, which is suitable for complex question-and-answer scenarios.

CN120258156AActive Publication Date: 2025-07-04BEIJING NORMAL UNIVERSITY

Patent Information

Application Number
CN202510759595.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-04
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Large language models are prone to hallucinations in intelligent dialogue scenarios, resulting in inaccurate output results, and it is difficult for existing methods to objectively and comprehensively detect and correct hallucinations.

Method used

Multiple rounds of heuristic answers, stimulating thinking and self-reflection results are generated through large language models, and logical semantic relationship detection is carried out in combination with small language models, error information is filtered, and inference chains are formed with incremental enhanced logic and semantic correct target answers are extracted.

Benefits of technology

The phased collaboration between large language models and small language models is realized, the limitations of self-reflection detection hallucinations of large language models are overcome, and the reliability and accuracy of the generated results are ensured. It is suitable for complex question-and-answer scenarios that require multi-step logical derivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258156A_ABST
    Figure CN120258156A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an intelligent dialogue method, system and device based on big language model illusion alleviation and a medium, and the method comprises the steps: outputting multiple rounds of heuristic answer results, stimulation thinking results and self-reflection results based on an input question through a big language model; generating a reasoning chain based on the heuristic answer result, the stimulation thinking result and the self-reflection result of the same round; performing logic semantic relationship detection on each reasoning chain through a small language model to obtain a target reasoning chain with a correct relationship; and extracting and outputting a target answer of the input question from the target reasoning chain. According to the intelligent dialogue method based on big language model illusion relief, the big language model is used for forming an inference chain, the small language model is used for analyzing the output of the big language model, inaccurate information can be accurately recognized and filtered out, staged cooperation between the two models is achieved, the limitation that the big language model conducts illusion detection in a self-reflection mode is overcome, and the intelligent dialogue efficiency is improved. And a result can be accurately output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an intelligent dialogue method, system, device and medium based on alleviating large language model hallucinations. Background Art

[0002] In recent years, large language models (LLMs) have demonstrated certain capabilities in complex reasoning tasks. Through chain of thought (CoT) prompting for multi-step reasoning, they can generate coherent and contextually relevant responses. However, LLMs have obvious drawbacks and are prone to hallucinations, that is, generating outputs that seem reasonable but are actually incorrect, which seriously affects the accuracy of the results.

[0003] In contrast, small language models based on bidirectional encoder representations from transformers (BERT) and its variants, etc., although more efficient in using computing resources and can give accurate outputs for specific tasks, have a narrow reasoning scope and perform poorly when dealing with complex multi-step problems that require in-depth context understanding.

[0004] Existing methods for solving LLMs hallucinations, such as self-consistency checking, retrieval-augmented generation, and confidence calibration, etc., but these methods all rely heavily on the self-reflection of LLMs and are difficult to objectively and comprehensively detect and correct hallucinations. The self-reflection ability of LLMs is based on its existing knowledge and learning patterns, and there are natural cognitive limitations. It cannot go beyond its training data and learning experience to discover deep-seated errors or hallucinations, and can only check and adjust within the known framework. This situation will lead to inaccurate output results in the intelligent dialogue scenario. Summary of the Invention

[0005] The present invention provides an intelligent dialogue method, system, device and medium based on alleviating large language model hallucinations to solve the problem in the prior art that large language model hallucinations will lead to inaccurate output results in the intelligent dialogue scenario.

[0006] In a first aspect, the present invention provides an intelligent dialogue method based on alleviating large language model hallucinations, including: Through a large language model, based on the input question, multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results are output; among them, the i-th round of heuristic answer results is generated based on the input question, the i-th round of stimulating thinking results is generated based on the input question and the i-th round of heuristic answer results, and the i-th round of self-reflection results is generated based on the input question, the i-th round of heuristic answer results, and the i-th round of stimulating thinking results; i is a positive integer; Based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round, corresponding inference chains are generated; Through a small language model, logical semantic relationship detection is performed on each inference chain to obtain multiple target inference chains with correct logical semantic relationships; Extract the target answer to the input question from the multiple target inference chains and output the target answer.

[0007] In one embodiment, the step of performing logical semantic relationship detection on each inference chain through a small language model to obtain multiple target inference chains with correct logical semantic relationships includes: Through a small language model, logical semantic relationship detection is performed on each inference chain to obtain the logical semantic relationship detection result of each inference chain; Score based on the logical semantic relationship detection result of each inference chain to obtain the logical semantic relationship score of each inference chain; among them, when the logical semantic relationship of the inference chain is incorrect, the logical semantic relationship score of the inference chain is 0, and when the logical semantic relationship of the inference chain is correct, the logical semantic relationship score of the inference chain is 1; Determine the inference chains with the product of all logical semantic relationship scores being 1 as the target inference chains with correct logical semantic relationships.

[0008] In one embodiment, the logical semantic relationship detection result includes the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result; When performing logical semantic relationship detection on each inference chain through a small language model to obtain the logical semantic relationship detection result of each inference chain, the following steps are executed for each inference chain: Through a small language model, using natural language inference technology, perform logical semantic relationship detection on the input question and the heuristic answer result in the inference chain to obtain the first logical semantic relationship detection result between the input question and the heuristic answer result; Through a small language model, using natural language inference technology, perform logical semantic relationship detection on the heuristic answer result and the stimulating thinking result in the inference chain to obtain the second logical semantic relationship detection result between the heuristic answer result and the stimulating thinking result; Using a small language model and natural language inference technology, perform logical semantic relationship detection on the results of stimulated thinking and self-reflection in the reasoning chain to obtain the third logical semantic relationship detection result between the results of stimulated thinking and the results of self-reflection.

[0009] In one embodiment, when scoring based on the logical semantic relationship detection results of each reasoning chain to obtain the logical semantic relationship score of each reasoning chain, the following steps are performed for each reasoning chain: If the first logical semantic relationship detection result is an entailment relationship, and the second logical semantic relationship detection result is an entailment relationship, and the third logical semantic relationship detection result is an entailment relationship, then determine that the logical semantic relationship score of the reasoning chain is 1; If any one of the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result is a neutral relationship or a contradiction relationship, then determine that the logical semantic relationship score of the reasoning chain is 0.

[0010] In one embodiment, the extracting the target answer to the input question from the multiple target reasoning chains and outputting the target answer includes: Using a regular expression, extract the target information for answering the input question from each target reasoning chain; If the target information is successfully extracted from any target reasoning chain, then standardize the format of the target information to obtain the initial answer corresponding to the target reasoning chain; If the target information is not successfully extracted from any target reasoning chain, then use a small language model fine-tuned for the task to generate the target information for answering the input question based on the target reasoning chain, and standardize the format of the target information to obtain the initial answer corresponding to the target reasoning chain; Perform consistency evaluation based on each of the initial answers to obtain the target answer to the input question, and output the target answer.

[0011] In one embodiment, the if the target information is successfully extracted from any target reasoning chain, then standardize the format of the target information to obtain the initial answer corresponding to the target reasoning chain includes: If the target information is successfully extracted from any target reasoning chain, then standardize the format of the target information to obtain the format-standardized information; Use a small language model fine-tuned for the task to verify the reliability of the format-standardized information; If the verification is successful, then determine the format-standardized information as the initial answer corresponding to the target reasoning chain; If the verification fails, a small language model fine-tuned by tasks generates an initial answer corresponding to the target reasoning chain based on the target reasoning chain.

[0012] In one embodiment, the obtaining of the target answer for answering the input question based on each of the initial answers includes: Determining the number of occurrences of the same answer among each of the initial answers; If the highest number of occurrences is greater than or equal to a preset occurrence threshold, determining the answer corresponding to the highest number of occurrences as the target answer for answering the input question.

[0013] In a second aspect, the present invention further provides an intelligent dialogue system for alleviating large language model hallucinations, including: A model reasoning module for outputting multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results based on an input question through a large language model; wherein, the i-th round of heuristic answer result is generated based on the input question, the i-th round of stimulating thinking result is generated based on the input question and the i-th round of heuristic answer result, and the i-th round of self-reflection result is generated based on the input question, the i-th round of heuristic answer result, and the i-th round of stimulating thinking result; A reasoning chain generation module for generating corresponding reasoning chains based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round; A logical semantic relationship detection module for detecting the logical semantic relationships of each reasoning chain through a small language model to obtain multiple target reasoning chains with correct logical semantic relationships; A target answer output module for extracting the target answer for answering the input question from the multiple target reasoning chains and outputting the target answer.

[0014] In a third aspect, the present invention provides an electronic device, where the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps of any one of the above-mentioned intelligent dialogue methods for alleviating large language model hallucinations.

[0015] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any one of the above-mentioned intelligent dialogue methods for alleviating large language model hallucinations.

[0016] The intelligent dialogue method, system, device, and medium for alleviating hallucinations of large language models provided by the present invention utilize large language models to first generate heuristic answer results to expand the reply perspective, then stimulate in-depth reasoning, and finally conduct self-reflection to correct potential problems, forming a progressively enhanced reasoning chain. Further, a small language model is used to analyze the output of the large language model from an independent perspective, and logical semantic detection is performed on each reasoning chain, which can accurately identify and filter out inaccurate information therein, realizing the phased collaboration of the large language model and the small language model, combining the reasoning depth of the large language model with the accuracy and efficiency of the small language model, overcoming the limitations of the large language model's self-reflection for detecting hallucinations, ensuring the reliability of the generated results, and finally extracting the target answer from the target reasoning chain that conforms to the logical semantic relationship, capable of accurately and efficiently outputting results, especially suitable for complex question-and-answer scenarios that require multi-step logical reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is one of the schematic flowcharts of the intelligent dialogue method for alleviating hallucinations of large language models provided by the present invention.

[0019] Figure 2 It is the second schematic flowchart of the intelligent dialogue method for alleviating hallucinations of large language models provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the accuracy ratio of different rounds of dialogue of the MS-HM collaborative method provided by the present invention on Qwen-14B Chat and GPT-3.5-turbo.

[0021] Figure 4 It is a schematic diagram of the qualitative results of the traditional reasoning method in physical tasks provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the structure of the intelligent dialogue system for alleviating hallucinations of large language models provided by the present invention.

[0023] Figure 6 It is a schematic diagram of the structure of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.

[0025] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein.

[0026] The following will be combined with Figures 1-6 to describe the intelligent dialogue method, system, device, and medium provided by the present invention for alleviating hallucinations of large language models.

[0027] It should be noted that the intelligent dialogue method for alleviating hallucinations of large language models provided in the embodiments of the present invention is implemented based on an intelligent dialogue system for alleviating hallucinations of large language models. The intelligent dialogue method for alleviating hallucinations of large language models provided in the embodiments of the present invention effectively alleviates the hallucination problem of large language models, improves the reasoning efficiency and accuracy, combines the reasoning depth of large language models with the accuracy and efficiency of small language models through three key stages: large language model-guided reasoning, small language model hallucination detection, and reply result standardization, realizes structured output, and gives full play to the synergistic advantages of large language models and small language models.

[0028] The embodiments of the present invention describe the intelligent dialogue method for alleviating hallucinations of large language models with an intelligent dialogue system for alleviating hallucinations of large language models as the execution subject.

[0029] Combined with Figure 1 and Figure 2 , Figure 1 is one of the flow diagrams of the intelligent dialogue method for alleviating hallucinations of large language models provided by the present invention, Figure 2 is the second flow diagram of the intelligent dialogue method for alleviating hallucinations of large language models provided by the present invention.

[0030] As Figure 1 shown, the intelligent dialogue method for alleviating hallucinations of large language models includes the following steps: Step 101: Based on the input question, output multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results through a large language model; Step 102: Generate corresponding reasoning chains based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round; Step 103: Use a small language model to detect the logical semantic relationships of each reasoning chain and obtain multiple target reasoning chains with correct logical semantic relationships; Step 104: Extract the target answer to the input question from the multiple target reasoning chains and output the target answer.

[0031] Specifically, the user inputs a question through an interactive page. The input question corresponds to a specific task, and the specific task is not limited here. Based on the input question Q, the language understanding ability and reasoning ability of the large language model are used to output an answer.

[0032] First, through the large language model, based on the input question Q, output multi-round heuristic answer results . It can be understood that the large language model generates multi-round independent answers for the input question Q, and each answer does not depend on the previous generation history. The heuristic answer result is an answer presented in a step-by-step, multi-angle, and guiding logic generated by the large language model based on the input question Q, rather than a direct final conclusion.

[0033] By modeling the generation process of multi-round heuristic answer results : .

[0034] Then, through the large language model, stimulate thinking based on the multi-round heuristic answer results and the input question Q to generate multi-round stimulating thinking results .

[0035] By modeling the generation process of multi-round stimulating thinking results : .

[0036] Finally, through the large language model, conduct self-reflection based on the multi-round stimulating thinking results , multi-round heuristic answer results and the input question Q to generate multi-round self-reflection results .

[0037] By modeling the generation process of multi-round self-reflection results : .

[0038] In actual operation, the value of the number of rounds k can be set according to the complexity of the task and the computing resources. For example, for simple tasks, it can be set to 2 - 3 rounds; for complex tasks, the value of k can be appropriately increased. The multi-round dialogue mechanism helps the model better understand the context information and improve the accuracy of answers. Although for some complex tasks, the effect of increasing the number of dialogue rounds needs to be further optimized, overall, this method can adapt to the complexity of different tasks and provide a comprehensive, reliable and flexible solution for complex reasoning tasks.

[0039] To simplify the calculation, it can be assumed that there is conditional independence in each step, that is is generated only depending on , is generated only depending on and Q, is generated only depending on , and Q. At this stage, the large language model decomposes complex tasks into manageable parts and generates initial reasoning steps and potential solutions.

[0040] Furthermore, the heuristic answer results, thought-stimulating results, and self-reflection results of the same round are concatenated to generate corresponding reasoning chains, obtaining k reasoning chains. That is: .

[0041] The above process is Figure 2 the guided reasoning process based on the large language model in

[0042] Furthermore, led by the small language model, using Natural Language Inference (NLI) technology, the logical semantic relationship of each reasoning chain is detected to obtain the logical semantic relationship detection result of each reasoning chain. Then, based on the logical semantic relationship detection result of each reasoning chain, the reasoning chains with problems in the logical semantic relationship are filtered out, and only the reasoning chains with normal logical semantic relationships are retained as the target reasoning chains, obtaining j target reasoning chains. That is to say, by executing the NLI task with the small language model, the error information or contradictory information in the content generated by the large language model can be filtered out, and only the correct information in the content generated by the large language model is retained, effectively alleviating the hallucination problem of the large language model.

[0043] The above process is Figure 2 the hallucination mitigation process based on the small language model in

[0044] Furthermore, by combining rule-based methods and small language models, target information for answering the input question is extracted from multiple target inference chains, and it is standardized to meet the requirements of downstream tasks. Then, the consistency of the standardized information is evaluated, and the target answers that meet the consistency requirements are output for downstream tasks.

[0045] The above process is Figure 2 the process of standardizing the response results based on rule-based and small language models in

[0046] The intelligent dialogue method for alleviating large language model hallucinations provided by the present invention uses a large language model to first generate heuristic answer results to expand the response perspective, then stimulates in-depth reasoning, and finally conducts self-reflection to correct potential problems, forming a progressively enhanced inference chain. Further, a small language model is used to analyze the output of the large language model from an independent perspective, perform logical semantic detection on each inference chain, accurately identify and filter out inaccurate information therein, realizing the phased cooperation between the large language model and the small language model, combining the inference depth of the large language model with the accuracy and efficiency of the small language model, overcoming the limitations of the large language model's self-reflection in detecting hallucinations, ensuring the reliability of the generated results, and finally extracting the target answer from the target inference chains that conform to the logical semantic relationship, capable of accurately and efficiently outputting the results, especially suitable for complex question-and-answer scenarios that require multi-step logical reasoning.

[0047] In some embodiments, based on step 103, the process of using a small language model to perform logical semantic relationship detection on each inference chain to obtain multiple target inference chains with correct logical semantic relationships includes: Using a small language model to perform logical semantic relationship detection on each inference chain to obtain the logical semantic relationship detection results of each inference chain; Scoring based on the logical semantic relationship detection results of each inference chain to obtain the logical semantic relationship scores of each inference chain; wherein, when the logical semantic relationship of the inference chain is incorrect, the logical semantic relationship score of the inference chain is 0, and when the logical semantic relationship of the inference chain is correct, the logical semantic relationship score of the inference chain is 1; Determining the inference chains whose product of all logical semantic relationship scores is 1 as the target inference chains with correct logical semantic relationships.

[0048] Specifically, NLI aims to judge the logical semantic relationship between two text segments. Therefore, the NLI technique can be executed by a small language model to detect the logical semantic relationship between the input question and the heuristic answer result in each inference chain, the logical semantic relationship between the heuristic answer result and the stimulated thinking result, and the logical semantic relationship between the stimulated thinking result and the self-reflection result.

[0049] Then, a score is calculated comprehensively based on the logical semantic relationships between the input question and the heuristic answer result, between the heuristic answer result and the stimulated thinking result, and between the stimulated thinking result and the self-reflection result. When one of these logical semantic relationships is incorrect, it is considered that the logical semantics of the entire reasoning chain is incorrect, and the score for the logical semantic relationship is 0. When all logical semantic relationships are correct, it is considered that the logical semantics of the entire reasoning chain is correct, and the score for the logical semantic relationship is 1.

[0050] After calculating the scores for the logical semantic relationships of each reasoning chain, a filtering mechanism is applied to retain only the reasoning chains whose product of all chain scores is equal to 1, that is, only the reasoning chains with a logical semantic relationship score of 1, as the target reasoning chains with correct logical semantics, namely: 。

[0051] Based on the above, when detecting the logical semantic relationship of each reasoning chain through a small language model to obtain the logical semantic relationship detection result of each reasoning chain, the following steps are performed for each reasoning chain: Through a small language model, using natural language inference technology, detect the logical semantic relationship between the input question and the heuristic answer result in the reasoning chain to obtain the first logical semantic relationship detection result between the input question and the heuristic answer result; Through a small language model, using natural language inference technology, detect the logical semantic relationship between the heuristic answer result and the stimulated thinking result in the reasoning chain to obtain the second logical semantic relationship detection result between the heuristic answer result and the stimulated thinking result; Through a small language model, using natural language inference technology, detect the logical semantic relationship between the stimulated thinking result and the self-reflection result in the reasoning chain to obtain the third logical semantic relationship detection result between the stimulated thinking result and the self-reflection result.

[0052] Specifically, small language models such as BERT can be applied to perform NLI tasks. When judging the logical relationship between two text fragments, NLI usually falls into three categories, namely entailment, contradiction, and neutral.

[0053] Among them, entailment means that the "premise" text supports the "hypothesis" text. For example, if the premise text is "The capacitance of the capacitor is 2 millifarads and the potential difference is 5 volts. According to the formula Q = C × V, the charge on the positive plate is 0.01 coulombs.", and the hypothesis text is "The charge on the positive plate is proportional to the capacitance and voltage.", then the calculation in the premise directly supports the conclusion of the hypothesis.

[0054] A contradiction means that there is a conflict between the "premise" text and the "hypothesis" text. For example, if the premise text is "The electric charge of a capacitor is only determined by its capacitance and has nothing to do with voltage.", and the hypothesis text is "The electric charge Q needs to be calculated using the formula Q = C×V.", then the premise negates the role of voltage, while the hypothesis clearly requires voltage for calculation, and there is a logical conflict between the two.

[0055] Neutrality means that the "premise" text neither supports nor negates the "hypothesis" text. For example, if the premise text is "The potential difference between the two plates of a capacitor is 5 volts.", and the hypothesis text is "This capacitor is made of ceramic dielectric.", then the premise does not mention the dielectric material, and the information in the hypothesis is neither supported nor negated.

[0056] When detecting the logical semantic relationship of each reasoning chain, the consistency between different stages of the reasoning chain is checked through NLI, that is, the input question Q is used as the premise (where is the hypothesis of Q), is used as the premise (where is the hypothesis), is used as the premise (where is the hypothesis). For the set premise and hypothesis, the logical semantic relationship between the two is detected, and the first logical semantic relationship detection result between the input question and the heuristic answer result, the second logical semantic relationship detection result between the heuristic answer result and the stimulating thinking result, and the third logical semantic relationship detection result between the stimulating thinking result and the self-reflection result are obtained respectively.

[0057] In practical applications, the detection efficiency can be improved by batch processing, and multiple reasoning chains can be input into the NLI model for calculation at one time.

[0058] Furthermore, when scoring based on the logical semantic relationship detection results of each reasoning chain to obtain the logical semantic relationship score of each reasoning chain, the following steps are performed for each reasoning chain: If the first logical semantic relationship detection result is an entailment relationship, and the second logical semantic relationship detection result is an entailment relationship, and the third logical semantic relationship detection result is an entailment relationship, then it is determined that the logical semantic relationship score of the reasoning chain is 1; If any of the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result is a neutral relationship or a contradiction relationship, then it is determined that the logical semantic relationship score of the reasoning chain is 0.

[0059] Specifically, if the first logical semantic relationship detection result is an entailment relationship, the second logical semantic relationship detection result is an entailment relationship, and the third logical semantic relationship detection result is an entailment relationship, then it is determined that the logical semantic relationship score of this inference chain is 1.

[0060] If any one of the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result is a neutral relationship or a contradictory relationship, then it is determined that the logical semantic relationship score of this inference chain is 0.

[0061] The defined scoring function is as follows: 。

[0062] Filter the inference chains according to the scores, remove the inference chains that do not meet the requirements, and retain the reliable inference chains for subsequent processing.

[0063] In the embodiments of the present invention, a small language model is used to execute the NLI-based hallucination detection mechanism, and the small language model is used to analyze the output of the large language model from an independent perspective, realizing the automatic detection and scoring of multi-stage logical semantic relationships in the inference chain, screening out the target inference chains with completely coherent logic, overcoming the limitations of the large language model's self-reflection and detection of hallucinations, ensuring the reliability of the generated results, and also ensuring the reliability of subsequent processing.

[0064] In some embodiments, based on step 104, extracting the target answer to the input question from the multiple target inference chains and outputting the target answer includes: Using regular expressions, extract the target information for answering the input question from each target inference chain; If the target information is successfully extracted from any target inference chain, then standardize the format of the target information to obtain the initial answer corresponding to the target inference chain; If the target information is not successfully extracted from any target inference chain, then use the small language model fine-tuned for the task to generate the target information for answering the input question based on the target inference chain, and standardize the format of the target information to obtain the initial answer corresponding to the target inference chain; Based on the consistency evaluation of each initial answer, obtain the target answer to the input question and output the target answer.

[0065] Specifically, use regular expressions and keyword matching techniques for structured extraction to extract the target information for answering the input question from the target inference chains that have passed the hallucination detection.

[0066] If the target information is successfully extracted from a certain target inference chain, the target information is formatted and standardized, and the output template is enforced, such as representing structured data following JSON fields or using a fixed option list in a classification task, to ensure that the result format is consistently available, and the initial answer corresponding to the target inference chain is obtained. For example, when processing text containing date information, a specific date format is matched through regular expressions and extracted for output according to the specified date format.

[0067] If the target information is not successfully extracted from a certain target inference chain, a small language model fine-tuned for the task is used to generate simplified target information for answering the input question based on the target inference chain, and the generated simplified target information is formatted and standardized to obtain the initial answer corresponding to the target inference chain. If the small language model fine-tuned for the task cannot generate the corresponding answer either, the target inference chain is marked for further investigation or manual intervention (in the experiment, the marked answer is regarded as an incorrect answer).

[0068] Among them, the small language model fine-tuned for the task is a small language model fine-tuned according to the current task. Before the conversation for hallucination mitigation, the small language model is fine-tuned accordingly for different tasks. For example, in a classification task, a BERT classifier is trained to adapt to specific classification labels and text features so that the small language model can accurately perform the classification task.

[0069] Based on the above content, if the target information is successfully extracted from any target inference chain, formatting and standardizing the target information to obtain the initial answer corresponding to the target inference chain includes: If the target information is successfully extracted from any target inference chain, the target information is formatted and standardized to obtain formatted and standardized information; The reliability of the formatted and standardized information is verified through a small language model fine-tuned for the task; If the verification is successful, the formatted and standardized information is determined as the initial answer corresponding to the target inference chain; If the verification fails, a small language model fine-tuned for the task is used to generate the initial answer corresponding to the target inference chain based on the target inference chain.

[0070] Specifically, if the target information is successfully extracted from a certain target inference chain, the target information is formatted and standardized to obtain formatted and standardized information. Then, the reliability of the formatted and standardized information is verified through a small language model fine-tuned for the task, and the reliability of the extraction result is re-verified to confirm whether the extracted classification label or key information is consistent with the text content and the input question.

[0071] If the verification is successful, the formatted and standardized information is determined as the initial answer corresponding to the target inference chain.

[0072] If the verification fails, a small language model fine-tuned by the task generates simplified target information for answering the input question based on the target reasoning chain, and standardizes the format of the generated simplified target information to obtain the initial answer corresponding to the target reasoning chain.

[0073] Furthermore, the consistency evaluation based on each of the initial answers to obtain the target answer for answering the input question includes: Determine the number of occurrences of the same answer among each of the initial answers; If the highest number of occurrences is greater than or equal to the preset number-of-occurrences threshold, determine the answer corresponding to the highest number of occurrences as the target answer for answering the input question.

[0074] Specifically, since the initial answers are format-standardized, there will be the same answers among multiple initial answers, and determine the number of occurrences of the same answers among each of the initial answers.

[0075] If the highest number of occurrences is greater than or equal to the preset number-of-occurrences threshold, i.e., S≥T, accept the answer corresponding to the highest number of occurrences and output it as the target answer for answering the input question Q.

[0076] If the highest number of occurrences is less than the preset number-of-occurrences threshold, i.e., S<T, reject the answer corresponding to the highest number of occurrences.

[0077] Among them, when setting the preset number-of-occurrences threshold T, appropriate values can be determined through multiple experiments according to the characteristics of different tasks and the requirements for result reliability. For example, in tasks with extremely high accuracy requirements, the value of T can be appropriately increased; in tasks with high efficiency requirements and tolerable error rates, the value of T can be decreased.

[0078] The embodiment of the present invention uses a hybrid standardization method that combines rule-based formatting and small language model-driven verification to efficiently extract or generate standardized answers from logically and semantically correct target reasoning chains, which not only ensures the basic structure and consistency of the output but also can be flexibly adjusted according to specific application requirements, enhancing the usability and consistency of the framework in diverse tasks.

[0079] The multi-stage hallucination mitigation synergy (MS-HM Synergy) method proposed by the present invention realizes the phased collaboration between LLMs and small models, combines the reasoning depth and creativity of LLMs with the accuracy and efficiency of small models, and in physical, chemical, mathematical, and logical benchmark tests, the accuracy rate is significantly improved compared with traditional methods such as the standard prompting method and chain of thought (CoT). The following shows the quantitative result comparison between the MS-HM synergy method and the baseline method in language reasoning tasks.

[0080] Table 1 shows the quantitative results of the accuracy of the Qwen-14B Chat large language model in language inference tasks. This method has a significant improvement compared with the previous state-of-the-art methods, with absolute advantages of 0.63%, 8.08%, 6.89%, 12.59% and 2.29% in logical benchmark tests, physics, chemistry, mathematics and classroom dialogue coding classification respectively. The Tree of Thought (ToT) and Socratic questioning methods require pre-constructing specific data structures, and the nature of these structures will affect their results. Therefore, when evaluating with the Qwen-14B Chat model, these methods are not directly compared with the MS-HM collaborative method.

[0081] Table 1

[0082] Table 2 shows the quantitative results of the accuracy of the GPT-3.5-turbo large language model in language inference tasks. This method outperforms the previous advanced methods in categories such as physics, chemistry, mathematics and classroom dialogue coding, with improvements of 1.28%, 0.48%, 41.66% and 2.86% respectively. This effectively highlights the advantages of this method. It should be noted that the classroom dialogue dataset is in Chinese, while the Socratic questioning method is mainly for English datasets. Therefore, applying the Socratic questioning method to different languages poses significant challenges and may not fully meet the original intention of the method. Therefore, the experimental results of the Socratic questioning method on the classroom dialogue dataset are not provided in Table 2.

[0083] Table 2

[0084] As can be seen from Table 1 and Table 2, when using the Qwen-14B Chat large language model, there are absolute advantages of 0.63%, 8.08%, 6.89%, 12.59% and 2.29% respectively compared with the traditional methods in logical benchmark tests, physics, chemistry, mathematics and classroom dialogue coding classification; when using the GPT-3.5-turbo large language model, there are also obvious improvements in categories such as physics, chemistry, mathematics and classroom dialogue coding, which are 1.28%, 0.48%, 41.66% and 2.86% respectively.

[0085] In addition, a detailed analysis of the performance of the MS-HM collaborative method in different rounds of conversations is also carried out, which can be combined with Figure 3 for analysis. Figure 3 is a schematic diagram of the ratio of the accuracy of different rounds of conversations of the MS-HM collaborative method provided by the present invention on Qwen-14B Chat and GPT-3.5-turbo. Figure 3(a)shows the accuracy of Qwen-14B Chat in different rounds of conversations based on the MS-HM collaborative method. Figure 3 (b)shows the accuracy of GPT-3.5-turbo in different rounds of conversations based on the MS-HM collaborative method. By comparing the accuracy of two-round and three-round conversations, it is found that increasing the number of conversation rounds generally improves the performance of the model. For example, in the Qwen-14B Chat model, when the number of conversation rounds increases from two to three, the average accuracy increases from 53.84% to 55.40%. Similarly, in the GPT-3.5-turbo model, the average accuracy increases from 50.20% to 53.30%. These results indicate that multi-round conversations help the model better understand context information, thereby improving the accuracy of its answers. However, for some complex tasks, such as classroom conversations, increasing the number of conversation rounds does not bring a significant performance improvement, and even leads to a slight decline in some cases. This may be attributed to the complexity of the conversation content of these tasks, which poses higher requirements for the model's context understanding and reasoning abilities. Therefore, the experimental results show that although increasing the number of conversation rounds generally improves the model performance, for specific tasks, it is still necessary to further optimize the conversation strategy and model design to better adapt to complex conversation scenarios.

[0086] Next, the qualitative results comparison between the MS-HM collaborative method and the baseline method in physical tasks is presented. First, as Figure 4 shown, Figure 4 is a schematic diagram of the qualitative results of the traditional reasoning method provided by the present invention in physical tasks. The traditional reasoning methods include Standard-Prompting, Chain of Thought (CoT), Tree of Thought (ToT), and Graph of Thought (GoT) methods. Figure 4 shows a schematic diagram of the qualitative results of the Standard-Prompting, Chain of Thought, Tree of Thought, and Graph of Thought methods in physical tasks. The correct answer to this example is B.

[0087] It can be seen that the MS-HM collaborative method can effectively generate prompts containing the information required to solve the original problem. In the guided reasoning stage, large language models (LLMs) use their language understanding ability to decompose the problem and generate relevant prompts. Subsequently, in the hallucination detection stage, the unreliable information in these prompts is identified and eliminated, ensuring the accuracy of the information used for reasoning. Finally, in the result standardization stage, the output is structured to make it coherent and easy to understand. By selectively using these refined prompts, the MS-HM collaborative method provides a reasonable explanation through reflective reasoning and arrives at the correct final answer. In contrast, the Standard-Prompting, Chain of Thought (CoT), Tree of Thought (ToT), and Graph of Thought (GoT) methods have poor reasoning paths, such as Figure 4As shown, the red part is the wrong answer inferred by the traditional method, resulting in model hallucinations. The Tree of Thoughts (ToT) method decomposes problem-solving into linear, structured steps. It attempts to ensure thorough reasoning but appears too rigid and inefficient for simple tasks like capacitor problems. Its repetitive, step-by-step nature lacks the flexibility required to handle simple problems. In the overall goal of this invention - leveraging the comprehensive advantages of the model, ToT cannot well adapt to the complexity of different problems and does not scale well for problems where a direct solution is faster. Additionally, it has no effective mechanism to detect and correct potential hallucinations in the reasoning process. The Graph of Thoughts (GoT) method organizes problem-solving in a non-linear manner and can flexibly explore interconnected steps. Although it may be effective for complex problems, it overcomplicates simple tasks by introducing unnecessary nodes, backtracking, and cognitive load. This makes it inefficient both visually and procedurally for direct formula-based solutions. Similar to ToT, GoT also lacks a method to comprehensively handle hallucinations and standardize outputs in a way that meets the requirements of downstream tasks. The MS-HM collaborative method has significant advantages in multi-step logical reasoning and overcomes the key limitations of traditional chain-of-thought (CoT) techniques. Different from sequential or predefined structured methods that are prone to error propagation and challenges in decomposing complex problems, the MS-HM collaborative method combines the complex reasoning ability of LLMs and the verification accuracy of small language models. Through a three-stage process of guided reasoning, hallucination detection, and result standardization, it achieves more accurate and flexible problem-solving. As shown in Table 3, the correct answer for this example is B.

[0088] The key limitations of thinking (CoT) techniques. Different from sequential or predefined structured methods that are prone to error propagation and challenges in decomposing complex problems, the MS-HM collaborative method combines the complex reasoning ability of LLMs and the verification accuracy of small language models. Through a three-stage process of guided reasoning, hallucination detection, and result standardization, it achieves more accurate and flexible problem-solving. As shown in Table 3, the correct answer for this example is B.

[0089] Table 3

[0090] Next, the structure of the intelligent dialogue system based on large language model hallucination mitigation provided by the present invention will be described. The intelligent dialogue system based on large language model hallucination mitigation described below can be mutually corresponding and referred to the intelligent dialogue method based on large language model hallucination mitigation described above.

[0091] Referring to Figure 5 , Figure 5 is a schematic structural diagram of the intelligent dialogue system based on large language model hallucination mitigation provided by the present invention.

[0092] The intelligent dialogue system based on large language model hallucination mitigation includes: The model inference module 510 is used to output multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results through a large language model based on the input question; among them, the i-th round of heuristic answer result is generated based on the input question, the i-th round of stimulating thinking result is generated based on the input question and the i-th round of heuristic answer result, and the i-th round of self-reflection result is generated based on the input question, the i-th round of heuristic answer result, and the i-th round of stimulating thinking result; The inference chain generation module 520 is used to generate corresponding inference chains based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round; The logical semantic relationship detection module 530 is used to detect the logical semantic relationship of each inference chain through a small language model to obtain multiple target inference chains with correct logical semantic relationships; The target answer output module 540 is used to extract the target answer to the input question from the multiple target inference chains and output the target answer.

[0093] The intelligent dialogue system based on hallucination mitigation of large language models provided by the present invention uses a large language model to first generate heuristic answer results to expand the reply perspective, then stimulate in-depth reasoning, and finally conduct self-reflection to correct potential problems, forming a progressively enhanced inference chain. Further, a small language model is used to analyze the output of the large language model from an independent perspective, detect the logical semantics of each inference chain, and can accurately identify and filter out inaccurate information therein, realizing the phased cooperation of the large language model and the small language model, combining the reasoning depth of the large language model with the accuracy and efficiency of the small language model, overcoming the limitations of the large language model's self-reflection and detection of hallucinations, ensuring the reliability of the generated results, and finally extracting the target answer from the target inference chains that conform to the logical semantic relationship, and can accurately and efficiently output the results, especially suitable for complex question-and-answer scenarios that require multi-step logical derivation.

[0094] Furthermore, the logical semantic relationship detection module 530 is further used for: Detect the logical semantic relationship of each inference chain through a small language model to obtain the logical semantic relationship detection result of each inference chain; Score based on the logical semantic relationship detection results of each inference chain to obtain the logical semantic relationship score of each inference chain; among them, when the logical semantic relationship of the inference chain is incorrect, the logical semantic relationship score of the inference chain is 0, and when the logical semantic relationship of the inference chain is correct, the logical semantic relationship score of the inference chain is 1; Determine the inference chains with the product of all logical semantic relationship scores being 1 as the target inference chains with correct logical semantic relationships.

[0095] Further, the logical semantic relationship detection module 530 is further configured to: Detect the logical semantic relationship between the input question and the heuristic answer result in the inference chain through a small language model using natural language inference technology, to obtain a first logical semantic relationship detection result between the input question and the heuristic answer result; Detect the logical semantic relationship between the heuristic answer result and the stimulated thinking result in the inference chain through a small language model using natural language inference technology, to obtain a second logical semantic relationship detection result between the heuristic answer result and the stimulated thinking result; Detect the logical semantic relationship between the stimulated thinking result and the self-reflection result in the inference chain through a small language model using natural language inference technology, to obtain a third logical semantic relationship detection result between the stimulated thinking result and the self-reflection result.

[0096] Further, the logical semantic relationship detection module 530 is further configured to: If the first logical semantic relationship detection result is an entailment relationship, and the second logical semantic relationship detection result is an entailment relationship, and the third logical semantic relationship detection result is an entailment relationship, then determine that the logical semantic relationship score of the inference chain is 1; If any one of the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result is a neutral relationship or a contradiction relationship, then determine that the logical semantic relationship score of the inference chain is 0.

[0097] Further, the target answer output module 540 is further configured to: Use a regular expression to extract target information for answering the input question from each target inference chain; If target information is successfully extracted from any target inference chain, then standardize the format of the target information to obtain an initial answer corresponding to the target inference chain; If target information is not successfully extracted from any target inference chain, then generate target information for answering the input question based on the target inference chain through a task-fine-tuned small language model, and standardize the format of the target information to obtain an initial answer corresponding to the target inference chain; Perform a consistency evaluation based on each of the initial answers to obtain a target answer for answering the input question, and output the target answer.

[0098] Further, the target answer output module 540 is further configured to: If target information is successfully extracted from any target inference chain, then standardize the format of the target information to obtain format-standardized information; A small language model fine-tuned by tasks verifies the reliability of the format-standardized information; If the verification is successful, the format-standardized information is determined as the initial answer corresponding to the target reasoning chain; If the verification fails, a small language model fine-tuned by tasks generates the initial answer corresponding to the target reasoning chain based on the target reasoning chain.

[0099] Furthermore, the target answer output module 540 is further configured to: Determine the occurrence times of the same answers in each of the initial answers; If the highest occurrence times is greater than or equal to a preset occurrence times threshold, determine the answer corresponding to the highest occurrence times as the target answer to answer the input question.

[0100] It should be noted that the intelligent dialogue system based on large language model hallucination mitigation provided by the present invention can execute the intelligent dialogue method based on large language model hallucination mitigation described in any of the above embodiments during specific operation, and this embodiment will not be elaborated herein.

[0101] Figure 6 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 6 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the intelligent dialogue method based on large language model hallucination mitigation. The method includes: outputting multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results based on the input question through a large language model; wherein, the i-th round of heuristic answer result is generated based on the input question, the i-th round of stimulating thinking result is generated based on the input question and the i-th round of heuristic answer result, and the i-th round of self-reflection result is generated based on the input question, the i-th round of heuristic answer result, and the i-th round of stimulating thinking result; i is a positive integer; generating a corresponding reasoning chain based on the heuristic answer result, stimulating thinking result, and self-reflection result of the same round; detecting the logical semantic relationship of each reasoning chain through a small language model to obtain multiple target reasoning chains with correct logical semantic relationships; extracting the target answer to answer the input question from the multiple target reasoning chains and outputting the target answer.

[0102] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0103] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the intelligent dialogue method for alleviating large language model hallucinations provided in the above-mentioned various embodiments. The method includes: based on an input question, outputting multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results through a large language model; wherein, the i-th round of heuristic answer result is generated based on the input question, the i-th round of stimulating thinking result is generated based on the input question and the i-th round of heuristic answer result, and the i-th round of self-reflection result is generated based on the input question, the i-th round of heuristic answer result, and the i-th round of stimulating thinking result; i is a positive integer; generating corresponding inference chains based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round; detecting the logical semantic relationships of each inference chain through a small language model to obtain multiple target inference chains with correct logical semantic relationships; extracting a target answer for answering the input question from the multiple target inference chains and outputting the target answer.

[0104] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the intelligent dialogue method for alleviating hallucinations based on large language models provided in the above various embodiments. The method includes: through a large language model, based on an input question, outputting multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results; wherein, the i-th round of heuristic answer result is generated based on the input question, the i-th round of stimulating thinking result is generated based on the input question and the i-th round of heuristic answer result, and the i-th round of self-reflection result is generated based on the input question, the i-th round of heuristic answer result, and the i-th round of stimulating thinking result; i is a positive integer; based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round, generating corresponding inference chains; through a small language model, detecting the logical semantic relationships of each inference chain to obtain multiple target inference chains with correct logical semantic relationships; extracting a target answer for answering the input question from the multiple target inference chains and outputting the target answer.

[0105] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An intelligent dialogue method based on alleviating hallucinations of large language models, characterized in that, The intelligent dialogue method based on alleviating large language model hallucinations includes: Through a large language model, based on the input question, output multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results; among them, the i-th round heuristic answer result is generated based on the input question, the i-th round stimulating thinking result is generated based on the input question and the i-th round heuristic answer result, and the i-th round self-reflection result is generated based on the input question, the i-th round heuristic answer result, and the i-th round stimulating thinking result; i is a positive integer; Based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round, generate corresponding reasoning chains; Through a small language model, perform logical semantic relationship detection on each reasoning chain to obtain multiple target reasoning chains with correct logical semantic relationships; Extract the target answer to the input question from the multiple target reasoning chains and output the target answer.

2. The intelligent dialogue method based on large language model hallucination mitigation according to claim 1, wherein The step of performing logical semantic relationship detection on each reasoning chain through a small language model to obtain multiple target reasoning chains with correct logical semantic relationships includes: Through a small language model, perform logical semantic relationship detection on each reasoning chain to obtain the logical semantic relationship detection result of each reasoning chain; Score based on the logical semantic relationship detection result of each reasoning chain to obtain the logical semantic relationship score of each reasoning chain; among them, when the logical semantic relationship of the reasoning chain is incorrect, the logical semantic relationship score of the reasoning chain is 0, and when the logical semantic relationship of the reasoning chain is correct, the logical semantic relationship score of the reasoning chain is 1; Determine the reasoning chains with the product of all logical semantic relationship scores being 1 as the target reasoning chains with correct logical semantic relationships.

3. The intelligent dialogue method for alleviating large language model hallucinations according to claim 2, wherein The logical semantic relationship detection result includes the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result; When performing logical semantic relationship detection on each reasoning chain through a small language model to obtain the logical semantic relationship detection result of each reasoning chain, the following steps are executed for each reasoning chain: Through a small language model, using natural language inference technology, perform logical semantic relationship detection on the input question and the heuristic answer result in the reasoning chain to obtain the first logical semantic relationship detection result between the input question and the heuristic answer result; Through a small language model, using natural language inference technology, perform logical semantic relationship detection on the heuristic answer result and the stimulating thinking result in the reasoning chain to obtain the second logical semantic relationship detection result between the heuristic answer result and the stimulating thinking result; Through a small language model, using natural language inference technology, perform logical semantic relationship detection on the stimulating thinking result and the self-reflection result in the reasoning chain to obtain the third logical semantic relationship detection result between the stimulating thinking result and the self-reflection result.

4. The intelligent dialogue method based on large language model hallucination mitigation according to claim 3, characterized in that, When scoring based on the logical semantic relationship detection result of each reasoning chain to obtain the logical semantic relationship score of each reasoning chain, the following steps are executed for each reasoning chain: If the first logical semantic relationship detection result is an implication relationship, and the second logical semantic relationship detection result is an implication relationship, and the third logical semantic relationship detection result is an implication relationship, then determine that the logical semantic relationship score of the inference chain is 1; If any one of the first logical semantic relationship detection result, the second logical semantic relationship detection result, and the third logical semantic relationship detection result is a neutral relationship or a contradiction relationship, then determine that the logical semantic relationship score of the inference chain is 0.

5. The intelligent dialogue method based on the alleviation of large language model hallucinations according to claim 1, wherein, The extracting a target answer to the input question from the multiple target inference chains and outputting the target answer includes: Using a regular expression, extracting target information for answering the input question from each target inference chain; If target information is successfully extracted from any target inference chain, then standardize the format of the target information to obtain an initial answer corresponding to the target inference chain; If target information is not successfully extracted from any target inference chain, then generate target information for answering the input question based on the target inference chain through a small language model fine-tuned for the task, and standardize the format of the target information to obtain an initial answer corresponding to the target inference chain; Based on a consistency evaluation of each of the initial answers, obtain a target answer for answering the input question, and output the target answer.

6. The intelligent dialogue method based on large language model hallucination mitigation according to claim 5, wherein, The if target information is successfully extracted from any target inference chain, then standardize the format of the target information to obtain an initial answer corresponding to the target inference chain includes: If target information is successfully extracted from any target inference chain, then standardize the format of the target information to obtain format-standardized information; Verify the reliability of the format-standardized information through a small language model fine-tuned for the task; If the verification is successful, then determine the format-standardized information as the initial answer corresponding to the target inference chain; If the verification fails, then generate an initial answer corresponding to the target inference chain based on the target inference chain through a small language model fine-tuned for the task.

7. The intelligent dialogue method based on large language model hallucination mitigation according to claim 5, characterized in that, The based on a consistency evaluation of each of the initial answers, obtain a target answer for answering the input question includes: Determine the number of occurrences of the same answer among each of the initial answers; If the highest number of occurrences is greater than or equal to a preset occurrence threshold, then determine the answer corresponding to the highest number of occurrences as the target answer for answering the input question.

8. An intelligent dialogue system based on the alleviation of large language model hallucinations, characterized in that, Includes: A model inference module, configured to output multi-round heuristic answer results, multi-round stimulating thinking results, and multi-round self-reflection results based on an input question through a large language model; wherein, the i-th round heuristic answer result is generated based on the input question, the i-th round stimulating thinking result is generated based on the input question and the i-th round heuristic answer result, and the i-th round self-reflection result is generated based on the input question, the i-th round heuristic answer result, and the i-th round stimulating thinking result; An inference chain generation module, configured to generate corresponding inference chains based on the heuristic answer results, stimulating thinking results, and self-reflection results of the same round; A logical semantic relationship detection module, which is used to detect the logical semantic relationship of each inference chain through a small language model to obtain multiple target inference chains with correct logical semantic relationships; A target answer output module, which is used to extract a target answer to the input question from the multiple target inference chains and output the target answer.

9. An electronic device, the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the intelligent dialogue method based on large language model hallucination mitigation according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the intelligent dialogue method based on large language model hallucination mitigation according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Large model illusion problem relieving method and device, equipment and storage medium

    CN119938857A

  • Large language model illusion detection method and system based on knowledge graph structure

    CN120104765A

  • Faithful generation of output text for multimodal applications

    WO2025054081A1

Cited By

  • Dialogue chain multidimensional semantic enhancement method based on MCP agent negotiation and voting mechanism

    CN120851040A

  • Large language model illusion detection method based on causal double chains and intelligent question and answer method

    CN120892928A

  • Causal double chain-based large language model illusion detection method and intelligent question answering method

    CN120892928B

  • Coal mining machine fault diagnosis intelligent agent, construction and use method, medium and equipment

    CN121052375A

  • Language model optimization method and device, equipment and medium

    CN121188155A