Black box large language model illusion detection and correction method and device

By designing the metamorphic relationship to generate subsequent problems and detect consistency, and combining multi-path voting to correct hallucination answers, the problem of illusion generation by black box large language model is solved, and effective hallucination detection and correction without relying on confidence is achieved.

CN120258147APending Publication Date: 2025-07-04SUN YAT SEN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510417005.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing large-model hallucination detection and correction methods cannot effectively deal with the hallucination content generated by the black box model, especially the factual hallucinations, and the dependence on confidence detection is unreliable.

Method used

By designing the metamorphosis relationship, we generate subsequent questions and detect the consistency between the source answers and subsequent answers, we use multi-path voting to correct illusory answers, avoid relying on model confidence, and use a black box large language model to generate source answers and subsequent answers, and correct them in combination with preset errors to correct the metamorphosis relationship.

Benefits of technology

In the case of unreliable or unavailable confidence, effectively identifying and correcting the hallucinations of the content generated by the model improves the accuracy of hallucinations detection and correction, and avoids dependence on the internal state of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258147A_ABST
    Figure CN120258147A_ABST
Patent Text Reader

Abstract

The invention discloses a black box large language model illusion detection and correction method and device. The method and device are used for solving the technical problem that an existing large model illusion detection and correction method cannot effectively cope with model generation illusion content. The method comprises the following steps: acquiring a source question, and based on the source question and a preset question answer metamorphic relationship, generating a source answer and a subsequent answer by adopting a black box large language model; performing illusion detection on the source answer according to the subsequent answer to generate an illusion detection result; and based on the source question and a preset error correction metamorphic relation, correcting the source answer according to the illusion detection result by adopting a preset correction technology and a black box large language model, and outputting a corrected answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for detecting and correcting hallucinations in black-box large language models. Background Art

[0002] With the rapid development of artificial intelligence, large model technology has become a new paradigm in machine learning and has been widely applied in various fields. However, the hallucination problem of large models hinders the application of large model technology in high-reliability scenarios.

[0003] Hallucination refers to the content generated by a large model that is untrue or does not conform to the input. In the past, smaller-scale generation models may also produce hallucinations. The hallucinations of small models usually manifest as generating meaningless content or nonsense, and such hallucinations can be manually distinguished by users. However, the hallucinations of large models usually generate content that conforms to grammar logic but may be inconsistent with facts or the input in details, and such hallucinations are difficult to distinguish. Applying a large model with hallucination problems to key task fields such as medicine, finance, and law may have serious consequences.

[0004] Existing methods for detecting and correcting hallucinations in large models usually use the confidence of the model in the generated content to detect and correct hallucinations. However, this method requires accessing the intermediate layer output of the model, and mainstream commercial large models are black boxes, and users cannot obtain the intermediate layer output of these black box models. In addition, the confidence itself cannot fully reflect the accuracy of the generated content, because the model may show high confidence in incorrect content and low confidence in correct content. Therefore, simply relying on confidence for hallucination detection and correction cannot effectively solve the technical problem of the model generating hallucinated content. Summary of the Invention

[0005] The present invention provides a method and device for detecting and correcting hallucinations in black-box large language models, which are used to solve the technical problem that existing methods for detecting and correcting hallucinations in large models cannot effectively handle the problem of the model generating hallucinated content.

[0006] A method for detecting and correcting hallucinations in a black-box large language model provided by the first aspect of the present invention includes:

[0007] Obtain a source question, and based on the source question and a preset question-answer metamorphosis relationship, use a black-box large language model to generate a source answer and a subsequent answer;

[0008] Perform hallucination detection on the source answer according to the subsequent answer to generate a hallucination detection result;

[0009] Based on the source question and a preset error correction metamorphosis relationship, use a preset correction technique and the black-box large language model to correct the source answer according to the hallucination detection result, and output a corrected answer.

[0010] Optionally, based on the source question and the preset question-answer metamorphosis relationship, using a black-box large language model to generate a source answer and subsequent answers, including:

[0011] Input the source question into the black-box large language model to generate a source answer;

[0012] Use the preset question-answer metamorphosis relationship to metamorphose the source question and the source answer to generate subsequent questions;

[0013] Input the subsequent questions into the black-box large language model to generate subsequent answers.

[0014] Optionally, detecting hallucinations in the source answer based on the subsequent answer to generate a hallucination detection result, including:

[0015] Preprocess the source answer and the subsequent answer respectively to output a target source answer and a target subsequent answer;

[0016] Judge whether the target source answer and the target subsequent answer are consistent;

[0017] If they are consistent, determine the hallucination detection result as no hallucination;

[0018] If they are inconsistent, determine the hallucination detection result as having hallucinations.

[0019] Optionally, the preset correction technique includes a preset large model and the cosine similarity method; based on the source question and the preset error correction metamorphosis relationship, using the preset correction technique and the black-box large language model to correct the source answer according to the hallucination detection result and output a corrected answer, including:

[0020] When the hallucination detection result is determined to be no hallucination, use the source answer as the corrected answer;

[0021] When the hallucination detection result is determined to be having hallucinations, use the preset error correction metamorphosis relationship to metamorphose the source question and the source answer to generate multiple error correction subsequent questions;

[0022] Take each of the error correction subsequent questions as the input of the black-box large language model and output the error correction subsequent answers corresponding to each of the error correction subsequent questions;

[0023] Use the preset large model to conduct a majority vote on each of the error correction subsequent answers and output a corrected answer;

[0024] Or use the cosine similarity method to conduct a majority vote on each of the error correction subsequent answers and output a corrected answer.

[0025] Optionally, using the cosine similarity method to perform majority voting on each of the corrected subsequent answers and output a corrected answer, including:

[0026] Perform feature embedding on each of the corrected subsequent answers to generate sentence embedding vectors corresponding to each of the corrected subsequent answers;

[0027] Calculate the cosine similarity of each of the sentence embedding vectors to determine the similarity between each pair of the corrected subsequent answers;

[0028] Using the similarity between each pair of the corrected subsequent answers as the matrix element value, construct a similarity matrix;

[0029] Compare each matrix element value in the similarity matrix with a preset similarity threshold respectively;

[0030] Take any matrix element value greater than the preset similarity threshold as the first matrix element value;

[0031] Take any matrix element value less than or equal to the preset similarity threshold as the second matrix element value;

[0032] Count the number of the first matrix element values corresponding to each matrix row in the similarity matrix to determine the number of the first matrix element values corresponding to each matrix row in the similarity matrix;

[0033] Take the corrected subsequent answer corresponding to the largest number of the first matrix element values as the corrected answer.

[0034] Optionally, the preset question-answer transformation relationships include a thought chain prompt transformation relationship, a multilingual translation transformation relationship, a question optimization transformation relationship, an external knowledge introduction transformation relationship, a skeptical question construction transformation relationship, and a multiple-choice question construction transformation relationship.

[0035] A black-box large language model hallucination detection and correction device provided in the second aspect of the present invention includes:

[0036] An acquisition module, configured to acquire a source question, and based on the source question and the preset question-answer transformation relationship, use a black-box large language model to generate a source answer and subsequent answers;

[0037] A detection module, configured to perform hallucination detection on the source answer according to the subsequent answers to generate a hallucination detection result;

[0038] A correction module, configured to correct the source answer according to the hallucination detection result based on the source question and the preset error correction transformation relationship, using a preset correction technique and the black-box large language model, and output a corrected answer.

[0039] A computer device provided by the third aspect of the present invention includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the black-box large language model hallucination detection and correction method as described in any one of the above.

[0040] A computer-readable storage medium provided by the fourth aspect of the present invention stores a computer program / instructions thereon. When the computer program / instructions are executed, the steps of the black-box large language model hallucination detection and correction method as described in any one of the above are implemented.

[0041] A computer program product provided by the fifth aspect of the present invention includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes programs / instructions. When the programs / instructions are executed by a computer, the computer executes the steps of the black-box large language model hallucination detection and correction method as described in any one of the above.

[0042] From the above technical solutions, it can be seen that the present invention has the following advantages:

[0043] The above technical solution of the present invention provides a black-box large language model hallucination detection and correction method. The source problem is obtained, and based on the source problem and the preset problem-answer metamorphosis relationship, the source answer and subsequent answers are generated using a black-box large language model; the source answer is subjected to hallucination detection according to the subsequent answers to generate a hallucination detection result; based on the source problem and the preset error correction metamorphosis relationship, the source answer is corrected using a preset correction technique and a black-box large language model according to the hallucination detection result, and a corrected answer is output; based on the above solution, by combining the preset problem-answer metamorphosis relationship and the preset error correction metamorphosis relationship, the process of detecting and correcting the source answer generated by the black-box large language model according to the source problem and outputting a corrected answer, the present invention does not need to calculate the confidence of the content generated by the model, avoiding the dependence on confidence, and can effectively identify and correct the hallucination problem of the content generated by the model in the case where the confidence is unreliable / unobtainable. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a flowchart of the steps of a black-box large language model hallucination detection and correction method provided by Embodiment 1 of the present invention;

[0046] Figure 2 This is the overall framework diagram of a black-box large language model hallucination detection and correction method provided in Embodiment 1 of the present invention;

[0047] Figure 3 This is the flowchart of metamorphic testing hallucination detection provided in Embodiment 1 of the present invention;

[0048] Figure 4 This is the structural block diagram of a black-box large language model hallucination detection and correction device provided in Embodiment 2 of the present invention. Detailed implementation manners

[0049] Embodiments of the present invention provide a black-box large language model hallucination detection and correction method and device, which are used to solve the technical problem that existing large model hallucination detection and correction methods cannot effectively deal with the hallucination content generated by the model.

[0050] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] Term explanation:

[0052] 1. Metamorphic Testing (MT)

[0053] Metamorphic Testing is a software testing technique aimed at verifying the output of a program or model by constructing a Metamorphic Relation (MR). When there is no clear test benchmark, Metamorphic Testing detects potential errors in a program or model through the metamorphic relationship between input and output (such as output consistency under small input perturbations).

[0054] 2. Metamorphic Relation (MR)

[0055] The Metamorphic Relation is the core concept in Metamorphic Testing, referring to a predictable change rule that should be satisfied between input and output. By defining the Metamorphic Relation, the correctness of a system or model can be verified without a clear test benchmark (such as a standard answer). The Metamorphic Relation can be used to generate new test inputs, thereby revealing potential errors in the system.

[0056] 3. Factual / Faithfulness Hallucination

[0057] A factual hallucination refers to the phenomenon where the answers generated by large language models do not conform to objective facts or real-world knowledge. A fidelity hallucination refers to the situation where the output of a large language model fails to be faithful to the input prompt or background information, that is, the generated content deviates from the context or logical basis provided in the input information.

[0058] 4. Black-box / White-box Model

[0059] A black-box model refers to a model whose internal structure, parameter weights, or training details cannot be known to the user. The user can only interact with the model by providing input and observing the final output.

[0060] A white-box model refers to a model whose internal structure, parameter weights, and training details can be fully known to the user. The user can directly view, modify, and debug the implementation of the model, and even retrain the model.

[0061] Please refer to Figure 1 , Figure 1 which is the step flowchart of a method for detecting and correcting hallucinations in a black-box large language model provided in Embodiment 1 of the present invention.

[0062] A method for detecting and correcting hallucinations in a black-box large language model provided by the present invention includes:

[0063] Step 101, obtain the source question, and based on the source question and the metamorphic relationship between the source question and the subsequent question answer, use a black-box large language model to generate the source answer and the subsequent answer.

[0064] It should be noted that, please refer to Figure 2 , the detection and correction of the content generated by the model in the present invention are divided into two parts: metamorphic test hallucination detection and multi-path voting hallucination correction. Due to the instability of hallucinatory answers, when a large language model has hallucinations, by designing a reasonable metamorphic relationship (MR), the model is made to re-execute the same task with different execution paths, and the answers generated are more likely to be inconsistent.

[0065] Please refer to Figure 3For metamorphic testing of hallucination detection, first design a metamorphic relationship (the preset question-answering metamorphic relationship), transform the source question into subsequent questions with the same semantics, and verify whether the source answer and the subsequent answer generated by the model are consistent to detect factual hallucinations in large language models. Considering the mechanism behind the question-answering ability of large language models, design an effective metamorphic relationship to make the large language model re-execute the same task with different execution paths, improving the accuracy of hallucination detection; for multi-path voting hallucination correction, when the source answer is detected as a hallucination answer, generate error-correction subsequent answers under different execution paths by designing a metamorphic relationship (the preset error-correction metamorphic relationship), and use the voting mechanism to integrate these error-correction subsequent answers, and select the error-correction subsequent answer with the highest frequency as the final answer, thereby correcting the hallucination answer; when the source answer is detected as not a hallucination answer, the source answer is the final corrected answer.

[0066] Furthermore, the present invention mathematically describes the process of detecting and correcting the content generated by the model. Given a source question Q, the source answer generated by the black-box large language model M is A = M(Q). Define the metamorphic relationship MR as the mapping rule from (Q, A) to the subsequent question Q': Q' = MR(Q, A).

[0067] The black-box large language model M generates a subsequent answer A' = M(Q') for Q'. If MR is reasonably designed, in the case of no hallucination, A and A' should satisfy the consistency relationship R(A, A'), where R is a consistency determination function based on semantics or logic. If R(A, A') does not hold, then A is determined to be a hallucination. Further, generate an error-correction subsequent question set {Q1', Q2',..., Qn'} through the preset error-correction metamorphic relationship ECMR, that is, {Q1', Q2',..., Qn'} = ECMR(Q, A), and the black-box large language model M generates the corresponding error-correction subsequent answer set {A1', A2',..., An'} for the error-correction subsequent question set {Q1', Q2',..., Qn'}, and finally correct the hallucination based on the preset correction technique.

[0068] Specifically, step 101 includes the following sub-steps S11 - S13:

[0069] Step S11: Input the source question into the black-box large language model to generate the source answer;

[0070] Step S12: Use the preset question-answering metamorphic relationship to perform metamorphosis on the source question and the source answer to generate subsequent questions;

[0071] Step S13: Input the subsequent question into the black-box large language model to generate the subsequent answer.

[0072] The preset question-answer metamorphic relations include the thought chain prompting metamorphic relation, the multilingual translation metamorphic relation, the question optimization metamorphic relation, the introduction of external knowledge metamorphic relation, the construction of skeptical question metamorphic relation, and the construction of multiple-choice question metamorphic relation.

[0073] It should be noted that the present invention considers the mechanisms behind the question-answering ability of large language models, including question understanding, knowledge recall, and knowledge reasoning abilities, etc. The question understanding ability refers to the ability of the model to understand the context and intention of a given question. The knowledge recall ability refers to the ability of the model to retrieve the knowledge required to solve the problem from memory. The knowledge reasoning ability refers to the ability of the model to use the recalled or provided knowledge to derive new information to solve the problem. Based on the core mechanisms of the model's question-answering ability, the present invention designs reasonable metamorphic relations (preset question-answer metamorphic relations, which include question metamorphic relations, answer metamorphic relations, and composite metamorphic relations used in the hallucination detection process). Specifically:

[0074] For the question metamorphic relation (Questioning Metamorphic Relation, QMR), it changes the model's question understanding and knowledge recall paths through metamorphic questions. It includes:

[0075] a) Thought chain prompting metamorphic relation (QMR1): Add thought chain prompting words (such as: "Please think step by step to solve this problem and give the answer") to the source question to construct subsequent questions.

[0076] b) Multilingual translation metamorphic relation (QMR2): Translate the source question into different languages (such as Spanish, German, Dutch), and then vote based on the model answers to select the answer that appears the most times as the subsequent answer.

[0077] c) Question optimization metamorphic relation (QMR3): Relying on the language processing ability of the large language model, by adding prompting words (such as: "I hope you can act as a copywriter and text polisher. I will send you the text, and you help me improve the version. I hope you can modify the words and sentence structures as much as possible. Keep the meaning of the text unchanged, but make it easier to understand. You just need to polish the content without explaining the questions and requirements raised in the content."), to prompt the large model to rewrite the source question, and optimize the question with a clearer expression method to ensure that the subsequent question is semantically consistent with the source question but more clearly expressed.

[0078] d) Introduction of external knowledge metamorphic relation (QMR4): Add external knowledge (such as Wikipedia content related to the question, etc.), and let the model retrieve information from it to answer the question.

[0079] For Answering Metamorphic Relation (AMR), subsequent questions are constructed by challenging or confusing the source answer, which changes the knowledge reasoning path of the model. It includes:

[0080] a) Constructing the challenging question metamorphic relation (AMR1): The subsequent question is formed as follows: "[Source question] + [Source answer], is this statement false?" Since the answers given by large language models tend to cater to users, choosing a negative questioning form is more likely to make the model rethink the original answer.

[0081] b) Constructing the multiple-choice question metamorphic relation (AMR2): First, obtain multiple distractors based on the source answer, such as prompting the large language model to "give three answers similar to [Source answer], requiring only seemingly similar but referring to different things", and then construct a multiple-choice question as the subsequent question ("[Source question], please select from the following options: [Option 1], [Option 2],...), and use the obtained distractors, the source answer, and "none of the above answers are correct" as options.

[0082] Furthermore, the present invention can also construct a Composite Metamorphic Relation (CMR) for the hallucination detection process by combining multiple basic MRs (question metamorphic relation and answering metamorphic relation), that is, integrating different metamorphic relations to construct corresponding metamorphic relations to guide the model to adopt more different execution paths, thereby improving the hallucination detection accuracy. It includes:

[0083] a) Combining the chain of thought prompt and the question optimization metamorphic relation (CMR1).

[0084] b) Combining the chain of thought prompt, multilingual translation, and the question optimization metamorphic relation (CMR2).

[0085] c) Combining all QMRs (CMR3).

[0086] It is worth mentioning that for the construction of the question metamorphic relation, the answering metamorphic relation, and the composite metamorphic relation for the hallucination detection process:

[0087] a) When constructing the multilingual translation metamorphic relation (QMR2), multiple other languages can be selected.

[0088] b) When constructing the question optimization metamorphosis relationship (QMR3), the above scheme uses the form of role-playing to prompt the large model. Different forms of prompt words can also be used to prompt the model to achieve question optimization. For example: few-shot prompting ("Please help me optimize the question with reference to the following examples: [source question]. Example 1: [example question 1]; Optimized question: [optimized question of example question 1]; Example 2: ……; Example 3: …… Example n: ……").

[0089] c) When constructing the external knowledge introduction metamorphosis relationship (QMR4), the external knowledge can be sourced from the constructed knowledge base or from other large models. For example, use the following prompt words to ask other large models: "Please provide the background knowledge related to the following question: [source question]".

[0090] d) When constructing the composite metamorphosis relationship (CMR), multiple basic MRs (question metamorphosis relationship and answer metamorphosis relationship) can be arbitrarily arranged and combined to construct the composite metamorphosis relationship, and the subsequent questions can be generated using the composite metamorphosis relationship.

[0091] Furthermore, based on the above-mentioned thinking chain prompt metamorphosis relationship, multilingual translation metamorphosis relationship, question optimization metamorphosis relationship, external knowledge introduction metamorphosis relationship, construction of skeptical question metamorphosis relationship, construction of multiple-choice question metamorphosis relationship, and composite metamorphosis relationship, the pre-set question-answer metamorphosis relationship is obtained, and the combination of the source question and the source answer is metamorphosed through the pre-set question-answer metamorphosis relationship to generate subsequent questions. Examples are shown in Tables 1-1 to 1-3 as follows:

[0092] Table 1-1 Example of metamorphosis test hallucination detection (QMR)

[0093]

[0094] Table 1-2 Example of metamorphosis test hallucination detection (AMR)

[0095]

[0096] Table 1-3 Example of metamorphosis test hallucination detection (CMR)

[0097]

[0098] Step 102: Perform hallucination detection on the source answer based on the subsequent answer to generate a hallucination detection result.

[0099] Specifically, Step 102 may include the following sub-steps S21-S24:

[0100] Step S21: Preprocess the source answer and the subsequent answer respectively, and output the target source answer and the target subsequent answer;

[0101] Step S22: Determine whether the target source answer and the target subsequent answer are consistent;

[0102] Step S23: If they are consistent, determine the hallucination detection result as no hallucination;

[0103] Step S24: If they are inconsistent, determine the hallucination detection result as having hallucination.

[0104] It should be noted that after selecting any one of the above transformation relationships and transforming the source question into the subsequent question Q’ = MR(Q, A), it is sent into the black-box large language model M to generate the subsequent answer A’ = M(Q’). The answer consistency test is used to judge the consistency between the source answer and the subsequent answer. If they are consistent, the source answer is not a hallucination; if they are inconsistent, the source answer is a hallucination.

[0105] Specifically, the consistency between A and A’ is verified through the following steps:

[0106] 1) Answer preprocessing:

[0107] For natural language questions, simplify the long text answer, such as prompting the large model "Please summarize it in one sentence" to extract the core semantic content of the answer.

[0108] For programming questions, assign the same test input that conforms to the question format to the generated code, obtain the output result by executing the code, and estimate the consistency of the generated code by comparing the consistency of the code output results.

[0109] 2) Consistency determination:

[0110] For the simplified natural language answer, automatically judge the consistency based on the large model. The prompt template is: "Are the above two answers consistent? Please answer with 'yes' or 'no', no explanation is required."

[0111] For programming questions, if the output results of the two generated codes are the same under the same test input, it is determined to be logically equivalent.

[0112] It is worth mentioning that for natural language questions, when performing answer preprocessing, an existing abstract generation model can be used to simplify the long text answer, or answer preprocessing can be skipped and the consistency determination step can be directly carried out. And when performing consistency determination, other methods can be used to determine semantic consistency, such as using a language model to obtain the embedding vectors of the two answers and calculating the cosine similarity of the two embedding vectors to judge whether the two answers are consistent.

[0113] For programming questions, answer preprocessing can be skipped and a code clone detector can be directly used to judge whether the two generated codes are consistent.

[0114] It is worth mentioning that multiple subsequent questions and subsequent answers can also be generated using the constructed metamorphic relationship, and the answer that appears most frequently is selected from multiple subsequent answers by majority voting as the final subsequent answer. Whether the source answer is an hallucination answer is determined by verifying the consistency between the final subsequent answer and the source answer.

[0115] Step 103: Based on the source question and the pre-set error correction metamorphic relationship, use the pre-set correction technology and the black-box large language model to correct the source answer according to the hallucination detection result, and output the corrected answer.

[0116] The pre-set correction technology includes a pre-set large model and the cosine similarity method.

[0117] It should be noted that those skilled in the art can choose any one of the pre-set large model and the cosine similarity method for hallucination correction as needed, and the present invention does not make specific limitations.

[0118] Specifically, step 103 may include the following sub-steps S31-S35:

[0119] Step S31: When the hallucination detection result is determined to be no hallucination, use the source answer as the corrected answer;

[0120] Step S32: When the hallucination detection result is determined to be hallucination, use the pre-set error correction metamorphic relationship to metamorphose the source question and the source answer to generate multiple error correction subsequent questions;

[0121] It should be noted that after detecting hallucinations, the present invention has the ability to correct hallucinations. Specifically, the present invention generates multiple error correction subsequent questions through the error correction metamorphic relationship (ECMR), that is, the pre-set error correction metamorphic relationship used in the hallucination correction process, so as to generate error correction subsequent answers under multiple execution paths, and correct the hallucination answer based on the pre-set large model or the cosine similarity method.

[0122] Furthermore, for the construction of the pre-set error correction metamorphic relationship, ECMR generates differentiated error correction subsequent questions by expanding or combining multiple basic metamorphic relationships to trigger diverse execution paths of the model. It includes:

[0123] a) Introduce a new question form metamorphic relation (ECMR1): Trigger the model to rethink the answer by questioning the source answer ("[Source question]+[Source answer], is this statement false? If incorrect, please give the correct answer"). In addition to negative questioning (AMR1), the following forms can also be introduced: affirmative questioning ("[Source question]+[Source answer], is this statement correct? If incorrect, please give the correct answer"), which allows the model to verify the correctness of its own answer through affirmative questioning. Neutral questioning ("[Source question]+[Source answer], please judge whether this statement is correct? If incorrect, please give the correct answer"), which uses a neutral expression to stimulate the model to objectively evaluate the answer.

[0124] b) Introduce a new language metamorphic relation (ECMR2): The original execution path of multilingual translation is to translate the question into three languages (such as Spanish, German, Dutch) (QMR2). On this basis, one more language (such as French) can be added or less commonly used languages (such as Arabic, Russian) can be selected to further expand the diversity of the execution path.

[0125] c) Introduce a new question optimization metamorphic relation (ECMR3): Based on question optimization (QMR3), new optimization questions can be constructed by changing the optimization prompt words. For example, prompt the model to only change the words in the sentence or only change the sentence structure to optimize the question.

[0126] d) Introduce a new external knowledge metamorphic relation (ECMR4): Based on introducing external knowledge (QMR4), the types and quantities of the introduced external knowledge can be further increased, and then different external knowledge can be used to construct follow-up error correction questions.

[0127] e) Composite Error Correction Metamorphic Relation (CECMR): The Composite Error Correction Metamorphic Relation combines ECMR2, ECMR3, and ECMR4 to guide the model to adopt more different execution paths, utilize the stability of the correct answer, and thus vote to obtain the correct answer from the answers obtained from multiple different execution paths to correct the model hallucination.

[0128] It is worth mentioning that for the construction of the pre-set error correction metamorphic relation for hallucination correction, it can also be completed through the following methods:

[0129] a) When constructing and introducing a new language metamorphic relation (ECMR2), multiple other languages can be selected.

[0130] b) When constructing the new problem optimization metamorphosis relationship (ECMR3), different prompting words can be used to prompt the model to optimize the problem. For example: few-shot prompting ("Please help me optimize the problem with reference to the following examples: [source problem]. Example 1: [example problem 1]; Optimized problem: [optimized problem of example problem 1]; Example 2: ……; Example 3: …… Example n: ……").

[0131] c) When constructing the new external knowledge metamorphosis relationship (ECMR4), the new external knowledge introduced can be sourced from the constructed knowledge base or from other large models. For example, use the following prompting words to ask other large models: "Please provide the background knowledge related to the following problem: [source problem]".

[0132] d) When constructing the composite error correction metamorphosis relationship (CECMR), multiple basic MRs (problem metamorphosis relationship and answer metamorphosis relationship) can be arbitrarily arranged and combined to construct a composite metamorphosis relationship, and multiple different composite metamorphosis relationships, or multiple implementation methods of the basic MRs in the same composite metamorphosis relationship (such as introducing new languages in ECMR2), can be used to construct multiple error correction follow-up questions for hallucination correction.

[0133] Step S33: Respectively use each error correction follow-up question as the input of the black-box large language model, and output the error correction follow-up answers corresponding to each error correction follow-up question;

[0134] Step S34: Use the pre-set large model to conduct a majority vote on each error correction follow-up answer, and output the corrected answer;

[0135] It should be noted that any one of the above ECMRs is selected to generate multiple error correction follow-up questions, and the corresponding multiple error correction follow-up answers are obtained. The multiple error correction follow-up answers are preprocessed, and the preprocessing process is the same as the principle of the steps for preprocessing the source answer and follow-up answers above. Subsequently, use the large model (i.e., the pre-set large model, which is a pre-set local large model, and can be any large model with excellent performance, or the target black-box large model that needs to detect and correct hallucinations) to conduct a majority vote on all the preprocessed error correction follow-up answers, and select the answer with the highest frequency of occurrence (from the perspective of semantic similarity) among the error correction follow-up answers as the final answer (i.e., the corrected answer) to correct the source answer. The prompting words for the large model are: "Please select the answer with the most occurrences from the following answers (note that only output the answer without any explanation): [error correction follow-up answer 1], [error correction follow-up answer 2], ……, [error correction follow-up answer n]". Examples are shown in Tables 2-1 to 2-2:

[0136] Table 2-1 Example of multi-path voting hallucination correction (ECMR)

[0137]

[0138] Table 2-2 Example of Multipath Voting Hallucination Correction (CECMR)

[0139]

[0140] Step S35: Or use the cosine similarity method to perform majority voting on each subsequent error correction answer and output the corrected answer.

[0141] Furthermore, step S35 may include the following sub-steps S351 - S358:

[0142] Step S351: Perform feature embedding on each subsequent error correction answer to generate a sentence embedding vector corresponding to each subsequent error correction answer;

[0143] Step S352: Calculate the cosine similarity of each sentence embedding vector to determine the similarity between each pair of subsequent error correction answers;

[0144] Step S353: Use the similarity between each pair of subsequent error correction answers as the matrix element value to construct a similarity matrix;

[0145] Step S354: Compare each matrix element value in the similarity matrix with a preset similarity threshold respectively;

[0146] Step S355: Take any matrix element value greater than the preset similarity threshold as the first matrix element value;

[0147] Step S356: Take any matrix element value less than or equal to the preset similarity threshold as the second matrix element value;

[0148] Step S357: Count the number of first matrix element values corresponding to each matrix row in the similarity matrix to determine the number of first matrix element values corresponding to each matrix row in the similarity matrix;

[0149] Step S358: Take the subsequent error correction answer corresponding to the largest number of first matrix element values as the corrected answer.

[0150] It should be noted that when selecting the final answer from multiple error correction follow-up answers {A1’, A2’, …, An’}, the similarity between two error correction follow-up answers can be calculated based on the edit distance or the cosine similarity of sentence embedding vectors obtained from pre-trained language models (such as BERT, GPT, etc.). That is, calculate the edit distance pairwise between each error correction follow-up answer and all other error correction follow-up answers, or calculate the cosine similarity pairwise between the sentence embedding vectors corresponding to each error correction follow-up answer and all other error correction follow-up answers, to determine the multiple similarities corresponding to each error correction follow-up answer, that is, the similarities pairwise between each error correction follow-up answer and all other error correction follow-up answers. Use the multiple similarities corresponding to all error correction follow-up answers to construct a similarity matrix (the matrix element values in each row of the matrix correspond to the multiple similarities corresponding to the error correction follow-up answers), and set a preset similarity threshold. Binarize the similarity matrix according to the threshold to convert it into a 0-1 matrix (set the matrix element value greater than the threshold to 1, that is, the first matrix element value, and vice versa to 0, that is, the second matrix element value), and take the error correction follow-up answer corresponding to the matrix row with the most 1s in the matrix (that is, the answer with the highest frequency of occurrence among the error correction follow-up answers) as the final answer (corrected answer). Among them, if there are multiple error correction follow-up answers with the highest frequency of occurrence, take any one of them as the final answer (corrected answer).

[0151] For example, currently there are three error correction follow-up answers, and each error correction follow-up answer corresponds to two similarities. Take the two similarities corresponding to the first error correction follow-up answer as the matrix element values of the first matrix row, take the two similarities corresponding to the second error correction follow-up answer as the matrix element values of the second matrix row, and take the two similarities corresponding to the third error correction follow-up answer as the matrix element values of the third matrix row to obtain a similarity matrix with three rows and two columns. Then set a preset similarity threshold, binarize the similarity matrix according to the threshold to convert it into a 0-1 matrix (set the matrix element value greater than the threshold to 1, that is, the first matrix element value, and vice versa to 0, that is, the second matrix element value), and take the error correction follow-up answer corresponding to the matrix row with the most 1s in the matrix (that is, the answer with the highest frequency of occurrence among the error correction follow-up answers) as the final answer (corrected answer). Among them, if there are multiple error correction follow-up answers with the highest frequency of occurrence, take any one of them as the final answer (corrected answer); for example, if the number of 1s in the first matrix row of the similarity matrix is 1, the number of 1s in the second matrix row is 1, and the number of 1s in the third matrix row is 0, then take any one of the error correction follow-up answers corresponding to the first matrix row and the error correction follow-up answer corresponding to the second matrix row as the final answer (corrected answer).

[0152] As a comparison of technical effects, it can be referenced in combination with the existing technology. With the development of large language models, there are currently three mainstream methods to alleviate model hallucinations:

[0153] Chain of Thought: In multi-step reasoning tasks (such as arithmetic or logical problems), large language models often make mistakes, leading to hallucinations. Some methods improve the performance of the models by guiding them to think step by step. For example, CoT (Chain-of-Thought) guides the language model to generate intermediate reasoning processes by prompting the large model to think step by step (such as adding a prompt after the question: "Please think step by step"), thereby improving the accuracy of the model in answering complex questions;

[0154] Self-Consistency Method: CoT-SC (Chain-of-Thought with Self-Consistency) guides the language model to generate reasoning steps through few-shot prompting (adding example questions and their detailed reasoning steps and answers as few-shot examples), and adjusts the "temperature parameter" of decoding to increase the randomness of the answers. Then, multiple reasoning paths are randomly sampled, and finally, the answer with the highest frequency of occurrence among the answers generated by different paths is selected as the final result.

[0155] Uncertainty Estimation: Some methods judge the reliability of the output content by measuring the confidence of the model in the generated content. If the model shows low confidence when generating content, it may indicate the occurrence of hallucinations. For example, the minimum, maximum, or average probability of all tokens in the model's output answer is used as the uncertainty estimate.

[0156] Furthermore, the method based on the chain of thought is overly dependent on the performance of the model itself, such as its reasoning ability. However, when the model has hallucinations, its reasoning path is unreliable. This makes it difficult for existing methods based on the chain of thought to effectively solve the model's hallucinations.

[0157] The method based on self-consistency only adjusts the "temperature parameter" of decoding to increase the randomness of the answers. However, the answers obtained by adjusting the "temperature parameter" of decoding have a high degree of repetition and cannot effectively increase the diversity of reasoning paths, making it difficult to apply to the unstable characteristics of the model's hallucinatory answers.

[0158] The method based on uncertainty estimation often requires access to the intermediate layer output of the model. However, mainstream commercial large models are black boxes, and users cannot obtain the intermediate layer output of these black box models, so they cannot use these methods to solve the model's hallucinations.

[0159] In addition, large language models distort facts and make non-factual statements (i.e., factual hallucinations) when processing question-and-answer tasks, which affects the wide application of large language models. Existing technologies usually use the confidence of the model in the generated content for hallucination detection and correction. These methods are difficult to apply in black box models and have low reliability.

[0160] In view of the above problems, the present invention proposes a method for detecting and correcting hallucinations in black-box large language models, aiming to detect and correct hallucinations in large language models by introducing metamorphic testing and utilizing the instability of hallucinatory answers. Specifically, the present invention utilizes the characteristic that hallucinatory answers of the model are unstable while non-hallucinatory answers are relatively stable, generates subsequent questions through metamorphic testing, and obtains subsequent answers. By detecting the consistency between the subsequent answers and the source answer, it detects whether the source answer is a hallucination, thus achieving hallucination detection. At the same time, when the source answer is a hallucination, it utilizes the characteristic that hallucinatory answers of the model are unstable while non-hallucinatory answers are relatively stable, generates multiple error-correction subsequent questions through metamorphic testing, and obtains multiple error-correction subsequent answers. Through the majority voting method, it selects the error-correction subsequent answer that accounts for the majority from multiple error-correction subsequent answers to correct the source answer, thus achieving hallucination correction. In addition, by considering the mechanisms behind the question-answering ability of large language models, including question understanding, knowledge recall, and knowledge reasoning abilities, etc., it designs metamorphic relationships for hallucination detection, including question metamorphic relationships, answer metamorphic relationships, and composite metamorphic relationships used in the hallucination detection process. And by considering the mechanisms behind the question-answering ability of large language models, including question understanding, knowledge recall, and knowledge reasoning abilities, etc., it designs error-correction metamorphic relationships.

[0161] In summary, the current best self-consistency-based method, CoT-SC, generates diverse answers by adjusting the "temperature parameter" and uses the most consistent answer to correct hallucinations. Compared with the prior art, the present invention adopts the idea of metamorphic testing, generates subsequent questions by designing a series of metamorphic relationships, and can more effectively guide the model to adopt different execution paths to solve the same problem, thereby better utilizing the instability of hallucinatory answers to detect and correct hallucinatory answers.

[0162] At the same time, the present invention avoids relying on the model's own reasoning ability or confidence. This enables the present invention to accurately identify and correct factual hallucinations even when the model's reasoning ability is insufficient or the confidence is unreliable / unobtainable. Therefore, compared with the method based on the chain of thought and the method based on uncertainty estimation, the present invention also has significant advantages.

[0163] In an embodiment of the present invention, the present invention provides a method for detecting and correcting hallucinations in a black-box large language model. The method includes obtaining a source question, and based on the source question and a preset question-answer metamorphic relationship, using the black-box large language model to generate a source answer and a subsequent answer; performing hallucination detection on the source answer according to the subsequent answer to generate a hallucination detection result; based on the source question and a preset error correction metamorphic relationship, using a preset correction technique and the black-box large language model to correct the source answer according to the hallucination detection result and output a corrected answer; based on the above solution, combining the preset question-answer metamorphic relationship and the preset error correction metamorphic relationship, detecting and correcting the source answer generated by the black-box large language model based on the source question, and outputting the process of the corrected answer. The present invention does not need to calculate the confidence of the content generated by the model, avoiding the dependence on the confidence. In the case where the confidence is unreliable / unobtainable, it can effectively identify and correct the hallucination problem of the content generated by the model.

[0164] Please refer to Figure 4 , Figure 4 which is a structural block diagram of a device for detecting and correcting hallucinations in a black-box large language model provided in the second embodiment of the present invention.

[0165] A device for detecting and correcting hallucinations in a black-box large language model provided by the present invention includes:

[0166] An acquisition module 401, configured to obtain a source question, and based on the source question and a preset question-answer metamorphic relationship, use the black-box large language model to generate a source answer and a subsequent answer;

[0167] A detection module 402, configured to perform hallucination detection on the source answer according to the subsequent answer to generate a hallucination detection result;

[0168] A correction module 403, configured to based on the source question and a preset error correction metamorphic relationship, use a preset correction technique and the black-box large language model to correct the source answer according to the hallucination detection result and output a corrected answer.

[0169] Further, the acquisition module 401 is specifically configured to:

[0170] Input the source question into the black-box large language model to generate a source answer;

[0171] Use the preset question-answer metamorphic relationship to perform metamorphosis on the source question and the source answer to generate a subsequent question;

[0172] Input the subsequent question into the black-box large language model to generate a subsequent answer.

[0173] Further, the detection module 402 is specifically configured to:

[0174] Preprocess the source answer and the subsequent answer respectively, and output a target source answer and a target subsequent answer;

[0175] Determine whether the target source answer and the target subsequent answer are consistent;

[0176] If they are consistent, determine the hallucination detection result as no hallucination;

[0177] If they are inconsistent, determine the hallucination detection result as having hallucinations.

[0178] Furthermore, the preset correction technology includes a preset large model and the cosine similarity method; the correction module 403 includes:

[0179] The first sub-module is used to take the source answer as the corrected answer when the hallucination detection result is determined to be no hallucination;

[0180] The second sub-module is used to, when the hallucination detection result is determined to be having hallucinations, perform metamorphosis on the source question and the source answer using the preset error correction metamorphosis relationship to generate multiple error correction subsequent questions;

[0181] The third sub-module is used to take each error correction subsequent question as the input of the black-box large language model and output the error correction subsequent answer corresponding to each error correction subsequent question;

[0182] The fourth sub-module is used to perform majority voting on each error correction subsequent answer using the preset large model and output the corrected answer;

[0183] The fifth sub-module is used to, or perform majority voting on each error correction subsequent answer using the cosine similarity method and output the corrected answer.

[0184] Furthermore, the fifth sub-module is specifically used for:

[0185] Perform feature embedding on each error correction subsequent answer to generate sentence embedding vectors corresponding to each error correction subsequent answer;

[0186] Perform cosine similarity calculation on each sentence embedding vector to determine the similarity between each pair of error correction subsequent answers;

[0187] Construct a similarity matrix with the similarity between each pair of error correction subsequent answers as the matrix element value;

[0188] Compare each matrix element value in the similarity matrix with the preset similarity threshold respectively;

[0189] Take any matrix element value greater than the preset similarity threshold as the first matrix element value;

[0190] Take any matrix element value less than or equal to the preset similarity threshold as the second matrix element value;

[0191] Count the number of the first matrix element values corresponding to each matrix row in the similarity matrix to determine the number of the first matrix element values corresponding to each matrix row in the similarity matrix;

[0192] Use the subsequent error correction response corresponding to the largest number of first matrix element values as the corrected response.

[0193] Furthermore, the preset question-answer transformation relationships include the thought chain prompt transformation relationship, the multilingual translation transformation relationship, the question optimization transformation relationship, the introduction of external knowledge transformation relationship, the construction of skeptical question transformation relationship, and the construction of multiple-choice question transformation relationship.

[0194] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described devices, modules, and sub-modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0195] The embodiments of the present invention also provide a computer device, including a memory and a processor, where a computer program is stored in the memory; when the computer program is executed by the processor, the processor is caused to execute the steps of the black-box large language model hallucination detection and correction method according to any one of the above embodiments.

[0196] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed, the steps of the black-box large language model hallucination detection and correction method according to any one of the above embodiments are implemented.

[0197] The embodiments of the present invention also provide a computer program product, including a computer program stored on a non-transitory computer-readable storage medium, where the computer program includes a program / instruction, and when the program / instruction is executed by a computer, the computer is caused to execute the steps of the black-box large language model hallucination detection and correction method according to any one of the above embodiments.

[0198] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0199] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for detecting and correcting hallucinations in black-box large language models, characterized in that, Including: Obtain the source question, and based on the source question and the preset question-answer metamorphosis relationship, use a black-box large language model to generate a source answer and subsequent answers; Perform hallucination detection on the source answer according to the subsequent answers to generate a hallucination detection result; Based on the source question and the preset error correction metamorphosis relationship, use a preset correction technique and the black-box large language model to correct the source answer according to the hallucination detection result, and output a corrected answer.

2. The method for detecting and correcting hallucinations of the black-box large language model according to claim 1, wherein, The step of using a black-box large language model to generate a source answer and subsequent answers based on the source question and the preset question-answer metamorphosis relationship includes: Input the source question into the black-box large language model to generate a source answer; Use the preset question-answer metamorphosis relationship to metamorphose the source question and the source answer to generate subsequent questions; Input the subsequent questions into the black-box large language model to generate subsequent answers.

3. The method for detecting and correcting hallucinations of the black-box large language model according to claim 1, characterized in that The step of performing hallucination detection on the source answer according to the subsequent answers to generate a hallucination detection result includes: Preprocess the source answer and the subsequent answers respectively to output a target source answer and a target subsequent answer; Determine whether the target source answer and the target subsequent answer are consistent; If they are consistent, determine the hallucination detection result as no hallucination; If they are inconsistent, determine the hallucination detection result as having hallucination.

4. The method for detecting and correcting hallucinations in a black-box large language model according to claim 1, characterized in that, The preset correction technique includes a preset large model and the cosine similarity method; the step of using the preset correction technique and the black-box large language model to correct the source answer according to the hallucination detection result based on the source question and the preset error correction metamorphosis relationship, and output a corrected answer includes: When the hallucination detection result is determined as no hallucination, use the source answer as the corrected answer; When the hallucination detection result is determined as having hallucination, use the preset error correction metamorphosis relationship to metamorphose the source question and the source answer to generate multiple error correction subsequent questions; Respectively take each of the error correction subsequent questions as the input of the black-box large language model, and output the error correction subsequent answers corresponding to each of the error correction subsequent questions; Use the preset large model to perform majority voting on each of the error correction subsequent answers, and output a corrected answer; Or use the cosine similarity method to perform majority voting on each of the error correction subsequent answers, and output a corrected answer.

5. The method for detecting and correcting hallucinations in a black-box large language model according to claim 4, characterized in that, The step of using the cosine similarity method to perform majority voting on each of the error correction subsequent answers and output a corrected answer includes: Perform feature embedding on each of the error correction subsequent answers to generate sentence embedding vectors corresponding to each of the error correction subsequent answers; Perform cosine similarity calculation on each of the sentence embedding vectors to determine the similarity between each pair of the error correction subsequent answers; Use the similarity between each pair of the error correction subsequent answers as the matrix element value to construct a similarity matrix; Compare each matrix element value in the similarity matrix with a preset similarity threshold respectively; Take any matrix element value greater than the preset similarity threshold as the first matrix element value; Take any matrix element value less than or equal to the preset similarity threshold as the second matrix element value; Count the number of first matrix element values corresponding to each matrix row in the similarity matrix, and determine the number of first matrix element values corresponding to each matrix row in the similarity matrix; Use the subsequent answer corresponding to the largest number of first matrix element values as the corrected answer.

6. The method for detecting and correcting hallucinations of the black-box large language model according to claim 1, characterized in that, The preset question-answer metamorphosis relationships include the thought chain prompt metamorphosis relationship, the multilingual translation metamorphosis relationship, the question optimization metamorphosis relationship, the introduction of external knowledge metamorphosis relationship, the construction of skeptical question metamorphosis relationship, and the construction of multiple-choice question metamorphosis relationship.

7. An apparatus for detecting and correcting hallucinations in a black-box large language model, characterized in that, It includes: An acquisition module for acquiring a source question and generating a source answer and subsequent answers based on the source question and the preset question-answer metamorphosis relationships using a black-box large language model; A detection module for performing hallucination detection on the source answer according to the subsequent answers and generating a hallucination detection result; A correction module for correcting the source answer according to the hallucination detection result using a preset correction technique and the black-box large language model based on the source question and the preset error correction metamorphosis relationships, and outputting a corrected answer.

8. A computer device, characterized in that, It includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the black-box large language model hallucination detection and correction method according to any one of claims 1-6.

9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed, the steps of the black-box large language model hallucination detection and correction method according to any one of claims 1-6 are implemented.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes programs / instructions. When the programs / instructions are executed by a computer, the computer executes the steps of the black-box large language model hallucination detection and correction method according to any one of claims 1-6.

Citation Information

Cited By

  • Group consensus large model illusion reduction method based on multi-model question

    CN120744065A

  • A group consensus large model illusion reduction method based on multi-model interrogation

    CN120744065B

  • Student behavior intelligent correction and feedback method and system based on action logic

    CN122415285A