Illusion detection method for model, medium, electronic equipment and program product

By performing fine-grained segmentation of the content generated by the large language model and using a lightweight expert model for detection, the efficiency and accuracy issues of faithfulness illusion recognition in model-generated content are solved, achieving efficient and reliable illusion detection.

CN121636713APending Publication Date: 2026-03-10BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and locate fidelity illusions in content generated by large language models, especially when applied across different domains where detection results are unreliable, and they consume significant computational resources and prolong inference time.

Method used

By performing fine-grained segmentation of the model-generated content, using a lightweight expert model to perform consistency checks on each content segment, and combining key information extraction and consistency detection modules, inconsistent or fictitious content segments can be identified and located.

Benefits of technology

It improves the efficiency and accuracy of model illusion detection, reduces computational resource consumption and response latency, and significantly enhances detection accuracy and robustness, making it suitable for online and large-scale production environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636713A_ABST
    Figure CN121636713A_ABST
Patent Text Reader

Abstract

The invention discloses an illusion detection method for a model, a medium, electronic equipment and a program product, and the method comprises the steps: obtaining a first content generated by a first model, and obtaining a first context used when the first model generates the first content; obtaining at least one first content segment from the first content; for each first content segment, using a second model to perform content detection on the first content segment based on the first context to obtain a content detection result corresponding to the first content segment, the content detection result being used for representing a consistency state of the first content segment and the first context; and outputting an illusion detection result for the first model based on the content detection result corresponding to each first content segment. The detection efficiency and the detection precision of the model illusion are effectively improved, and the content segments with the model illusion problem can be accurately identified and positioned based on the consistency state between each content segment and the context.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of models, in particular to a hallucination detection method for a model, a medium, an electronic device and a program product. BACKGROUND

[0002] With the development of Large Language Models (LLM) technology, the model hallucination problem has gradually become the focus of attention. The so-called model hallucination refers to the generation of information that is inconsistent with facts, fictitious or misleading by the LLM when generating content.

[0003] In related technologies, the rule-based method measures the overlap between the LLM generated content and the context, which is difficult to capture the semantic level of the model hallucination problem. Although the self-evaluation method based on LLM can detect from the semantic level, the reasoning is time-consuming, low in efficiency and low in accuracy. SUMMARY

[0004] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0005] In a first aspect, the present disclosure provides a hallucination detection method for a model, the method comprising: obtaining first content generated by a first model, and obtaining a first context used by the first model when generating the first content; obtaining at least one first content segment from the first content; for each first content segment, using a second model to perform content detection on the first content segment based on the first context, to obtain a content detection result corresponding to the first content segment, the content detection result being used to represent a consistency state of the first content segment and the first context; outputting a hallucination detection result for the first model based on the content detection result corresponding to each first content segment.

[0006] In a second aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the method in the first aspect.

[0007] In a third aspect, the present disclosure provides an electronic device comprising: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method in the first aspect.

[0008] In a fourth aspect, the present disclosure provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of the first aspect.

[0009] By the above technical solution, the first content generated by the first model and the first context used by the first model to generate the first content are obtained, at least one first content segment is obtained from the first content, then for each first content segment, the second model is used to perform content detection on the first content segment based on the first context, to obtain a content detection result representing the consistency state of the first content segment and the first context, and finally the hallucination detection result for the first model is output based on the content detection result corresponding to each first content segment. By finely granularly splitting the model-generated content, the second model can perform fine-grained detection on each content segment based on the context, which not only effectively improves the detection efficiency and detection accuracy of the model hallucination, but also accurately identifies and locates the content segment with the model hallucination problem based on the consistency state between the content segment and the context.

[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which: Figure 1 FIG. 1 is a flowchart of a hallucination detection method for a model according to an exemplary embodiment.

[0012] Figure 2 FIG. 2 is a structural diagram of a hallucination detection system for a model according to an exemplary embodiment.

[0013] Figure 3 FIG. 3 is a process diagram of a hallucination detection method for a model according to an exemplary embodiment.

[0014] Figure 4 FIG. 4 is a content diagram of an input prompt word according to an exemplary embodiment.

[0015] Figure 5 FIG. 5 is a structural diagram of a hallucination detection apparatus for a model according to an exemplary embodiment.

[0016] Figure 6 FIG. 6 is a structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood.

[0018] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0019] The term “comprising” and variations thereof as used herein are open-ended, and mean “including but not limited to”. The term “based on” means “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related terms are defined in the description that follows.

[0020] It should be noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly stated in the context.

[0022] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0023] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0024] For example, in response to receiving an active request of a user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide personal information to the software or hardware, such as an electronic device, an application program, a server, or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.

[0025] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, in which the prompt information may be presented in a textual manner. In addition, the pop-up window may also carry selection controls for the user to select “agree” or “disagree” to provide personal information to the electronic device.

[0026] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0027] At the same time, it can be understood that the data (including but not limited to the data itself, the obtaining or use of the data) involved in the technical solution should comply with the requirements of the corresponding laws and regulations and relevant provisions.

[0028] Model hallucination refers to the generation of information inconsistent with facts, fictitious, or misleading by an LLM when generating content. For example, inputting “which is the longest river in the world” to the LLM, the LLM outputs an answer “xx river”, but actually xx river is not the longest river. Or, when you ask the LLM to introduce the latest progress of a certain research direction, it can be reasonable and provide detailed information such as reference title and author, but when you search, you will find that those documents do not exist. These are typical manifestations of hallucination problems in actual business scenarios. Large language models are essentially probability language models: statistical distributions learned from massive amounts of data are used to predict tokens to ensure that the output is coherent and contextually appropriate, but not factually accurate. As LLMs are widely used in search, question answering, medical, financial, and other critical fields, such hallucination answers that are similar to “talking nonsense in a serious manner” not only affect user experience, but also may cause serious actual risks.

[0029] In the related art, model hallucination is mainly divided into two categories: factual hallucination and fidelity hallucination. Factual hallucination refers to the inconsistency between the generated content of the model and the real world facts, and fidelity hallucination refers to the inconsistency between the generated content and the input prompt or context. The hallucination detection method for the model provided in the present disclosure mainly aims at the problem of fidelity hallucination.

[0030] For the problem of fidelity hallucination, in the related art, a rule-based method usually uses statistical indicators or rules such as named entity recognition and keyword matching to measure the degree of overlap between the generated content of the LLM and the context, so as to evaluate the hallucination risk. However, such a method relies on surface text matching and is difficult to capture semantic contradictions or false information. At the same time, for hallucinations caused by synonym replacement, sentence transformation or context logical reasoning, the rule method is prone to false positives or false negatives, and the overall accuracy and robustness are low. The LLM self-evaluation method requires inputting all generated content and context into the LLM, which consumes a lot of computing resources and has a long reasoning delay, and is not suitable for online or large-scale production environments. The detection method based on the trained discriminant model is prone to false positives due to the lack of domain background knowledge when applied across domains, resulting in unreliable detection results.

[0031] Therefore, the present disclosure provides a model hallucination detection method, medium, electronic device and program product to solve the above technical problems.

[0032] Figure 1 is a flowchart of a model hallucination detection method according to an exemplary embodiment. As shown in Figure 1 the method can include the following steps: S101: obtaining first content generated by a first model, and obtaining a first context used by the first model to generate the first content.

[0033] wherein the first context includes first knowledge recalled from inside the model by the first model and / or second knowledge recalled from outside the model by the first model.

[0034] For example, the first knowledge is obtained by the first model during the training stage, and the second knowledge is retrieved by the first model during the model inference process, for example, from a pre-set knowledge base, which is not limited in the present disclosure.

[0035] wherein the first model is a target model that needs to be detected for model hallucination, which can be a large language model or other types of business models, such as a question and answer large model acting as an intelligent customer service, etc., which can be determined according to actual conditions, and the present disclosure does not limit this.

[0036] For example, assuming that a large language model is input with "what is xxx", the large language model generates corresponding reply content based on recalled content (including internal knowledge and / or external knowledge), which is the first content, and the recalled knowledge is the first context, wherein the recalled content can be a complete document or a knowledge fragment, which can be determined according to actual conditions, and the present disclosure does not limit this.

[0037] S102: obtaining at least one first content segment from the first content.

[0038] It should be understood that in real-world applications, the content generated by the model is often quite long and may contain multiple real or fake content fragments. If the entire long text is evaluated directly, it may easily mask inconsistencies or errors in certain parts of the text.

[0039] In this embodiment, the long text response generated by the model is split into several independent content segments so that each content segment can be detected separately. This enables more granular and accurate detection of fidelity illusions, which helps to improve the recall rate.

[0040] For example, such as Figure 2 The model illusion detection system 20 shown can split long texts into multiple fine-grained content fragments through the content splitting module 21.

[0041] For example, a text processing method based on regular expressions can be used to split the model's response according to punctuation marks, which is simple and efficient. For instance, a long text response generated by the model can be split by a period to obtain multiple content segments. Of course, other symbols can also be used to split the long text response generated by the model; this disclosure does not limit this. Alternatively, long text can be split by combining text length; for example, short sentences with a length less than a preset threshold can be not split, but treated as a single content segment along with the preceding and following clauses. Alternatively, a pre-trained splitting model can be used to split the long text response generated by the model, and so on; this disclosure does not limit this. Each content segment is treated as a semantically independent statement unit to achieve efficient structured processing of long text responses.

[0042] S103: For each first content segment, the second model is used to perform content detection on the first content segment based on the first context to obtain the content detection result corresponding to the first content segment. The content detection result is used to characterize the consistency state between the first content segment and the first context.

[0043] The consistency state includes at least one of the following: a semantically consistent state that represents the semantic consistency between the first content segment and the first context; a semantically conflicting state that represents the semantic conflict between the first content segment and the first context; and an unsupported state that represents the lack of basis for the first content segment in the first context.

[0044] In this embodiment, the second model is a detection model used to detect model illusions, which can also be understood as a consistency detection model. It is used to detect semantic consistency or conflict between the first content fragment and the first context, or to determine whether the first content fragment can be supported by the first context, i.e., whether the first content fragment is fictitious content lacking evidence. For example, if no corresponding content can be found in the context, it can be determined that the content fragment is unsupported or fictitious. The aforementioned second model can be a lightweight expert model trained by a model to achieve efficient content detection.

[0045] For example, such as Figure 2 The model illusion detection system 20 shown can detect the consistency status between the first content fragment and the first context through the consistency detection module 23.

[0046] In this embodiment, a second context used to generate the first content fragment can be determined first in the first context. Then, the second model performs content detection on the first content fragment based on the second context to obtain the content detection result corresponding to the first content fragment. This simplifies the content that the second model needs to analyze and detect, and further improves the detection efficiency of model illusions.

[0047] Assuming the first context includes document A, document B, and document C, and content fragment 1 in the first content is generated based on document A, then when detecting content fragment 1, document A and content fragment 1 are fed together into the second model for illusion detection. If the first content is generated by combining documents A, B, and C, and the generated content corresponding to a single content fragment cannot be determined, then for each content fragment, content detection is performed based on the entire context. The specific details can be determined according to the actual business scenario, and this disclosure does not impose any restrictions.

[0048] S104: Output the illusion detection results for the first model based on the content detection results corresponding to each first content segment.

[0049] For example, the content detection results corresponding to each content fragment can be summarized to obtain the hallucination detection results of the first model. For example, the content detection results corresponding to each content fragment can be used as the hallucination detection results of the first model. Statistics can be performed on the content detection results of different types, and the hallucination detection results of the first model can be determined based on the statistical results. For example, "number of consistent fragments: 1, number of inconsistent fragments: 2, number of fictitious fragments: 1". It is also possible to determine that the model-generated content has a model hallucination risk if any content fragment is inconsistent with the context or is fictitious content, etc. This disclosure does not limit this.

[0050] By employing the above method, the model-generated content is split into fine-grained segments so that the second model can perform fine-grained detection on each content segment based on the context. This not only effectively improves the detection efficiency and accuracy of model illusions, but also accurately identifies and locates content segments with model illusion problems based on the consistency between each content segment and the context.

[0051] In one possible manner, obtaining at least one first content fragment from the first content includes: splitting the first content to obtain at least one initial content fragment; performing content recognition on each initial content fragment to obtain a content recognition result; and characterizing the content recognition result in the at least one initial content fragment as a content fragment containing second content, and determining it as a first content fragment.

[0052] For example, such as Figure 3 As shown, taking the question-answering big data model as an example, the big data model responds to the user's input question by recalling the context from the knowledge base and generating the corresponding response content based on the recalled context. The generated content of the big data model is then broken down into multiple initial content fragments. Each initial content fragment can be identified to determine whether it contains key information. The content fragment containing key information among the multiple initial content fragments is identified as the first content fragment, which is equivalent to filtering the initial content fragments.

[0053] It should be understood that in real-world business scenarios, not all content generated by the model contains material requiring fidelity illusion detection. For example, some text, while grammatically correct and semantically fluent, may only involve subjective expressions, general descriptions, or sentence structures lacking substantial information, such as subjective descriptions or irrelevant chatter and clichés. Therefore, as... Figure 3 As shown, by identifying and filtering out content fragments that lack factual statement characteristics, and retaining content fragments containing key information (which can be entities or verifiable facts), the detection of hallucination risks can be made more targeted and accurate, reducing both invalid detections and false positives. In specific implementations, different strategies can be selected based on actual business needs; this disclosure does not impose any restrictions on this.

[0054] For example, such as Figure 2 The model illusion detection system 20 shown can identify and filter out content fragments that do not include key information such as entity objects and verifiable facts through the key information extraction module 22.

[0055] In one possible manner, the second content includes an entity object. Content recognition is performed on the initial content fragment to obtain a content recognition result, including: performing content recognition on the initial content fragment based on a preset recognition rule, and determining that the initial content fragment is a content recognition result containing the first entity object if the initial content fragment contains a first entity object that satisfies the preset recognition rule.

[0056] For example, in business scenarios where the focus is on whether the model-generated content contains fictitious or erroneous entity information, a regular expression-based entity recognition method can be used to quickly extract explicit information such as names of people, places, organizations, and times from the responses. This allows for the filtering of content fragments containing entity objects for subsequent detection. Specific preset recognition rules can be set according to requirements, and this disclosure does not impose any restrictions on them. This approach can satisfy the needs of targeted detection while reducing the content to be detected by the second model, effectively improving the efficiency and accuracy of model illusion detection.

[0057] In one possible approach, the second model is used to perform content detection on the first content fragment based on the first context, including: if the content recognition result of the first content fragment indicates that it contains a first entity object, the second model is used to perform entity information verification on the first entity object based on the first context.

[0058] Accordingly, when performing content detection using the second model, entity objects can be targeted for detection based on context, such as detecting whether they contain fictitious or erroneous entity information, thereby meeting the need for targeted detection of entity objects and effectively improving the accuracy of model illusion detection.

[0059] In one possible manner, the second content includes factual content. Content recognition is performed on the initial content fragment to obtain a content recognition result, including: performing semantic understanding on the initial content fragment through a third model, and determining the content recognition result representing that the initial content fragment contains the first factual content when the initial content fragment contains the first factual content.

[0060] For example, for businesses concerned with whether model responses contain false objective information, a third model can be introduced to perform semantic-level objectivity identification and classification of content fragments, thereby filtering out content fragments containing factual content for subsequent detection. For instance, if the initial content fragment is "The latest development in HH technology is the MM document released by the KK organization on xxxx-xx-xx," this content is verifiable factual content.

[0061] This not only meets the needs of targeted detection but also reduces the detection content of the second model, effectively improving the efficiency and accuracy of model illusion detection.

[0062] For example, the third model can be a lightweight discriminative model. Iterative training of the initial discriminative model can be performed by constructing sample content fragments containing factual content and sample content fragments not containing factual content until preset training completion conditions are met, such as model convergence and discrimination accuracy exceeding a preset threshold, to obtain a discriminative model that can identify whether a content fragment contains factual content.

[0063] In one possible approach, the second model is used to perform content detection on the first content fragment based on the first context, including: if the content recognition result of the first content fragment indicates that it contains first factual content, the second model is used to perform factual verification on the first factual content based on the first context.

[0064] Accordingly, when performing content detection using the second model, for content fragments containing factual content, the factual content can be verified in particular to determine whether the factual content is correct, thereby meeting the detection requirements for verifying factual content and effectively improving the accuracy of the model's illusion detection.

[0065] In this embodiment, as Figure 3 In the consistency detection step shown, the second model can use a lightweight expert model to determine the consistency status between the content generated by the first model and the context.

[0066] The second model can use a pre-trained large language model as a base, with the output layer set as a three-class linear layer to predict the consistency state between content fragments and context. The preset classification results include three categories: consistent, fictitious (unsupported), and conflicting (inconsistent), which correspond to three situations: the model-generated content is consistent with the context, the content cannot be found in the context, and the model-generated content has semantic conflicts with the context.

[0067] In one possible manner, the second model is trained as follows: A pre-trained large language model and training samples labeled with sample detection results are obtained. The training samples include sample content fragments and the sample context used by the first model when generating the sample content fragments. A second prompt word is constructed based on the training samples and a preset prompt word template, which includes preset classification results corresponding to each consistency state. The following steps are repeated until the preset model training completion conditions are met, and the trained large language model is used as the second model: The second prompt word is input into the large language model to obtain the predicted detection results output by the large language model. The model loss is calculated based on the predicted detection results and the sample detection results, and the model parameters of the large language model are updated based on the model loss.

[0068] For example, a training sample can be directly generated based on the sample content and sample context, or the sample content can be split into sample content fragments, and then a training sample can be generated based on the sample content fragments and sample context. This disclosure does not impose any restrictions on this.

[0069] For example, to enable the second model to more fully understand semantic relationships and accurately identify semantically inconsistent or fictitious content, this embodiment uses system prompts and formats the training samples as follows: Figure 4 The unified input structure is shown. The preset prompt word template includes preset classification results for three categories: correspondence, fiction, and conflict, and prompts the model to judge whether there is a risk of hallucination based on the generated content and context. The specific prompt word template can be constructed according to the needs, and this disclosure does not impose any restrictions on it.

[0070] During the training phase, formatted training samples are input into the initial large language model. For each sample, the model base generates a hidden state vector corresponding to the input sequence. The last hidden state vector of the last token (word segmentation) in the input sequence is taken as the input of the classification layer. After linear mapping and normalized exponential function calculation, the probability distributions corresponding to the three categories are obtained, and the predicted detection results are output. Then, cross-entropy is used to calculate the loss between the predicted detection results and the sample detection results, and backpropagation is performed to update the model parameters. The model training steps are repeated until the model convergence or the prediction accuracy is greater than a preset threshold, etc., to obtain the trained second model. Thus, the consistency state between the generated content fragment and the context can be predicted based on the second model.

[0071] In one possible approach, a second model is used to perform content detection on the first content fragment based on the first context to obtain the content detection result corresponding to the first content fragment. This includes: constructing a first prompt word based on a preset prompt word template, the first content fragment, and the first context. The preset prompt word template includes preset classification results corresponding to each consistency state. The first prompt word is input into the second model, and the second model performs semantic understanding on the first content fragment and the first context in the first prompt word to obtain the predicted probability corresponding to each preset classification result. The preset classification result with the highest predicted probability is determined as the content detection result corresponding to the first content fragment.

[0072] For example, similar to the model training phase, a structure like... Figure 4 The prompt words shown are input into the second model for classification and prediction. The preset classification result with the highest prediction probability is determined as the content detection result corresponding to the first content segment.

[0073] In practical applications, due to the time-sensitive and closed nature of the knowledge inherent in large language models, they cannot fully cover the latest or specific domain knowledge. Therefore, model application systems typically combine retrieval-augmented generation (RAG) technology, introducing external knowledge sources (such as databases, enterprise documents, or web page content) as context during the generation process, enabling LLMs to generate more reliable answers based on external evidence. However, when the content generated by the model conflicts with or cannot be supported by the provided context, a so-called illusion of fidelity is formed.

[0074] To identify such hallucination risks, this embodiment inputs the first content fragment obtained by filtering in the aforementioned key information extraction step, combined with its corresponding context, into a specially trained lightweight expert model to evaluate the degree of consistency between the content fragment and the context. The lower the consistency of the content fragment, the higher its hallucination risk, thereby achieving efficient identification and location of inconsistent or fictitious content in the model-generated content.

[0075] In this embodiment, as Figure 2 As shown, a model illusion detection system is provided, consisting of a content segmentation module, a key information extraction module, and a consistency detection module. Through the collaborative efforts of these modules, a method for detecting model fidelity illusions is implemented. This system primarily addresses the model illusion problem that easily arises within LLM (Layered Model) generation, where inconsistencies with the context can easily occur in real-world business scenarios. First, fine-grained content segmentation breaks down the long text content generated by the model into a set of independent content fragments. Then, a key information extraction mechanism filters out statements with objective factual attributes or semantic core value, filtering out redundancy and subjective components. Finally, a consistency detection model is used to determine the semantic consistency between the extracted key information and the supporting context, thereby achieving efficient and accurate identification and location of fidelity illusion risks, significantly improving the accuracy and robustness of model illusion detection.

[0076] By combining regularity-based text analysis with a lightweight consistency detection model, the aforementioned method effectively reduces computational resource consumption and response latency in the hallucination detection process, overcoming the limitations of large-model self-evaluation methods, which suffer from high computational resource consumption, slow inference, and difficulty in practical application. Furthermore, a multi-layered detection mechanism is constructed by introducing a content segmentation module and a key information extraction module. The content segmentation module significantly refines the granularity of hallucination detection, preventing local factual errors in long texts from being masked by the overall content. Simultaneously, the key information extraction module identifies and filters entity objects and factual content, improving the targeting and effectiveness of the detection. Based on the collaborative work of these modules, this method can achieve more refined and reliable fidelity hallucination recognition under low latency conditions, effectively solving the problems of insufficient detection accuracy and low recall that are commonly found in discriminative models in practical applications.

[0077] Figure 5 This is a schematic diagram illustrating the structure of a hallucination detection device for a model, according to an exemplary embodiment. Figure 5 As shown, the hallucination detection device 500 for the model includes: The first acquisition module 501 is used to acquire the first content generated by the first model, and to acquire the first context used by the first model when generating the first content; The second acquisition module 502 is used to acquire at least one first content fragment from the first content; The detection module 503 is used to perform content detection on each of the first content segments using a second model based on the first context, and obtain the content detection result corresponding to the first content segment. The content detection result is used to characterize the consistency state between the first content segment and the first context. The output module 504 is used to output the hallucination detection results for the first model based on the content detection results corresponding to each of the first content segments.

[0078] Optionally, the consistency state includes at least one of the following: The semantic consistency state represents the semantic consistency between the first content fragment and the first context; A semantic conflict state that characterizes the semantic conflict between the first content segment and the first context; This characterizes the unsupported state of the first content fragment, which lacks basis in the first context.

[0079] Optionally, the detection module 503 is used for: A first prompt word is constructed based on a preset prompt word template, the first content fragment, and the first context. The preset prompt word template includes preset classification results corresponding to each of the consistency states. The first prompt word is input into the second model. The second model performs semantic understanding on the first content fragment and the first context in the first prompt word to obtain the predicted probability corresponding to each preset classification result. The preset classification result with the highest predicted probability is determined as the content detection result corresponding to the first content fragment.

[0080] Optionally, the second model is trained in the following manner: Obtain a pre-trained large language model and training samples labeled with sample detection results, wherein the training samples include sample content fragments and the sample context used by the first model when generating the sample content fragments; A second prompt word is constructed based on the training samples and the preset prompt word template, wherein the preset prompt word template includes preset classification results corresponding to each of the consistency states; Repeat the following steps until the preset model training completion conditions are met, and use the trained large language model as the second model: The second prompt word is input into the large language model to obtain the prediction and detection results output by the large language model; The model loss is calculated based on the predicted detection results and the sample detection results, and the model parameters of the large language model are updated based on the model loss.

[0081] Optionally, the first context includes first knowledge recalled by the first model from within the model and / or second knowledge recalled by the first model from outside the model.

[0082] Optionally, the second acquisition module 502 is used for: The first content is split to obtain at least one initial content fragment; For each of the initial content segments, content recognition is performed on the initial content segment to obtain the content recognition result; The content recognition result of the at least one initial content segment is characterized as a content segment containing the second content, and is determined as the first content segment.

[0083] Optionally, the second content includes entity objects, and the second acquisition module 502 is used for: Based on preset recognition rules, the initial content fragment is identified. If the initial content fragment contains a first entity object that satisfies the preset recognition rules, the initial content fragment is identified as a content recognition result that contains the first entity object.

[0084] Optionally, the detection module 503 is used for: If the content recognition result of the first content fragment indicates that the first entity object is included, the second model is used to verify the entity information of the first entity object based on the first context.

[0085] Optionally, the second content includes factual content, and the second acquisition module 502 is used to: The initial content fragment is semantically understood using a third model. If the initial content fragment contains first factual content, the content recognition result representing that the initial content fragment contains the first factual content is determined.

[0086] Optionally, the detection module 503 is used for: If the content recognition result of the first content fragment indicates that it contains the first factual content, the second model is used to perform factual verification on the first factual content based on the first context.

[0087] Regarding the hallucination detection device for the model in the above embodiments, the method logic executed by each functional module has been described in detail in the section on methods, and will not be repeated here.

[0088] Based on the same concept, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of any of the above-described hallucination detection methods for a model.

[0089] Based on the same concept, this disclosure also provides an electronic device that may include: A storage device on which computer programs are stored; A processing device for executing a computer program stored in a storage device to implement the steps of any of the above-described hallucination detection methods for a model.

[0090] Based on the same concept, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described hallucination detection methods for a model.

[0091] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0092] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0093] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0094] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0095] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0096] In some implementations, communication can be conducted using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0097] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0098] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire first content generated by a first model, and acquire a first context used by the first model when generating the first content; acquire at least one first content fragment from the first content; for each first content fragment, perform content detection on the first content fragment using a second model based on the first context to obtain a content detection result corresponding to the first content fragment, the content detection result being used to characterize the consistency state between the first content fragment and the first context; and output a hallucination detection result for the first model based on the content detection results corresponding to each first content fragment.

[0099] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0101] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.

[0102] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0103] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0104] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0105] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0106] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A hallucination detection method for a model, the method comprising: The method comprises: obtaining first content generated by a first model, and obtaining a first context used by the first model to generate the first content; obtaining at least one first content segment from the first content; for each first content segment, performing content detection on the first content segment based on the first context by using a second model to obtain a content detection result corresponding to the first content segment, the content detection result being used to represent a consistency state of the first content segment and the first context; outputting a hallucination detection result for the first model based on the content detection result corresponding to each first content segment.

2. The method for hallucination detection against a model according to claim 1, wherein, The consistency state comprises at least one of: a semantic consistency state representing semantic consistency of the first content segment and the first context; a semantic conflict state representing semantic conflict of the first content segment and the first context; an unsupported state representing lack of basis of the first content segment in the first context.

3. The method for hallucination detection against a model according to claim 2, wherein, The content detection on the first content segment based on the first context by using the second model to obtain the content detection result corresponding to the first content segment comprises: constructing a first prompt word based on a preset prompt word template, the first content segment and the first context, the preset prompt word template comprising a preset classification result corresponding to each consistency state; inputting the first prompt word into the second model, performing semantic understanding on the first content segment and the first context in the first prompt word by the second model to obtain a prediction probability corresponding to each preset classification result, and determining a preset classification result with the maximum prediction probability as the content detection result corresponding to the first content segment.

4. The method for hallucination detection against a model according to claim 2, wherein, The second model is obtained by training in the following manner: obtaining a pre-trained large language model and training samples labeled with sample detection results, the training samples comprising a sample content segment and a sample context used by the first model to generate the sample content segment; constructing a second prompt word based on the training samples and a preset prompt word template, the preset prompt word template comprising a preset classification result corresponding to each consistency state; repeating the following steps until a preset model training completion condition is met, and using a large language model trained to completion as the second model: inputting the second prompt word into the large language model to obtain a predicted detection result output by the large language model; calculating a model loss based on the predicted detection result and the sample detection result, and updating model parameters of the large language model based on the model loss.

5. The method for hallucination detection against a model according to any one of claims 1-4, characterized in that, The first context comprises first knowledge recalled from inside the model by the first model and / or second knowledge recalled from outside the model by the first model.

6. The method for hallucination detection against a model according to any one of claims 1-4, wherein, The at least one first content segment obtained from the first content comprises: splitting the first content to obtain at least one initial content segment; for each initial content segment, performing content recognition on the initial content segment to obtain a content recognition result; The content recognition result of the at least one initial content segment indicates a content segment containing the second content.

7. The hallucination detection method for a model according to claim 6, wherein, The second content includes an entity object, and the content recognition on the initial content segment includes: The content recognition on the initial content segment is based on a preset recognition rule, and in a case where the initial content segment contains a first entity object satisfying the preset recognition rule, the initial content segment is determined to be a content recognition result indicating that the initial content segment contains the first entity object.

8. The hallucination detection method for a model according to claim 7, wherein, The content detection on the first content segment based on the first context by the second model includes: In a case where the content recognition result of the first content segment indicates that the first content segment contains the first entity object, the second model is used to verify entity information of the first entity object based on the first context.

9. The hallucination detection method for a model according to claim 6, wherein, The second content includes factual content, and the content recognition on the initial content segment includes: The initial content segment is subjected to semantic understanding by a third model, and in a case where the initial content segment contains first factual content, a content recognition result indicating that the initial content segment contains the first factual content is determined.

10. The hallucination detection method for a model according to claim 9, wherein, The content detection on the first content segment based on the first context by the second model includes: In a case where the content recognition result of the first content segment indicates that the first content segment contains the first factual content, the second model is used to verify the first factual content based on the first context.

11. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by the processing device to implement the steps of the method of any one of claims 1-10.

12. An electronic device, comprising: The computer program is executed by the processing device to implement the steps of the method of any one of claims 1-10. The computer program is executed by the processing device to implement the steps of the method of any one of claims 1-10. The computer program is executed by the processing device to implement the steps of the method of any one of claims 1-10.

13. A computer program product comprising a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Text comparison method, computer equipment and computer storage medium

    CN115017879A

  • Large model illusion detection method and device, electronic equipment and storage medium

    CN119622464A

  • Large model illusion correction method and device, equipment, medium and program product

    CN120086348A