Semantic drift distribution perception-based large language model hallucination detection method and system

CN122366574BActive Publication Date: 2026-09-22CHANGCHUN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610831649.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-22
Estimated Expiration
2046-06-10

AI Technical Summary

Technical Problem

[0003]现有幻觉检测方法主要分为以下几类:第一类是基于模型内部状态的白盒方法,通过访问模型的注意力权重、激活值或词嵌入矩阵进行检测,此类方法要求对模型架构具有深度访问权限,无法适用于商业API场景;第二类是基于一致性的黑盒方法,通过多次采样或温度调节比较输出分布,但此类方法依赖输出端的温度参数控制,属于半白盒方法;第三类是基于单一度量的无监督方法,如通过余弦相似度单一阈值判断,此类方法特征维度单一,判别能力有限

Benefits of technology

(1)完全黑盒检测:仅通过标准API调用获取大语言模型回复,无需访问模型内部权重、温度参数或激活值,适用于所有形式的大语言模型服务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122366574B_ABST
    Figure CN122366574B_ABST
Patent Text Reader

Abstract

The application discloses a hallucination detection method and system based on semantic drift distribution perception for a large language model, and belongs to the technical field of artificial intelligence and natural language processing. Without accessing internal parameters of the model, accurate and efficient hallucination detection is realized by inputting semantic disturbance and combining an independent semantic encoding model. The method specifically comprises the following steps: receiving an input pair containing an original query and a reply of the large language model, generating a set of disturbed queries; performing vector encoding on the reply and the set of disturbed queries by using a pre-trained semantic encoding model to obtain a reply embedding vector and a sequence of disturbed embedding vectors; calculating the cosine distance between the reply embedding vector and each disturbed embedding vector one by one to form a distance sequence; extracting multi-dimensional statistical features from the distance sequence to form a feature vector; inputting the feature vector into a pre-trained supervised classifier to output a hallucination detection result, and the supervised classifier is trained based on labeled hallucination data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and natural language processing technology, specifically relating to a method and system for hallucination detection based on semantic drift distribution perception of large language models. Background Technology

[0002] Large Language Models (LLMs) have demonstrated exceptional capabilities in text generation, question-answering systems, and information retrieval. However, their outputs suffer from factual errors, a phenomenon known as "hallucination." Hallucination refers to content generated by the model that contradicts objective facts or contains false information beyond the scope of known knowledge. In high-risk fields such as medicine, law, and finance, hallucinations can lead to serious consequences. Therefore, hallucination detection using large language models has become a core research topic in AI safety and trustworthy AI.

[0003] Existing hallucination detection methods can be mainly divided into the following categories: The first category is white-box methods based on the internal state of the model, which detect hallucinations by accessing the model's attention weights, activation values, or word embedding matrices. These methods require deep access to the model architecture and are not suitable for commercial API scenarios. The second category is black-box methods based on consistency, which compare the output distribution through multiple samplings or temperature adjustment. However, these methods rely on the temperature parameter control at the output end and belong to semi-white-box methods. The third category is unsupervised methods based on a single metric, such as judging by a single threshold of cosine similarity. These methods have a single feature dimension and limited discriminative ability.

[0004] The shortcomings of the existing technology are: (1) it cannot perform effective detection under completely black box conditions (only through standard API calls); (2) the feature expression dimension is single and fails to make full use of the multidimensional statistical information of semantic drift distribution; (3) it adopts unsupervised threshold judgment and cannot use labeled data to improve classification accuracy. Summary of the Invention

[0005] To address the aforementioned shortcomings, this invention proposes a completely black-box hallucination detection method based on multidimensional semantic drift statistical features and a supervised classifier. Without accessing the model's internal parameters, it achieves accurate and efficient hallucination detection by combining semantic perturbation at the input end with an independent semantic coding model.

[0006] The method is specifically as follows: S1, Disturbance Generation Step: Receive data containing the original query. and large language model response The input pair for the original query Perform semantic preservation perturbation operations to generate A set of perturbation queries ; S2, Embedding Encoding Step: Using a pre-trained semantic coding model, the response r and the perturbation query set are processed. Perform vector encoding separately to obtain the response embedding vector. and perturbation embedding vector sequence ; S3. Distance sequence calculation steps: Calculate the response embedding vector one by one. With each perturbation embedding vector The cosine distances between them form a distance sequence. ; S4. Multidimensional statistical feature extraction step: From the distance sequence Extract multidimensional statistical features to construct a feature vector. ; S5. Hallucination classification step: The feature vector... Input a pre-trained supervised classifier and output hallucination detection results. The supervised classifier is trained based on labeled hallucination data.

[0007] Furthermore, the semantic preservation perturbation operation includes at least one of synonym replacement and sentence-level template rewriting.

[0008] Furthermore, in the semantic preservation perturbation operation, for Chinese queries, a word segmentation-based synonym replacement is adopted, and a bilingual dictionary containing common words and their corresponding synonyms is maintained. The segmented words are then processed with a preset probability. Random replacement is performed; for English queries, synonym replacement based on word form transformation is used; for queries in any language, sentence template-based rewriting is used to embed the core semantics of the query into a preset query template for re-expression.

[0009] Furthermore, the semantic coding model employs multiple languages. Model.

[0010] Furthermore, distance sequence Specifically: , .

[0011] Furthermore, the feature vector Including the mean ,variance kurtosis Minimum distance Maximum distance Range Dispersion within the perturbation Coefficient of variation and skewness .

[0012] Furthermore, the mean This characterizes the overall degree of semantic drift. This is a function for taking the average value; variance Characterizes the degree of drift fluctuation. To take the variance function; Kudo Characterizes the shape of distance distribution. It is the excess kurtosis function; minimum distance Representing minimal semantic drift, The function is for finding the minimum value; Maximum distance Represents the maximum semantic drift. This is a function to find the maximum value. Range Characterizes the width of the distribution; Dispersion within the perturbation , where is the average cosine distance within the perturbation embedding vector set; coefficient of variation , representing the degree of normalized volatility, where, Cosine distance sequence standard deviation Distance sequence The mean, These are extremely small positive numbers, used to prevent the denominator from being zero; Skewness Characterizing the symmetry of distance distribution, It is a third-order normalized moment function.

[0013] This invention also provides a large language model hallucination detection system based on semantic drift distribution perception, the system comprising: The unit that performs the perturbation generation step: receives the original query. and large language model response The input pair for the original query Perform semantic preservation perturbation operations to generate A set of perturbation queries ; The unit performing the embedding encoding step: using a pre-trained semantic encoding model to process the response r and the perturbation query set. Perform vector encoding separately to obtain the response embedding vector. and perturbation embedding vector sequence ; The unit performing the distance sequence calculation step: calculates the response embedding vector one by one. With each perturbation embedding vector The cosine distances between them form a distance sequence. ; The unit for performing the multidimensional statistical feature extraction step: from the distance sequence Extract multidimensional statistical features to construct a feature vector. ; The unit for performing the hallucination classification step: The feature vector Input a pre-trained supervised classifier and output hallucination detection results. The supervised classifier is trained based on labeled hallucination data.

[0014] The beneficial effects of the method are as follows: (1) Complete black-box detection: The response of the large language model is obtained only through standard API calls. There is no need to access the internal weights, temperature parameters or activation values ​​of the model. It is applicable to all forms of large language model services.

[0015] (2) Multidimensional feature representation: Innovatively extracts 9-dimensional statistical feature vectors from the cosine distance sequence, covering mean, variance, kurtosis, extreme value, range, intraclass dispersion, coefficient of variation and skewness. Compared with single measurement methods, it significantly improves the discriminative ability of features. The Cohen's d effect size of the main features all exceed 1.2, which belongs to the large effect level.

[0016] (3) High supervised classification accuracy: A supervised classifier is used to replace the unsupervised threshold judgment, achieving detection performance of AUC-ROC: 0.8590 and Macro-F1: 0.7825 on real labeled datasets.

[0017] (4) Low computational overhead: The method of the present invention has a single sample detection latency of about 38.4 milliseconds in a CPU environment, which meets the real-time requirements of online deployment and does not require GPU acceleration. Attached Figure Description

[0018] Figure 1 This is a flowchart of the hallucination detection method based on semantic drift distribution perception of a large language model in an embodiment of the present invention; Figure 2 The ROC curve in this embodiment of the invention is shown, with the horizontal axis representing the false positive rate (FPR) and the vertical axis representing the true positive rate (TPR). The area under the curve is AUC = 0.8590. Figure 3 This is a 9-dimensional feature importance ranking diagram in an embodiment of the present invention, showing the gain importance score of each feature by the XGBoost classifier. Detailed Implementation

[0019] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0020] Example 1 like Figure 1 The diagram shown is a flowchart of the method described in this invention. Figure 1 The method of this invention will be described in detail below. The hallucination detection method based on a large language model with semantic drift distribution perception specifically includes the following steps: (1) Disturbance generation steps: Receive containing the original query and large language model response The input pair for the original query Perform semantic preservation perturbation operations to generate A set of perturbation queries The semantic preservation perturbation operation includes at least one of synonym replacement and sentence-level template rewriting; In the semantic preservation perturbation operation, for Chinese queries, a word segmentation-based synonym replacement is adopted. A bilingual dictionary containing common words and their corresponding synonyms is maintained, and the segmented words are processed with a preset probability. Random word replacement is performed, using a conventional data augmentation technique in natural language processing (Random Word Replacement), which is existing technology. For English queries, synonym replacement based on lemmatization is used. This lemmatization-based synonym replacement combines lemmatization with WordNet Synonym Substitution, also existing technology in natural language processing. For queries in any language, sentence template-based rewriting is used, specifically template-based paraphrase generation, also existing technology in natural language processing. The core semantics of the query are then embedded into a preset query template for re-expression.

[0021] Word segmentation is a fundamental technology in Natural Language Processing (NLP), referring to the algorithm's automatic segmentation of continuous Chinese text into word sequences. It is a prerequisite step for all Chinese AI functions.

[0022] (2) Embedding encoding step: Using a pre-trained semantic encoding model to encode the response r and the n perturbation query sets Perform vector encoding separately to obtain the response embedding vector. and perturbation embedding vector sequence The pre-trained semantic encoding model is independent of the large language model; the obtained response embedding vector is subjected to L2 normalization, so that the cosine distance is equivalent to 1 minus the inner product of the normalized vector, thereby reducing the computational complexity.

[0023] Pre-trained semantic encoding models for multilingual languages The model outputs a dense vector representation of fixed dimensions, makes no assumptions about the large language model to be detected, and achieves complete black-box detection. (3) Steps for calculating the distance sequence: Calculate the response embedding vector one by one With each perturbation embedding vector The cosine distances between them form a distance sequence: , ; (4) Steps for extracting multidimensional statistical features: From the distance sequence Extract the following multidimensional statistical features to form a feature vector. : mean This characterizes the overall degree of semantic drift. This is a function for taking the average value; variance This characterizes the degree of drift fluctuation; To take the variance function; Kudo The distance distribution shape is represented by kurtosis(), which is the excess kurtosis function used to measure the steepness of the probability distribution and the characteristics of the extreme values ​​at the tail. The larger the value, the heavier the tail and the more extreme values ​​there are.

[0024] minimum distance This represents the minimum semantic drift; The function is for finding the minimum value; Maximum distance This represents the maximum semantic drift; Range Characterizes the width of the distribution; Dispersion within the perturbation , where is the average cosine distance within the perturbation embedding vector set; coefficient of variation This characterizes the normalized volatility; where std(D) is the standard deviation of the cosine distance sequence D. Let D be the mean of the sequence. For very small positive numbers (e.g.) This is used to prevent the denominator from being zero, i.e., the mean tends to zero when all perturbation responses have completely identical semantics.

[0025] Skewness , characterizing the symmetry of the distance distribution. It is a third-order normalized moment function that characterizes the symmetry of the distance distribution. Positive values ​​indicate that the distribution is right-skewed (there are many large cosine distances), and negative values ​​indicate left-skewed.

[0026] (5) Hallucination classification steps: The feature vector Input a pre-trained supervised classifier and output hallucination detection results. The supervised classifier is trained based on labeled hallucination data.

[0027] The supervised classifier is a gradient boosting decision tree classifier, trained on a labeled illusion dataset, which supports the output detection of arbitrarily large language models without modifying or accessing the internal parameters of the large language model being detected.

[0028] The method further includes a feature importance analysis step: using the supervised classifier to output importance scores for each dimension of features, where the minimum distance... and mean The features with the highest importance scores, reaching 0.3666 and 0.1198 respectively, are the dominant discriminant features; the feature importance analysis helps to interpret the hallucination detection decision.

[0029] Example 2 This embodiment illustrates the specific application of the method in Embodiment 1.

[0030] Case 1: English QA Hallucination Detection Experiment This case study validates the effectiveness of the method of the present invention on a subset of the QA dataset of the HaluEval public benchmark dataset.

[0031] 1.1 Experimental Environment and Data Experimental environment: AMD Ryzen 9 5950X processor, Python 3.12, PyTorch 2.11.0.

[0032] Experimental dataset: HaluEval (pminervini / HaluEval, QA subset), with a total of 10,000 samples, including hallucination labels ("yes" / "no"). 2,000 samples were evenly sampled from it (1,000 non-hallucination samples and 1,000 hallucination samples).

[0033] Semantic encoding model: paraphrase-multilingual-MiniLM-L12-v2 (384-dimensional, multilingual), model parameters approximately 117MB, output 384-dimensional dense vectors, supports 50+ languages.

[0034] Perturbation parameters: Each query generates n=5 semantic perturbation variants, and the perturbation methods include synonym replacement (replacement probability p=0.25) and sentence template rewriting.

[0035] 1.2 Feature Distribution Analysis Table 1 shows the statistical differences of the 9-dimensional feature vector of the present invention between hallucination samples and non-hallucination samples, and the feature discrimination ability is quantified by Cohen's d effect size and two-sample t test.

[0036] Table 1: Discriminative power analysis of 9-dimensional feature vectors

[0037] Note: *** indicates p<0.001, ** indicates p<0.01, ns indicates no significant difference; Cohen's d>0.8 indicates a large effect, 0.5~0.8 indicates a moderate effect.

[0038] As shown in Table 1, the mean μ (d= 1.6363), minimum distance (d= 1.7463) and maximum distance (d= The large effect of the coefficient of variation (CV) (d=1.2798) indicates that the response distance of non-phantom responses to semantic perturbations is generally greater than that of phantom responses. The physical significance of this phenomenon is that the semantic association between phantom responses and the original query is looser; when the query is semantically perturbed, the relative position of the phantom response changes less, and the cosine distance is even smaller. The large effect of the coefficient of variation (CV) (d=1.066) suggests that phantom responses exhibit greater relative fluctuations under different perturbations.

[0039] 1.3 Classification Performance Evaluation Table 2 shows the classification performance metrics of the method of this invention on the test set (400 records). The dataset was split into 80% training set (1600 records) and 20% test set (400 records), and a stratified random split was used to maintain the class ratio.

[0040] Table 2: Classification performance indicators of the method of the present invention (test set, n=400)

[0041] Note: The above data was obtained through testing on a real-world labeled dataset (a subset of HaluEval QA). Table 3: Confusion Matrix (Test Set, n=400)

[0042] As shown in Tables 2 and 3, the AUC-ROC of the method of the present invention reaches 0.8590 under completely black-box conditions. Figure 2As shown, the hallucination detection F1 score reached 0.7852, and the single-sample detection latency was only 38.4 milliseconds, verifying that the method of this invention meets the practical deployment requirements in both effectiveness and real-time performance. In the 5-fold cross-validation, the AUC-ROC mean was 0.8650, and the standard deviation was only 0.0104, indicating that the method has good generalization stability.

[0043] Case 2: Chinese Hallucination Detection Scenario This case illustrates the application of the method of the present invention in the scenario of illusion detection using a large Chinese language model.

[0044] Special configuration for Chinese processing: (1) In the perturbation generation step, after segmenting the Chinese input, synonym replacement is performed, and the replacement probability is set to 0.30; (2) Maintain a dictionary containing commonly used Chinese words and their synonyms, covering major word classes such as verbs, nouns, and adjectives; (3) Use Chinese query templates for sentence-level template rewriting to ensure that the rewritten perturbation sentences conform to Chinese grammatical habits; (4) Use the paraphrase-multilingual-MiniLM-L12-v2 model for semantic encoding. This model natively supports Chinese encoding and does not require additional adaptation.

[0045] The detection process for Chinese and English queries is completely identical. The language detection module automatically determines the input language (if the proportion of Chinese characters exceeds 20%, it is determined to be Chinese) and selects the corresponding perturbation strategy. This case verifies the multilingual versatility of the method of this invention.

[0046] Case Study 3: Deployment of Online Testing Services This case study describes a deployment scheme for the method of the present invention as an online detection service.

[0047] Deployment Architecture: The method of this invention is encapsulated as an independent detection microservice, accepting standard HTTP requests. The input format is a JSON object {query: "original query", response: "large language model response"}, and the output format is {label:0 / 1, confidence: 0~1, feature_vector: [f1,...,f9]}. Here, label is the hallucination detection label, where 0 indicates no hallucination and 1 indicates hallucination; confidence is the classification confidence, ranging from [0,1], with higher values ​​indicating greater confidence in the classifier's judgment; feature_vector is a nine-dimensional statistical feature vector [f1,...,f9], corresponding to mean, variance, kurtosis, minimum, maximum, range, within-group standard deviation, coefficient of variation, and skewness, used for audit tracing.

[0048] Performance metrics: In a CPU environment (AMD Ryzen 9 5950X), the processing time for a single detection request is approximately 38.4 milliseconds, of which SBERT encoding takes approximately 35 milliseconds, XGBoost inference takes approximately 0.1 milliseconds, and the remainder is for preprocessing and post-processing. With 16 cores running concurrently, the system's theoretical throughput exceeds 400 requests / second, meeting the real-time detection requirements of a medium-sized production environment.

[0049] Resource consumption: The runtime memory consumption of the method of this invention is about 500MB (mainly for SBERT model weights), and disk storage is about 200MB. It does not require a GPU, can be deployed on a standard cloud computing instance, and has low operating costs.

[0050] The results of the feature importance analysis are shown in Table 4: Table 4: Ranking of Feature Importance in XGBoost Classifiers (Gain Metric)

[0051] Figure 3 This is a 9-dimensional feature importance ranking chart, showing the XGBoost classifier's gain importance score for each feature.

Claims

1. A hallucination detection method based on a large language model with semantic drift distribution perception, characterized in that, The method includes the following steps: S1, Disturbance Generation Step: Receive data containing the original query. and large language model response The input pair for the original query Perform semantic preservation perturbation operations to generate A set of perturbation queries ; S2, Embedding Encoding Step: Using a pre-trained semantic coding model, the response r and the perturbation query set are processed. Perform vector encoding separately to obtain the response embedding vector. and perturbation embedding vector sequence ; S3. Distance sequence calculation steps: Calculate the response embedding vector one by one. With each perturbation embedding vector The cosine distances between them form a distance sequence. ; S4. Multidimensional statistical feature extraction step: From the distance sequence Extract multidimensional statistical features to construct a feature vector. ; Feature vector Including the mean ,variance kurtosis Minimum distance Maximum distance Range Coefficient of variation and skewness ; S5. Hallucination classification step: The feature vector... Input a pre-trained supervised classifier and output hallucination detection results. The supervised classifier is trained based on labeled hallucination data.

2. The hallucination detection method based on a large language model with semantic drift distribution perception according to claim 1, characterized in that, The semantic preservation perturbation operation includes at least one of synonym replacement and sentence-level template rewriting.

3. The hallucination detection method based on a large language model with semantic drift distribution perception according to claim 2, characterized in that, In the semantic preservation perturbation operation, for Chinese queries, a word segmentation-based synonym replacement is adopted. A bilingual dictionary containing common words and their corresponding synonyms is maintained, and the segmented words are processed with a preset probability. Random replacement is performed; for English queries, synonym replacement based on word form transformation is used; for queries in any language, sentence template-based rewriting is used to embed the core semantics of the query into a preset query template for re-expression.

4. The hallucination detection method based on a large language model with semantic drift distribution perception according to claim 3, characterized in that, Semantic coding model adopts multiple languages Model.

5. The hallucination detection method based on a large language model with semantic drift distribution perception according to claim 4, characterized in that, Distance sequence Specifically: , , Indicates the first perturbation embedding vector With response embedding vector The cosine distance between them.

6. The hallucination detection method based on a large language model with semantic drift distribution perception according to claim 5, characterized in that, mean This characterizes the overall degree of semantic drift. This is a function for taking the average value; variance Characterizes the degree of drift fluctuation. To take the variance function; Kudo Characterizes the shape of distance distribution. It is the excess kurtosis function; minimum distance Representing minimal semantic drift, The function is for finding the minimum value; Maximum distance Represents the maximum semantic drift. This is a function to find the maximum value. Range Characterizes the distribution width; coefficient of variation , characterizing the degree of normalized volatility, where, Cosine distance sequence standard deviation Distance sequence The mean, These are extremely small positive numbers, used to prevent the denominator from being zero; Skewness Characterizing the symmetry of distance distribution, It is a third-order normalized moment function.

7. A large language model-based hallucination detection system based on semantic drift distribution perception, characterized in that, The system includes: The unit that performs the perturbation generation step: receives the original query. and large language model response The input pair for the original query Perform semantic preservation perturbation operations to generate A set of perturbation queries ; The unit performing the embedding encoding step: using a pre-trained semantic encoding model to process the response r and the perturbation query set. Perform vector encoding separately to obtain the response embedding vector. and perturbation embedding vector sequence ; The unit performing the distance sequence calculation step: calculates the response embedding vector one by one. With each perturbation embedding vector The cosine distances between them form a distance sequence. ; The unit for performing the multidimensional statistical feature extraction step: from the distance sequence Extract multidimensional statistical features to construct a feature vector. ; Feature vector Including the mean ,variance kurtosis Minimum distance Maximum distance Range Coefficient of variation and skewness ; The unit for performing the hallucination classification step: The feature vector Input a pre-trained supervised classifier and output hallucination detection results. The supervised classifier is trained based on labeled hallucination data.

Citation Information

Patent Citations

  • Big language model illusion detection method

    CN120973939A

  • Knowledge graph-based big language model illusion detection method and device, and medium

    CN121809451A

  • Question and answer processing method, device and equipment, computer readable storage medium and computer program product

    CN121958509A