An illusion recognition method and system, an electronic device, and a readable storage medium

CN122693771APending Publication Date: 2026-09-04HUNAN ZHITONG STAR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611178042.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-05
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

但基于增量监督微调(Incremental SFT)的幻觉修正方法不能显式告诉模型“历史错误在哪”,容易遗忘原正确能力,错误模式反复出现;仅拟合输出表面,特征层偏差未消除

Benefits of technology

本申请通过根据包含对话输入序列、模型历史幻觉输出序列以及人工修正输出序列三元组的幻觉识别样本,构建幻觉识别样本数据集;根据对话输入序列,构建标准质检对话数据集;将对话输入序列和模型历史幻觉输出序列,构建为幻觉侧输入;将对话输入序列和人工修正输出序列,构建为正确侧输入;构建包含线性投隐层、非线性编码器、层归一化以及L2归一化的学生幻觉纠偏分支和教师幻觉纠偏分支,学生幻觉纠偏分支和教师幻觉纠偏分支的结构一样;采用大语言模型对幻觉侧输入进行处理,得到第一中间向量;将第一中间向量输入至学生幻觉纠偏分支,得到幻觉表征;采用大语言模型对正确侧输入进行处理,得到第二中间向量;将第二中间向量输入至学生幻觉纠偏分支,得到学生正确表征;将第二中间向量输入至教师幻觉纠偏分支,得到教师锚点表征;基于正确侧输入和人工修正输出序列,构建标准生成损失;基于幻觉表征、学生正确表征以及教师锚点表征,构建温度加权MSE对齐损失、自适应边界余弦排斥损失以及正交投影惩罚损失;基于模型历史幻觉输出序列中的幻觉类型,构建排序正则化损失;对标准生成损失、温度加权MSE对齐损失、自适应边界余弦排斥损失、排序正则化损失以及正交投影惩罚损失进行加权求和,得到总损失函数;采用标准质检对话数据集和标准生成损失,迭代更新大语言模型中的参数;采用幻觉识别样本数据集和总损失函数,迭代更新大语言模型中的参数、学生幻觉纠偏分支以及门控网络;通过指数移动平均方法迭代更新教师幻觉纠偏分支;完成迭代更新后,得到训练好的大语言模型;通过训练好的大语言模型进行对话质检幻觉识别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122693771A_ABST
    Figure CN122693771A_ABST
Patent Text Reader

Abstract

The application discloses an illusion identification method and system, an electronic device and a readable storage medium. The method adopts a large language model, a student illusion correction branch and a teacher illusion correction branch, extracts illusion representation, student correct representation and teacher anchor point representation, performs weighted summation on standard generation loss, temperature weighted MSE alignment loss, adaptive boundary cosine repulsion loss, sorting regularization loss and orthogonal projection penalty loss to obtain a total loss function, uses standard quality inspection dialogue data set and standard generation loss to iteratively update parameters in the large language model, uses illusion identification sample data set and the total loss function to iteratively update the parameters in the large language model, the student illusion correction branch and the gating network, iteratively updates the teacher illusion correction branch through the exponential moving average method, and completes the iteration and update. After the iteration and update, the large language model trained is used for dialogue quality inspection illusion identification. The application can improve the illusion identification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, electronic device and readable storage medium for hallucination recognition. Background Technology

[0002] Existing related technologies include illusion correction methods based on incremental supervised fine-tuning (SFT), reinforcement learning methods based on human feedback, and contrastive learning methods. However, illusion correction methods based on incremental supervised fine-tuning (SFT) cannot explicitly tell the model "where the historical errors are," easily forgetting the original correct abilities, and error patterns recurring; they only fit the output surface, and feature layer biases are not eliminated. Contrastive learning methods heavily rely on large batches, hard example mining, and boundary hyperparameters, which easily conflict with generation tasks in large language models (LLM), leading to training divergence; furthermore, the negative examples in contrastive learning are mostly artificially constructed perturbations, not personalized real errors of the model, resulting in weak specificity.

[0003] Reinforcement learning methods based on human feedback are highly unstable during training, with a high risk of reward cheating and pattern collapse. Their engineering systems are complex, requiring the construction of environments, reward models, and sampling pipelines, resulting in significant implementation costs. Some methods use uniform weights and magnitudes for all preference pairs in their preference loss, regardless of whether the error is a fatal misjudgment or a minor misclassification; the optimization effort is exactly the same, and it only affects the output probability, lacking direct geometric constraints on the hidden layer representation space. This prevents the explicit elimination of collinear components along the correct direction in erroneous representations, leading to easy recurrence of error patterns in out-of-distribution scenarios. The inability to eradicate residual and recurring erroneous representations makes it impossible to accurately construct and distinguish rewards for different types of errors, failing to meet the compliance requirements of e-commerce quality inspection. E-commerce quality inspection checks whether responses in dialogues based on large language models are accurate, compliant, and consistent with facts (such as product information and order status).

[0004] Therefore, existing technologies have relatively low accuracy in identifying large language model illusions, and there are still serious large language model illusions, resulting in low accuracy in dialogue quality inspection. Summary of the Invention

[0005] This application aims to provide a method, system, electronic device, and readable storage medium for hallucination recognition, which can improve the accuracy of hallucination recognition and the precision of dialogue quality inspection.

[0006] In a first aspect, embodiments of this application provide a method for hallucination recognition, the method comprising: A hallucination recognition sample dataset is constructed based on hallucination recognition samples containing triplet pairs of dialogue input sequences, model historical hallucination output sequences, and manually corrected output sequences; a standard quality inspection dialogue dataset is constructed based on the dialogue input sequences. The dialogue input sequence and the model history hallucination output sequence are used to construct the hallucination-side input; the dialogue input sequence and the manually corrected output sequence are used to construct the correct-side input. Construct student hallucination correction branches and teacher hallucination correction branches that include a linear hidden layer, a nonlinear encoder, layer normalization, and L2 normalization. The student hallucination correction branches and the teacher hallucination correction branches have the same structure. The hallucination input is processed using a large language model to obtain a first intermediate vector; the first intermediate vector is then input into the student hallucination correction branch to obtain a hallucination representation. The correct side input is processed using the large language model to obtain a second intermediate vector; the second intermediate vector is input into the student hallucination correction branch to obtain the student's correct representation; the second intermediate vector is input into the teacher hallucination correction branch to obtain the teacher's anchor representation. Based on the correct side input and the manually corrected output sequence, a standard generation loss is constructed; based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation, a temperature-weighted MSE alignment loss, an adaptive boundary cosine repulsion loss, and an orthogonal projection penalty loss are constructed; based on the hallucination type in the model's historical hallucination output sequence, a ranking regularization loss is constructed; the standard generation loss, the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, the ranking regularization loss, and the orthogonal projection penalty loss are weighted and summed to obtain the total loss function; The parameters in the large language model are iteratively updated using the standard quality inspection dialogue dataset and the standard generation loss; the parameters in the large language model, the student hallucination correction branch, and the gating network are iteratively updated using the hallucination recognition sample dataset and the total loss function; the teacher hallucination correction branch is iteratively updated using the exponential moving average method; after completing the iterative update, a trained large language model is obtained; and the trained large language model is used for dialogue quality inspection hallucination recognition.

[0007] In some implementations, the step of processing the hallucination-side input using a large language model to obtain a first intermediate vector includes: Building a large language model based on multiple Transformer layers; The input to the hallucination side is processed using a large language model, and the output vector of the penultimate layer in the multi-layer Transformer is used as the first intermediate vector.

[0008] In some implementations, the construction of temperature-weighted MSE alignment loss, adaptive boundary cosine repulsion loss, and orthogonal projection penalty loss based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation includes: Based on the teacher anchor point representation and the student correct representation, a temperature-weighted MSE alignment loss is constructed. Based on the illusion representation and the teacher anchor representation, an adaptive boundary cosine repulsion loss and an orthogonal projection penalty loss are constructed.

[0009] In some implementations, constructing the temperature-weighted MSE alignment loss based on the teacher anchor representation and the student correct representation includes: The teacher anchor representation is input into the gating network to obtain the dimension importance vector; The difference between the teacher anchor point representation and the student correct representation is calculated to obtain the difference result; Based on the dimensionality importance vector and the difference results, a temperature-weighted MSE alignment loss is constructed.

[0010] In some implementations, constructing an adaptive boundary cosine repulsion loss based on the hallucination representation and the teacher anchor representation includes: Assign learnable embedding vectors to the hallucination types in the historical hallucination output sequence of the model; The embedded vector is processed by linear projection and activation function to obtain adaptive boundary values; Calculate the cosine similarity between the hallucination representation and the teacher anchor representation; Based on the adaptive boundary value and the cosine similarity, an adaptive boundary cosine repulsion loss is constructed.

[0011] In some implementations, constructing the orthogonal projection penalty loss based on the hallucination representation and the teacher anchor representation includes: Calculate the squared magnitude of the projection vector between the hallucination representation and the teacher anchor representation, and construct the orthogonal projection penalty loss.

[0012] In some implementations, constructing the ranking regularization loss based on the hallucination types in the model's historical hallucination output sequence includes: Based on the hallucination types in the historical hallucination output sequence of the model, construct an ordered set of pairs that satisfy a strict partial order relation; Based on the ordered set of pairs and the adaptive boundary value corresponding to each illusion type, a sorting regularization loss is constructed.

[0013] Secondly, embodiments of this application also provide a hallucination recognition system, the system comprising: The dataset construction module is used to construct an illusion recognition sample dataset based on illusion recognition samples containing triplets of dialogue input sequences, model historical illusion output sequences, and manually corrected output sequences; and to construct a standard quality inspection dialogue dataset based on the dialogue input sequences. An input sequence construction module is used to construct the illusion-side input from the dialogue input sequence and the model history illusion output sequence; and to construct the correct-side input from the dialogue input sequence and the manually corrected output sequence. The bias correction branch construction module is used to construct student hallucination bias correction branches and teacher hallucination bias correction branches, which include linear hidden layers, nonlinear encoders, layer normalization, and L2 normalization. The student hallucination bias correction branches and the teacher hallucination bias correction branches have the same structure. The first data processing module is used to process the hallucination side input using a large language model to obtain a first intermediate vector; and input the first intermediate vector into the student hallucination correction branch to obtain a hallucination representation. The second data processing module is used to process the correct side input using the large language model to obtain a second intermediate vector; input the second intermediate vector into the student hallucination correction branch to obtain the student correct representation; and input the second intermediate vector into the teacher hallucination correction branch to obtain the teacher anchor point representation. The loss function construction module is used to construct a standard generation loss based on the correct side input and the manually corrected output sequence; to construct a temperature-weighted MSE alignment loss, an adaptive boundary cosine repulsion loss, and an orthogonal projection penalty loss based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation; to construct a ranking regularization loss based on the hallucination type in the model's historical hallucination output sequence; and to perform a weighted summation of the standard generation loss, the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, the ranking regularization loss, and the orthogonal projection penalty loss to obtain the total loss function. The model illusion recognition module is used to iteratively update the parameters in the large language model using the standard quality inspection dialogue dataset and the standard generation loss; iteratively update the parameters in the large language model, the student illusion correction branch, and the gating network using the illusion recognition sample dataset and the total loss function; iteratively update the teacher illusion correction branch using the exponential moving average method; after completing the iterative update, a trained large language model is obtained; and illusion recognition for dialogue quality inspection is performed using the trained large language model.

[0014] Thirdly, embodiments of this application also provide an electronic device, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform an illusion recognition method as described above.

[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform an illusion recognition method as described above.

[0016] Compared with the prior art, this application has the following beneficial effects: This application constructs a hallucination recognition sample dataset based on hallucination recognition samples containing triples of dialogue input sequences, model historical hallucination output sequences, and manually corrected output sequences; constructs a standard quality inspection dialogue dataset based on dialogue input sequences; constructs the hallucination-side input using the dialogue input sequences and model historical hallucination output sequences; constructs the correct-side input using the dialogue input sequences and manually corrected output sequences; constructs student hallucination correction branches and teacher hallucination correction branches, each with the same structure, including linear hidden layers, nonlinear encoders, layer normalization, and L2 normalization; processes the hallucination-side input using a large language model to obtain a first intermediate vector; inputs the first intermediate vector into the student hallucination correction branch to obtain a hallucination representation; processes the correct-side input using a large language model to obtain a second intermediate vector; inputs the second intermediate vector into the student hallucination correction branch to obtain a correct student representation; and inputs the second intermediate vector into the teacher hallucination correction branch. The process involves obtaining teacher anchor representations; constructing a standard generation loss based on correct input and manually corrected output sequences; constructing temperature-weighted MSE alignment loss, adaptive boundary cosine repulsion loss, and orthogonal projection penalty loss based on hallucination representations, correct student representations, and teacher anchor representations; constructing ranking regularization loss based on hallucination types in the model's historical hallucination output sequences; weighting and summing the standard generation loss, temperature-weighted MSE alignment loss, adaptive boundary cosine repulsion loss, ranking regularization loss, and orthogonal projection penalty loss to obtain the total loss function; iteratively updating the parameters in the large language model using a standard quality-checking dialogue dataset and the standard generation loss; iteratively updating the parameters, student hallucination correction branch, and gating network in the large language model using a hallucination recognition sample dataset and the total loss function; iteratively updating the teacher hallucination correction branch using an exponential moving average method; obtaining the trained large language model after iterative updates; and performing hallucination recognition for dialogue quality checks using the trained large language model.

[0017] Thus, constructing the hallucination recognition sample dataset and the standard quality inspection dialogue dataset lays a solid data foundation for subsequent model training. Hallucination representations, student correct representations, and teacher anchor representations are extracted through hallucination-side input and correct-side input, providing a good data foundation for constructing the loss function and training the model. By constructing teacher hallucination correction branches and student hallucination correction branches based on a shared large language model, feature layer bias in the incremental supervised fine-tuning method of constrained branches is eliminated. The teacher hallucination correction branch is updated from the student hallucination correction branch through exponential moving average, providing stable and evolving teacher anchor representations for correct representations and solving the problem of random initialization of teacher anchor representations. By constructing standard generation loss, temperature-weighted MSE alignment loss, adaptive boundary cosine repulsion loss, ranking regularization loss, and orthogonal projection penalty loss, and then training the large language model through the total loss function, the accuracy of hallucination recognition and the precision of dialogue quality inspection can be improved by using the trained large language model for hallucination recognition in dialogue quality inspection. Attached Figure Description

[0018] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart illustrating an embodiment of the hallucination recognition method provided in this application; Figure 2 This is a schematic diagram of the representation extraction process in the preferred embodiment of the hallucination recognition method provided in this application; Figure 3 This is a schematic diagram of the structure of an embodiment of the hallucination recognition system provided in this application; Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation

[0019] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0020] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.

[0021] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0022] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0023] To address the problem of severe large language model illusion in related technologies, which leads to low accuracy in dialogue quality inspection, this application proposes an illusion recognition method, system, electronic device, and readable storage medium.

[0024] Reference Figure 1 This application provides a schematic flowchart of a hallucination recognition method. This hallucination recognition method is applied to an electronic device, which may be a server or a mobile terminal, etc. Figure 1 As shown, the hallucination recognition method may include the following steps S101 to S107.

[0025] Step S101: Construct a hallucination recognition sample dataset based on hallucination recognition samples containing triplet pairs of dialogue input sequences, model historical hallucination output sequences, and manually corrected output sequences; construct a standard quality inspection dialogue dataset based on dialogue input sequences.

[0026] In this step, each hallucination recognition sample consists of a triplet, which includes the dialogue input sequence, the model's historical hallucination output sequence, and the manually corrected output sequence. All hallucination recognition samples are combined into a hallucination recognition sample dataset. The dialogue input sequence is standard question-and-answer data, and all dialogue input sequences are combined into a standard quality control dialogue dataset.

[0027] Step S102: Construct the illusion-side input by combining the dialogue input sequence and the model history illusion output sequence; construct the correct side input by combining the dialogue input sequence and the manually corrected output sequence.

[0028] In this step, the dialogue input sequence and the model's historical hallucination output sequence are concatenated to construct the hallucination-side input; the dialogue input sequence and the manually corrected output sequence are concatenated to construct the correct side input.

[0029] Step S103: Construct student hallucination correction branches and teacher hallucination correction branches, which include linear hidden layers, nonlinear encoders, layer normalization, and L2 normalization. The structures of student hallucination correction branches and teacher hallucination correction branches are the same.

[0030] In this step, the hallucination correction branch network structure is defined, including linear hidden layers, a nonlinear encoder, layer normalization, and L2 normalization. The nonlinear encoder includes a dimension-reduced linear layer, a ReLU activation function, and a dimension-upgrading linear layer. The structures of the student hallucination correction branch and the teacher hallucination correction branch are the same, both employing the hallucination correction branch network structure.

[0031] Step S104: Process the hallucination side input using a large language model to obtain the first intermediate vector; input the first intermediate vector into the student hallucination correction branch to obtain the hallucination representation.

[0032] In this step, a large language model is constructed based on multiple Transformer layers. This model is then used to process the hallucination-side input, with the output vector of the penultimate layer of the Transformer layers serving as the first intermediate vector. This first intermediate vector is then input into the student's hallucination correction branch to obtain the hallucination representation.

[0033] Step S105: Process the correct side input using a large language model to obtain the second intermediate vector; input the second intermediate vector into the student hallucination correction branch to obtain the student's correct representation; input the second intermediate vector into the teacher hallucination correction branch to obtain the teacher's anchor point representation.

[0034] In this step, a large language model is used to process the correct side input, and the output vector of the penultimate layer in the multi-layer Transformer is used as the second intermediate vector. Then, the second intermediate vector is input into the student hallucination correction branch to obtain the student's correct representation; the second intermediate vector is input into the teacher hallucination correction branch to obtain the teacher's anchor representation.

[0035] Step S106: Based on the correct side input and manually corrected output sequence, construct the standard generation loss; based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation, construct the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, and the orthogonal projection penalty loss; based on the hallucination type in the model's historical hallucination output sequence, construct the ranking regularization loss; and perform a weighted summation of the standard generation loss, the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, the ranking regularization loss, and the orthogonal projection penalty loss to obtain the total loss function.

[0036] In this step, a temperature-weighted MSE alignment loss is constructed based on teacher anchor representations and student correct representations; an adaptive boundary cosine repulsion loss and an orthogonal projection penalty loss are constructed based on illusion representations and teacher anchor representations.

[0037] The above-mentioned standard generation loss is constructed based on the correct side input and manually corrected output sequence, including: ; in, Indicates the standard generation loss. This indicates the number of tokens in the manually corrected output sequence. A token is the smallest unit of text in a Large Language Model (LLM). This represents the probability distribution derived through LLM. Indicates the first digit in the manually corrected output sequence. One token, Indicates the position in the manually corrected output sequence. The token sequence up to this point.

[0038] The above-mentioned temperature-weighted MSE alignment loss, based on teacher anchor representations and student correct representations, includes: The teacher anchor representation is input into a gating network to obtain a dimension importance vector; the difference between the teacher anchor representation and the student's correct representation is calculated to obtain the difference result; based on the dimension importance vector and the difference result, a temperature-weighted MSE alignment loss is constructed, including: ; in, This represents the temperature-weighted MSE alignment loss. The dimensions representing teacher anchor representations and student correct representations. Indicates the first A dimensional importance vector, Indicates the first Each student correctly represented, Indicates the first Teacher anchor point representation.

[0039] The above-mentioned adaptive boundary cosine repulsion loss, based on hallucination representation and teacher anchor representation, includes: Assign learnable embedding vectors to hallucination types in the model's historical hallucination output sequence; process the embedding vectors through linear projection and activation functions to obtain adaptive boundary values; calculate the cosine similarity between hallucination representations and teacher anchor representations; construct an adaptive boundary cosine repulsion loss based on the adaptive boundary values ​​and cosine similarity, including: ; in, This represents the adaptive boundary cosine repulsion loss. This indicates taking the maximum value. Represents the cosine function. Indicates hallucination manifestation, This represents the teacher anchor point representation. This represents adaptive boundary values.

[0040] The above-mentioned orthogonal projection penalty loss, based on hallucination representations and teacher anchor representations, includes: Calculate the squared magnitude of the projection vector between the hallucination representation and the teacher anchor representation, and construct the orthogonal projection penalty loss, including: ; in, This represents the penalty loss for orthogonal projection. Indicating hallucinations Teacher anchor representation Projected length in the direction, This represents the squared magnitude of the projection vector. This represents the square of the L2 norm.

[0041] The above-mentioned ranking regularization loss is constructed based on the hallucination types in the model's historical hallucination output sequence, including: Based on the hallucination types in the model's historical hallucination output sequence, an ordered set of pairs satisfying a strict partial order relation is constructed; based on the ordered set of pairs and the adaptive boundary value corresponding to each hallucination type, a ranking regularization loss is constructed, including: ; in, This represents the ranking regularization loss. Indicates the first Adaptive boundary values ​​corresponding to each hallucination type Indicates the first Adaptive boundary values ​​corresponding to each hallucination type This represents the minimum interval hyperparameter. It represents an ordered set of pairs.

[0042] Step S107: Using the standard quality inspection dialogue dataset and standard generation loss, iteratively update the parameters in the large language model; using the illusion recognition sample dataset and total loss function, iteratively update the parameters, student illusion correction branch, and gating network in the large language model; iteratively update the teacher illusion correction branch using the exponential moving average method; after completing the iterative update, a trained large language model is obtained; the trained large language model is used for dialogue quality inspection illusion recognition.

[0043] In this step, for the standard quality inspection dialogue dataset The samples in the middle, frozen , and Only calculate And iteratively update the large language model In The parameters. For the illusion recognition sample dataset. Calculate the total loss function from the samples in the dataset. Iterative updates of the large language model In Parameters, Student Hallucination Correction Branch and gating network Teacher Branch The teacher hallucination correction branch is not involved in gradient updates, but is only updated synchronously and iteratively using the exponential moving average (EMA) method. After the iterative update is completed, a well-trained large language model is obtained; the well-trained large language model is then used for dialogue quality inspection and hallucination recognition.

[0044] To facilitate understanding by those skilled in the art, a set of preferred embodiments is provided below: This embodiment provides a dual-track decoupled Large Language Model (LLM) dialogue quality inspection illusion recognition method, which specifically includes the following steps: 1. Input layer and structure layer.

[0045] Suppose that an illusion recognition sample consists of a triple, including: the dialogue input sequence Model historical illusion output sequence and manually corrected output sequence The hallucination samples in the model's historical hallucination output sequence are labeled with hallucination types, including: missed detection, misjudgment, and misclassification. Missed detection hallucination occurs when the model fabricates a judgment that is not in violation, identifying obvious violation-related dialogue errors as compliant. Misjudgment hallucination occurs when the model fabricates a judgment that is in violation, identifying completely compliant dialogue errors as violations. Misclassification hallucination occurs when the model correctly identifies a violation, but fabricates the violation level or category, leading to inaccurate severity judgment. The splicing operation is defined as the hallucination-side input. and correct side input .

[0046] Concatenate the following template to form the input sequence for a Large Language Model (LLM). or : [CLS] Please determine if the following customer service conversation violates any rules, and provide the classification and reason. Conversation: Judgment result: [SEP]. Among them, [CLS] and [SEP] use special tokens already existing in the LLM vocabulary, or use custom separator tokens. After concatenation, they are converted into token sequences by a token segmenter and the necessary padding is added to the maximum batch length. Attention mask is generated synchronously, marking the actual token position as 1 and the padding position as 0.

[0047] Subsequently, a shared Large Language Model (LLM) backbone network was constructed (the Large Language Model is built based on Transformer, and the Transformer architecture includes attention layers and multilayer perceptron (MLP) layers), with the following parameters: (Including LoRA parameters), this embodiment uses a pre-trained large language model base (qwen3.5-4B), extending the scope of the LoRA parameters to the attention layer and MLP layer, with the LoRA rank set to 16; its hidden layer output function is denoted as... Take the last latent vector of the second-to-last token from the penultimate layer as the global representation: ; in, This refers to the vector dimension of an LLM, typically 4096. This represents an operation that retrieves the hidden vector, specifically retrieving the vector of the last token. This indicates retrieving the output vector of the second-to-last layer in a multi-layer transformer architecture. To obtain a fixed-dimensional vector representing the entire input sequence, the latent vector of the last token corresponding to the effective length of each sample is retrieved. Specifically, the attention_mask is used to locate the true end position index last_idx of each sample, and the latent vector is extracted. , This represents the index of the sample in each batch of training, and we get... L represents the number of hidden layers in the large language model. The reason for using tail token aggregation instead of average pooling is that LLM is an autoregressive decoder structure, and the tail token (dim=4096 means a 4096-dimensional vector) gathers all the information from the preceding text through the attention mechanism, making it naturally suitable as a global representation. Furthermore, taking the tail token per sample reduces the noise introduced by the padding region.

[0048] Define the illusion correction branch network The network structure includes linear projected hidden layers (Linear, ), nonlinear encoders (including dimension reduction linear layers) ReLU activation function and upgraded linear layer The algorithm performs layer normalization (LayerNorm, 4096) and L2 normalization (L2 Normalize), finally outputting a representation vector (dim=4096). The weight parameters The parameters of the student hallucination correction branch are denoted as trainable parameters. The parameters of the teacher's hallucination correction branch are denoted as The structure of the teacher hallucination correction branch parameters is completely identical to that of the student branch. At the start of training, the weights of the teacher hallucination correction branch are strictly copied from the weights of the student hallucination correction branch: After each training iteration (or N iterations), perform parameter updates: ; in, These are the parameters of the teacher hallucination correction branch after the current step is updated via gradient. This means assigning the value to the right of the arrow to the parameter on the left of the arrow to perform a parameter update. These are the parameters of the teacher hallucination correction branch after the gradient update in the previous step. This represents the parameters of the student hallucination correction branch after the current step is updated by the gradient. It is the momentum coefficient, preferably 0.999.

[0049] when At that time, the teacher's hallucination correction branch only absorbs changes of one-thousandth of the student's hallucination correction branch at each step, thus exhibiting significant inertia, effectively filtering out single-step gradient noise, and providing a high-quality and stable anchor point.

[0050] 2. Feature extraction process.

[0051] Reference Figure 2 Given an illusion recognition sample, forward propagation computes the following three representation vectors: (1) Hallucination representation .

[0052] ; This representational input is hallucinatory side input. The intermediate vector is calculated using a Large Language Model (LLM). The latent vector of the tail token (i.e., the output vector of the second-to-last layer of the multi-layer transformers (i.e., the first intermediate vector)) is then fed into the student hallucination correction branch. Output Used for subsequent calculation of repulsion loss and orthogonal loss .

[0053] (2) Students' correct representation .

[0054] ; This representation input is the correct side input. The intermediate vector is obtained through LLM calculation. The hidden vector of the tail token (i.e., the output vector of the second-to-last layer of the multi-layer transformers (i.e., the second intermediate vector)) is then fed into the student hallucination correction branch. Output Used for subsequent calculation of alignment loss .

[0055] (3) Teacher anchor representation .

[0056] ; This representation input is the correct side input. The intermediate vector is obtained through LLM calculation. The hidden vector of the tail token (i.e., the output vector of the second to last layer of the multi-layer transformers (i.e., the second intermediate vector)) is then fed into the teacher illusion correction branch. Output This is used for subsequent calculation of alignment loss, and the teacher illusion correction branch does not perform gradient updates to ensure anchor point stability. The student illusion correction branch updates the teacher illusion correction branch using EMA (i.e., exponential moving average, which updates weights by weighted sliding).

[0057] 3. Loss function design.

[0058] The total loss function consists of the following five components: ; in, , , as well as . It has the highest weight because directly eliminating erroneous patterns is the primary goal of illusion correction, and its hinge form ensures that it will not over-reject. Secondly, it is used to gather the correct features to the standard, but it should not be too large to avoid limiting the model's generalization. and Minimum, as an auxiliary penalty, provides fine-grained orthogonal correction and boundary sorting.

[0059] (1) Standard generation loss .

[0060] The conventional autoregressive language modeling loss calculates the cross-entropy of the model's prediction of the next token given the dialogue input and context. It remains active throughout training to maintain the general quality control capabilities of the Large Language Model (LLM), including understanding customer service dialogues, mastering violation judgment rules, and generating ratings and justifications. This loss ensures that this embodiment does not cause a degradation in model generation quality or the loss of basic judgment capabilities when performing representation-level hallucination correction on hallucination samples.

[0061] ; in, This represents the probability distribution obtained through LLM, where T represents the number of tokens in the manually corrected output sequence. Indicates the position in the manually corrected output sequence. The token sequence up to now, The correct answer (i.e., the manually corrected output sequence) is the first... A token.

[0062] (2) Temperature-weighted MSE alignment loss for channel importance perception .

[0063] Standard MSE loss treats all feature dimensions equally; however, in semantic representation, different dimensions contribute differently to correct violation judgments. For example, some dimensions represent customer service tone, while others represent the degree of matching of legal clauses. Indiscriminately aligning all dimensions may force noise on unimportant dimensions to be aligned, impairing generalization. Furthermore, students' correct representations may be poor in the early stages of training, and overly rigid alignment can cause gradient oscillations. Therefore, channel-level importance weights and temperature softening are introduced (here, temperature refers to the softening of the gating network, not a temperature coefficient, but the output of the gating network can be considered as softening importance).

[0064] Set up a gated network The structure is a single-layer fully connected plus : ; in, and These represent the weight matrix and bias, respectively. The input to the gated network is the teacher anchor representation. Output dimension importance vector , This represents the sigmoid activation function. Using the sigmoid ensures that the weights of each channel are within the specified range. Within this framework, the network is allowed to adaptively adjust based on the characteristics of the anchor points themselves. Dimension weights that contribute highly to the correct pattern will approach 1, while dimension weights that contribute little or are useless will approach 0.

[0065] Temperature-weighted MSE alignment loss calculation: ; in, Representing dimension, Indicates the first The numerical values ​​of the importance vector in each dimension. Indicates the first The numerical value correctly represented by each student Indicates the first The numerical value represented by each teacher anchor point.

[0066] This design achieves a "soft attention" alignment: biases in important dimensions are amplified (weights close to 1), and the model prioritizes correcting these dimensions; while biases in unimportant dimensions are suppressed (weights close to 0), allowing the model to maintain a degree of flexibility. Differentiated alignment weights are assigned to different dimensions, with more sensitivity to important dimensions, improving alignment efficiency and preventing overfitting of unimportant features. This focuses more on the core elements of illusion correction than uniform MSE.

[0067] (3) Adaptive Boundary Cosine Repulsion Loss Based on Hallucination Harm .

[0068] The original cosine repulsion loss only maximizes the cosine distance between the erroneous representation and the anchor point (equivalent to minimizing similarity), but it doesn't provide a bottom line for the distance. Once the illusory representation has been pushed too far (e.g., the cosine similarity has dropped to 0.2), the loss approaches 0, the gradient vanishes, and it can no longer reinforce the distance. More importantly, different types of illusions have different levels of business harm: false negatives (FN) are more serious than slight classification shifts and should therefore be subject to stronger repulsion. Therefore, this loss introduces an adaptive minimum bound determined by the type of illusion. This ensures that the rejection mechanism is activated only when the similarity is above the boundary, and the boundary value is dynamically adjusted by the task requirements.

[0069] Adaptive boundaries are generated as follows: Label each hallucination sample (i.e., the hallucination sample in the model's historical hallucination output sequence) with the hallucination type. Each type is assigned a learnable embedding vector. The embedding vector is randomly initialized, and then adaptive boundary values ​​are obtained through linear projection and sigmoid. : ; Among them, coefficient This limits the maximum value of the boundary to 0.5. If a certain type of hallucination requires stronger repulsion (such as FN), its embedding can be learned to output a smaller value. (e.g., 0.1) means that the cosine similarity between the hallucination representation and the teacher anchor representation must be below 0.1 to be considered a violation. Mild hallucinations can produce larger outputs. (e.g., 0.4), allowing for the preservation of more similarity. This represents the sigmoid activation function. Represents the weight matrix , Represents the bias vector , Represents the dimension of a vector or matrix. Embedded vectors and network parameters are optimized end-to-end, with boundary values ​​entirely driven by business data.

[0070] The adaptive boundary cosine repulsion loss takes the form of: ; This is a hinge-loss form, where the cosine similarity... Greater than At that time, losses occurred. Gradient-driven Decreasing the distance between students' correct representations and illusory representations leads to them becoming increasingly dissimilar; when It is already less than When the loss is 0, no more unnecessary repulsion is applied, thus avoiding the collapse of the representation space caused by infinite pushing.

[0071] In adaptive boundary design based on hallucination severity, different types of hallucinations should satisfy a clear order of boundary size: missed detection (FN) has the lowest tolerance, therefore the boundary size should be... The minimum should be minimized; the tolerance for false positives (FP) is the highest, and the boundary value is minimized. The maximum should be achievable; the level error (level_err) should be centered. If relying entirely on end-to-end learning, inverted boundary sizes may occur due to noise in the training data, imbalanced samples between illusion types, or differences in optimization dynamics (e.g., overtake This disrupts business priors. Ranking regularization loss, as a structured prior injection method, forces boundary values ​​to maintain the order of severity in the form of soft constraints, while preserving the adaptive capability of data-driven processes.

[0072] Let's rank error types (i.e., hallucination types) by severity, defining the ordered class relation as follows: FN (missed detection) has the highest severity and should have the smallest boundary; level_err (leveling error) has a medium severity; FP (false positive) has the lowest severity and should have the largest boundary. Based on this, we construct an ordered set of pairs that satisfy a strict partial order relation. : ; For each pair ,Require The ranking regularization loss is defined as: ; in, Indicates the first Adaptive boundary values ​​corresponding to each hallucination type Indicates the first The adaptive boundary values ​​corresponding to each illusion type are obtained by scaling / clipping the sigmoid output. The minimum interval hyperparameter is recommended to be 0.1, which is used to prevent the boundaries from being too close together and to ensure that there is enough space to represent the differences in hazard levels.

[0073] (4) Orthogonal projection penalty loss .

[0074] Using only cosine repulsion may produce a false sense of separation; the illusory representation is pushed away along a direction parallel to the correct anchor point. That is, the projected component of the illusory representation in the anchor point direction remains large, only the overall angle increases. This representation still retains components similar to the correct decision logic, and the model may be reactivated by specific contexts during actual reasoning, leading to recurring misjudgments. Orthogonal projection penalty aims to completely eliminate the component in the erroneous representation in the correct anchor point direction, forcing the illusory representation not only to move away but also orthogonal to the correct direction, achieving a cleaner separation.

[0075] set up and All have been L2 normalized. Vectors and The projected length in the direction is (i.e., cosine similarity). The squared magnitude of the projected vector is... Therefore, penalizing the square of the projection lengths forces them to be orthogonal: ; Complementary repulsion loss Directly reduce the cosine similarity to control the overall angle. Further elimination of parallel components will penalize any non-zero projections even when cosine similarity is already low. Combining these two approaches ensures both angular separation and directional independence, eradicating erroneous patterns from both angular distance and directional orthogonality dimensions.

[0076] 4. Training paradigm.

[0077] Each training batch consists of 70% general normal samples. And 30% artificial hallucination correction sample Hybrid composition. Gradient update rule: For standard quality inspection dialogue datasets Samples in: frozen , and Only calculate And update In The parameters. For the illusion recognition sample dataset. Samples in the dataset: Calculate the total loss function ,renew In Parameters, Student Hallucination Correction Branch and gating network Teacher hallucination correction branch It does not participate in gradient updates, but only synchronizes via EMA. After training, it only retains [the relevant data] during the inference phase. Remove all illusion correction branches and gating network modules.

[0078] 5. Training the overall process.

[0079] Step 1: Construct a standard quality inspection dialogue dataset and hallucination recognition sample dataset The standard quality control dialogue dataset includes a training dataset of standard data such as dialogue input sequences (question-and-answer dialogues). In the hallucination recognition sample dataset, each hallucination sample contains a hallucination error type label (i.e., the model's historical hallucination output sequence), a dialogue input sequence, and a manually corrected output sequence.

[0080] Step 2: Load the pre-trained Large Language Model (LLM) and add LoRA parameters; initialize the student hallucination correction branch and the small gating network; copy the weights to the teacher hallucination correction branch and freeze them all.

[0081] Step 3: In each training step, a batch is drawn from the pool. For normal samples in the standard quality inspection dialogue dataset, the standard generation loss is calculated forward only and the relevant parameters are updated. For hallucination samples in the hallucination recognition sample dataset, three representations are extracted in the manner described above. The total loss function is calculated and the student hallucination correction branch and the gating network are updated simultaneously.

[0082] Step 4: Execute EMA to update the parameters of the teacher hallucination correction branch.

[0083] Step 5: After training is complete, save only the LoRA weights of the large language model for online deployment.

[0084] Compared with the prior art, the technical solution of this embodiment has the following advantages: 1. High accuracy in hallucination recognition and significantly improved accuracy in dialogue quality control. The adaptive boundary rejection loss is graded according to the severity of three types of hallucinations: missed detection, false detection, and misclassification. The orthogonal projection penalty loss completely eliminates collinear remnants along the correct direction in hallucination representations, while the temperature-weighted MSE alignment loss protects key discrimination dimensions. This triple representation constraint mechanism improves the recognition accuracy of large language models and reduces hallucinations in large language models.

[0085] 2. The penalty standards are strictly aligned with human standards, ensuring auditable business compliance. A tiered exclusion mechanism is constructed using routine human review cases and their illusion type tags, ensuring that the sensitivity of the large language model to fatal errors such as missed illusions is completely consistent with that of quality control experts. Correct anchor points always represent human correction results, ensuring that the large language model's penalty standards are synchronized with the rules. Constraints at each stage, such as exclusion strength, dimensional importance, and orthogonality, have clear physical meanings. Illusion types are directly linked to penalty boundaries, facilitating auditing and rule backtracking, and effectively meeting the e-commerce dialogue quality inspection and compliance regulatory requirements.

[0086] 3. Zero additional cost for inference deployment and zero degradation of general quality inspection capabilities. After training, all illusion correction branches and gating networks are removed, and there are no additional structures on the inference side. Speed, memory usage, and latency are equivalent to the original SFT model, making it fully adaptable to high-concurrency online dialogue quality inspection. Hybrid batch conditional training ensures that the main generation task is not disturbed, the basic dialogue quality inspection judgment capability does not degrade, and catastrophic forgetting is completely avoided.

[0087] 4. Training is stable and controllable, with extremely low long-term maintenance costs. All losses are calculated deterministically in pairs, eliminating the need for negative sampling, reinforcement learning exploration, or reward models, fundamentally preventing RL-style crashes, reward hacks, and slow convergence. Engineering modifications only require adding a few modules to the standard SFT code, which can be continuously updated using daily review data generated by the business, forming a self-operating closed loop of "model illusion → manual correction → representation constraints → illusion reduction," resulting in near-zero long-term maintenance costs.

[0088] Reference Figure 3 This application also provides an illusion recognition system, which includes: The dataset construction module 100 is used to construct a hallucination recognition sample dataset based on hallucination recognition samples containing triplets of dialogue input sequences, model historical hallucination output sequences, and manually corrected output sequences; and to construct a standard quality inspection dialogue dataset based on dialogue input sequences. The input sequence construction module 200 is used to construct the illusion-side input from the dialogue input sequence and the model history illusion output sequence; and to construct the correct-side input from the dialogue input sequence and the manually corrected output sequence. The bias correction branch construction module 300 is used to construct student hallucination bias correction branches and teacher hallucination bias correction branches, which include linear hidden layers, nonlinear encoders, layer normalization, and L2 normalization. The student hallucination bias correction branches and teacher hallucination bias correction branches have the same structure. The first data processing module 400 is used to process the hallucination side input using a large language model to obtain the first intermediate vector; the first intermediate vector is input to the student hallucination correction branch to obtain the hallucination representation; The second data processing module 500 is used to process the correct side input using a large language model to obtain the second intermediate vector; input the second intermediate vector into the student hallucination correction branch to obtain the student's correct representation; and input the second intermediate vector into the teacher hallucination correction branch to obtain the teacher's anchor point representation. The loss function construction module 600 is used to construct the standard generation loss based on the correct side input and the manually corrected output sequence; to construct the temperature-weighted MSE alignment loss, adaptive boundary cosine repulsion loss, and orthogonal projection penalty loss based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation; to construct the ranking regularization loss based on the hallucination type in the model's historical hallucination output sequence; and to obtain the total loss function by weighted summation of the standard generation loss, the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, the ranking regularization loss, and the orthogonal projection penalty loss. The model illusion recognition module 700 is used to iteratively update the parameters in the large language model using a standard quality inspection dialogue dataset and a standard generation loss; iteratively update the parameters, student illusion correction branch, and gating network in the large language model using an illusion recognition sample dataset and a total loss function; iteratively update the teacher illusion correction branch using an exponential moving average method; after completing the iterative update, a trained large language model is obtained; and illusion recognition for dialogue quality inspection is performed using the trained large language model.

[0089] It should be noted that since the hallucination recognition system in this embodiment is based on the same inventive concept as the hallucination recognition method described above, the corresponding content in the method embodiment is also applicable to this system embodiment, and will not be described in detail here.

[0090] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations. The acquisition, storage, use and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.

[0091] Reference Figure 4This application also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in memory, and the processor executes at least one program to implement the illusion recognition method described above in this disclosure.

[0092] This electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0093] The electronic devices according to embodiments of this application will now be described in detail.

[0094] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure. The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and called by the processor 1600 to execute the illusion recognition method of the embodiments of this disclosure.

[0095] The input / output interface 1800 is used to implement information input and output. The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900); The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.

[0096] This disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the above-described hallucination recognition method.

[0097] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0098] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.

[0099] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0101] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0102] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0103] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0104] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. The embodiments of this application have been described in detail above with reference to the accompanying drawings, but this application is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of this application.

[0108] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.

Claims

1. A method for hallucination recognition, characterized in that, The method includes: A hallucination recognition sample dataset is constructed based on hallucination recognition samples containing triplets of dialogue input sequences, model historical hallucination output sequences, and manually corrected output sequences; a standard quality inspection dialogue dataset is constructed based on the dialogue input sequences. The dialogue input sequence and the model history hallucination output sequence are used to construct the hallucination-side input; the dialogue input sequence and the manually corrected output sequence are used to construct the correct-side input. Construct student hallucination correction branches and teacher hallucination correction branches that include a linear hidden layer, a nonlinear encoder, layer normalization, and L2 normalization. The student hallucination correction branches and the teacher hallucination correction branches have the same structure. The hallucination input is processed using a large language model to obtain a first intermediate vector; the first intermediate vector is then input into the student hallucination correction branch to obtain a hallucination representation. The large language model is used to process the correct side input to obtain a second intermediate vector; the second intermediate vector is input to the student hallucination correction branch to obtain the student's correct representation; the second intermediate vector is input to the teacher hallucination correction branch to obtain the teacher's anchor representation. Based on the correct side input and the manually corrected output sequence, a standard generation loss is constructed; based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation, a temperature-weighted MSE alignment loss, an adaptive boundary cosine repulsion loss, and an orthogonal projection penalty loss are constructed; based on the hallucination type in the model's historical hallucination output sequence, a ranking regularization loss is constructed; the standard generation loss, the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, the ranking regularization loss, and the orthogonal projection penalty loss are weighted and summed to obtain the total loss function; The parameters in the large language model are iteratively updated using the standard quality inspection dialogue dataset and the standard generation loss; the parameters in the large language model, the student hallucination correction branch, and the gating network are iteratively updated using the hallucination recognition sample dataset and the total loss function; the teacher hallucination correction branch is iteratively updated using the exponential moving average method; after completing the iterative update, a trained large language model is obtained; and the trained large language model is used for dialogue quality inspection hallucination recognition.

2. The hallucination recognition method according to claim 1, characterized in that, The process of using a large language model to process the hallucination-side input to obtain a first intermediate vector includes: Building a large language model based on multiple Transformer layers; The input to the hallucination side is processed using a large language model, and the output vector of the penultimate layer in the multi-layer Transformer is used as the first intermediate vector.

3. The hallucination recognition method according to claim 1, characterized in that, The method for constructing temperature-weighted MSE alignment loss, adaptive boundary cosine repulsion loss, and orthogonal projection penalty loss based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation includes: Based on the teacher anchor point representation and the student correct representation, a temperature-weighted MSE alignment loss is constructed. Based on the illusion representation and the teacher anchor representation, an adaptive boundary cosine repulsion loss and an orthogonal projection penalty loss are constructed.

4. The hallucination recognition method according to claim 3, characterized in that, The construction of a temperature-weighted MSE alignment loss based on the teacher anchor representation and the student correct representation includes: The teacher anchor representation is input into the gating network to obtain the dimension importance vector; The difference between the teacher anchor point representation and the student correct representation is calculated to obtain the difference result; Based on the dimensionality importance vector and the difference results, a temperature-weighted MSE alignment loss is constructed.

5. The hallucination recognition method according to claim 3, characterized in that, The adaptive boundary cosine repulsion loss, constructed based on the hallucination representation and the teacher anchor representation, includes: Assign learnable embedding vectors to the hallucination types in the historical hallucination output sequence of the model; The embedded vector is processed by linear projection and activation function to obtain adaptive boundary values; Calculate the cosine similarity between the hallucination representation and the teacher anchor representation; Based on the adaptive boundary value and the cosine similarity, an adaptive boundary cosine repulsion loss is constructed.

6. The hallucination recognition method according to claim 3, characterized in that, The construction of the orthogonal projection penalty loss based on the hallucination representation and the teacher anchor representation includes: Calculate the squared magnitude of the projection vector between the hallucination representation and the teacher anchor representation, and construct the orthogonal projection penalty loss.

7. The hallucination recognition method according to claim 1, characterized in that, The ranking regularization loss is constructed based on the hallucination types in the historical hallucination output sequence of the model, including: Based on the hallucination types in the historical hallucination output sequence of the model, construct an ordered set of pairs that satisfy a strict partial order relation; Based on the ordered set of pairs and the adaptive boundary value corresponding to each illusion type, a sorting regularization loss is constructed.

8. A hallucination recognition system, characterized in that, The system includes: The dataset construction module is used to construct an illusion recognition sample dataset based on illusion recognition samples containing triplets of dialogue input sequences, model historical illusion output sequences, and manually corrected output sequences; and to construct a standard quality inspection dialogue dataset based on the dialogue input sequences. An input sequence construction module is used to construct the illusion-side input from the dialogue input sequence and the model history illusion output sequence; and to construct the correct-side input from the dialogue input sequence and the manually corrected output sequence. The bias correction branch construction module is used to construct student hallucination bias correction branches and teacher hallucination bias correction branches, which include linear hidden layers, nonlinear encoders, layer normalization, and L2 normalization. The student hallucination bias correction branches and the teacher hallucination bias correction branches have the same structure. The first data processing module is used to process the hallucination side input using a large language model to obtain a first intermediate vector; and input the first intermediate vector into the student hallucination correction branch to obtain a hallucination representation. The second data processing module is used to process the correct side input using the large language model to obtain a second intermediate vector; input the second intermediate vector into the student hallucination correction branch to obtain the student correct representation; and input the second intermediate vector into the teacher hallucination correction branch to obtain the teacher anchor point representation. The loss function construction module is used to construct a standard generation loss based on the correct side input and the manually corrected output sequence; to construct a temperature-weighted MSE alignment loss, an adaptive boundary cosine repulsion loss, and an orthogonal projection penalty loss based on the hallucination representation, the student's correct representation, and the teacher's anchor point representation; to construct a ranking regularization loss based on the hallucination type in the model's historical hallucination output sequence; and to perform a weighted summation of the standard generation loss, the temperature-weighted MSE alignment loss, the adaptive boundary cosine repulsion loss, the ranking regularization loss, and the orthogonal projection penalty loss to obtain the total loss function. The model illusion recognition module is used to iteratively update the parameters in the large language model using the standard quality inspection dialogue dataset and the standard generation loss; iteratively update the parameters in the large language model, the student illusion correction branch, and the gating network using the illusion recognition sample dataset and the total loss function; iteratively update the teacher illusion correction branch using the exponential moving average method; after completing the iterative update, a trained large language model is obtained; and illusion recognition for dialogue quality inspection is performed using the trained large language model.

9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor to enable the at least one control processor to perform the illusion recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the illusion recognition method as described in any one of claims 1 to 7.