Real-time hallucination detection system and method for large model text generation based on probe technology

CN122819447APending Publication Date: 2026-09-25UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610783554.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]针对上述技术的不足,本发明的目的在于提供一种基于探针技术的大模型文本生成实时幻觉检测方法及系统,用以解决现有技术中长文本幻觉检测精度低、延迟高的问题

Benefits of technology

1、实体级流式检测框架将幻觉检测转化为级序列标注任务,支持生成过程实时干预,提升长文本检测能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819447A_ABST
    Figure CN122819447A_ABST
Patent Text Reader

Abstract

The application provides a large model text generation real-time hallucination detection system and method based on a probe technology.The system is composed of an entity-level labeling module, a probe training optimization module and a real-time detection intervention module.The entity-level labeling module constructs a level data set containing true and false fact trajectory comparison through a reverse fact induction mechanism.The probe training optimization module extracts the hidden state difference between adjacent levels of the large model as a new feature, trains a lightweight probe to identify the offset signal of the fact information during the internal circulation of the model, and uses regularization to ensure the stability of the model semantic space.The real-time detection intervention module monitors the interlayer difference trajectory in real time during the streaming decoding process, and realizes non-invasive generation correction through a spatial orthogonal projection algorithm.The application can effectively filter word sense noise, significantly improve the detection accuracy, and maintain extremely low reasoning delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to factual monitoring technology for long text generation in natural language processing. Specifically, this invention proposes a real-time entity-level illusion detection method based on a large language model, aiming to solve the problem of balancing real-time performance and accuracy in the long text generation process faced by traditional methods. Background Technology

[0002] With the widespread application of large language models in text generation tasks, the illusion problem (i.e., generated content that does not conform to objective facts) has become a key bottleneck restricting its application in high-risk domains. Existing illusion detection methods mainly face the following two limitations: 1) Short text limitation: Existing methods such as semantic entropy are mainly designed for short texts and cannot effectively handle entity-level illusions in long texts with multiple paragraphs. These methods have weak error detection capabilities for long texts, especially when there are intertwined correct / incorrect claims between entities in long texts, making accurate detection difficult. 2) Delayed verification problem: Existing verification methods, such as... Such methods typically require step-by-step verification using external knowledge bases, resulting in high latency (usually over 10 seconds) and high computational costs. For streaming generation tasks, this approach cannot meet real-time requirements.

[0003] Therefore, traditional hallucination detection methods cannot meet the requirements of modern long text generation tasks, especially the real-time requirements in high-risk scenarios such as medicine and law, and there is an urgent need for an efficient and accurate real-time detection solution. Summary of the Invention

[0004] To address the shortcomings of the aforementioned technologies, the present invention aims to provide a real-time hallucination detection method and system for large-model text generation based on probe technology, in order to solve the problems of low accuracy and high latency in long text hallucination detection in the prior art.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The core technical solution of this invention is to construct an integrated system that includes entity-level annotation, probe training and optimization, and real-time detection intervention. Through a lightweight architecture and regularization mechanism, it solves the problems of real-time performance and accuracy in long text illusion detection.

[0006] A real-time hallucination detection system based on probe technology for large-scale text generation includes: Entity-level annotation module, probe training and optimization module, and real-time detection and intervention module; The entity-level annotation module is connected to the probe training and optimization module via data communication, and the probe training and optimization module interacts with the real-time detection and intervention module via signal communication. The entity-level annotation module is configured to: receive text generated by the target large language model, generate authenticity labels for each token through fact-checking, and construct an annotation dataset containing a comparison between real trajectories and hallucinatory trajectories; The probe training optimization module is configured to: extract the hidden state difference between two adjacent layers of the target large language model during the token generation process as input features, and use the real labels in the labeled dataset as supervision signals to train a lightweight probe to predict the probability that the current token belongs to a hallucination. The real-time detection and intervention module is configured to: extract the hidden state difference in real time when the target large language model performs streaming decoding to generate each token, input it into the trained lightweight probe to calculate the illusion probability, compare the illusion probability with a preset threshold, and perform generation intervention or correction operation according to the comparison result.

[0007] Preferably, the entity-level annotation module includes a long text generation unit, an entity recognition annotation unit, a label generation unit, and a noise control unit; The output of the long text generation unit is connected to the input of the entity recognition and annotation unit, the output of the entity recognition and annotation unit is connected to the input of the label generation unit, and the output of the label generation unit is connected to the input of the noise control unit. The long text generation unit is configured with a target large language model, the entity recognition and annotation unit has a built-in search-enhanced large language model annotator, and the noise control unit is configured with a span-accurate matching mechanism. The long text generation unit takes the dataset and the target large language model as input and outputs the generated training text; the entity recognition and annotation unit performs entity span annotation and realism label annotation on the training text; the label generation unit maps the entity span annotation and realism label to a token-level label sequence; the noise control unit filters the token-level label sequence and outputs a high-confidence labeled dataset.

[0008] The annotation data output from the entity-level annotation module is input into the lightweight probe architecture unit, and simultaneously fed into the composite loss function unit and the KL regularization unit to collaboratively optimize the probe parameters.

[0009] The entity-level annotation module is trained using an adversarial dataset of the activation trajectory control group.

[0010] The probe training optimization module includes a lightweight probe architecture unit and a composite loss function unit. The regularization unit; the lightweight probe architecture unit is used to output the illusion probability, and the input feature of the lightweight probe is the hidden state difference between adjacent layers of the target large language model; the lightweight probe architecture unit is configured with either a linear probe or a low-rank adapter probe.

[0011] Preferably, the long text generation unit calls Text generation from long text datasets; The noise control unit, by configuring a span-accurate matching mechanism, eliminates non-aligned samples caused by text truncation or word segmentation deviation, ensuring strong consistency between labels and activation features; The entity recognition and annotation unit uses a large language model annotator to verify the authenticity of entity spans in the generated text and maps the annotation results to the corresponding entities. The sequence is labeled with entity types including names, organizations, and dates, and the labeling results are divided into two categories: "support" and "illusion".

[0012] Entity recognition and annotation unit through Complete the identification and span labeling of entities such as names and organizations, and verify their authenticity by combining web search, and output the "support" or "illusion" label.

[0013] The core advantage of this design lies in leveraging authoritative models and datasets to provide highly reliable labeled data for subsequent detection tasks, laying the foundation for probe training.

[0014] Preferably, the lightweight probe architecture unit has its input end connected to the first... Layer and First The hidden state differencer of the layer has its output connected to a binary classification head. The composite loss function unit employs a function that incorporates entity-level weighted cross-entropy to enhance the probe's ability to capture core entity hallucination signals. The Regularization unit, by introducing Divergence constraints ensure that the lightweight probe preserves the manifold structure of the original semantic representation space of the target large language model when extracting hallucination features.

[0015] Preferably, the loss function expression of the composite loss function unit is: ; in The weighted parameter for the travel entity is λt, which is the dynamic annealing coefficient ranging from 0 to 1; The The total loss expression for the regularization unit is: ; Where β is the KL regularization parameter, used to balance the probe detection performance with the original model generation capability.

[0016] The lightweight probe architecture unit offers two optional schemes: linear probes or probes. Linear probes are deployed in the intermediate layer of the model. The probability of hallucination is output through a linear classification head; The probe is injected into multiple network layers to enhance detection capabilities by fine-tuning local weights.

[0017] Both schemes achieve the goal of "low overhead and high accuracy", among which The probe performs better when migrating across models, and the memory increment can be controlled within 5%.

[0018] Preferably, the real-time detection and intervention module includes a streaming detection unit, a selective termination unit, a dynamic threshold adjustment unit, and a probability distribution reshaping unit; The streaming detection unit is configured to calculate the hidden state difference in real time and map it to the illusion probability during the decoding cycle of each token, and output the basic threshold. The selective termination unit has a preset threshold. The configuration is set to trigger generation termination and return an "Unverified" flag when the maximum illusion probability within the entity span exceeds a threshold t. The dynamic threshold adjustment unit receives the base threshold output by the streaming detection unit and dynamically scales the soft threshold based on the vertical domain sensitivity of the generated text. With hard threshold ; The dynamic threshold adjustment unit adjusts the threshold. Balancing detection accuracy with response generation rate has the advantage of enabling real-time monitoring and flexible intervention in the generation process, thus preventing the spread of hallucinatory information.

[0019] The probability distribution reshaping unit is configured to, when the hallucination probability exceeds But not exceeding Triggered Orthogonal projection correction of space, when the probability exceeds The process will trigger an abort and return an "unverified" flag.

[0020] The input to the streaming detection unit is the inter-layer activation difference residual (hidden state difference) of the target large language model and the trained probe. The output is the real-time hallucination probability and it is passed to the selective termination unit. The dynamic threshold adjustment unit dynamically adjusts the threshold according to the scene risk level and configures the selective termination unit.

[0021] The detection system provided by this invention has a clear workflow. First, it generates an entity-level annotation module. A multi-level labeled dataset is used; secondly, a lightweight probe is trained based on this dataset, and a composite loss function is applied. The model is optimized using regularization; finally, during text generation, the real-time detection and intervention module completes the process. The probability calculation and dynamic intervention of hallucination levels achieve a closed loop of "data-model-application".

[0022] A real-time hallucination detection method for large-model text generation based on probe technology, using the aforementioned real-time hallucination detection system for large-model text generation based on probe technology, includes the following steps: Step 1: Using the entity-level annotation module, generate a dataset with comparative features between real fact trajectories and reverse pseudo-fact trajectories. Level dataset; Step 2: Extract the hidden state difference of the target large language model as input, train the lightweight probe through the probe training optimization module, and use the composite loss function and KL regularization to establish the mapping relationship between inter-layer perturbation and factual illusion, so that the trained probe can output token-level illusion probability. Step 3: During the streaming inference process of the target large language model, the hidden state difference is monitored in real time. The token-level illusion probability is calculated through the real-time detection intervention module. The illusion probability is compared with the dynamically set soft threshold and hard threshold. If the illusion probability is lower than the soft threshold, the current token is output normally. If the illusion probability is between the soft threshold and the hard threshold, the Logits space orthogonal projection correction is performed. If the illusion probability is not lower than the hard threshold, the generation is terminated and the "unverified" flag is returned.

[0023] Compared with the prior art, the advantages of the present invention are: 1. The entity-level streaming detection framework transforms hallucination detection into... The sequence labeling task supports real-time intervention in the generation process and improves the ability to detect long texts. 2. Low-cost annotation and generalization design: Training cross-model probes using single-model labeled data improves annotation efficiency by 80%; 3. KL regularized probes achieve a Pareto balance between detection accuracy and the behavior of the original model; 4. Breakthrough in long text detection performance, solving the problem of low recall rate in traditional methods. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the architecture of the travel science popularization long text hallucination detection system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the entity-level annotation pipeline workflow in the travel domain of this invention. Figure 3 This is a schematic diagram illustrating the deployment and training principle of the travel scenario probe in this invention; Figure 4 This is a logic diagram for adaptive switching of probe types in travel scenarios according to the present invention; Figure 5 This is a flowchart of the real-time hallucination detection and intervention process for long-text travel science popularization articles in this invention; Figure 6 This is a comparison chart of the experimental results for detecting hallucinations in travel scenarios according to the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] like Figures 1-6 This embodiment discloses a real-time hallucination detection method and system for large-scale text generation based on probe technology, focusing on the specific application scenario of generating long-form travel science popularization texts. In such scenarios, long-form texts contain a large amount of knowledge-intensive content, and the appearance of hallucinations may mislead users and cause various risks.

[0027] This embodiment addresses the core needs of this scenario and provides a practical implementation method based on the core technical logic of this invention. All technical details and parameter configurations are adapted to the characteristics of the travel science popularization scenario, ensuring detection accuracy, real-time performance, and practicality. At the same time, the core technology of this invention can be flexibly transferred to other high-risk long text scenarios.

[0028] Based on the specific characteristics of long-text scenarios in travel science popularization, this embodiment determines the selection and exclusive configuration of each core component as follows. All configurations are adapted to the data characteristics and hallucination detection requirements in the field of travel science popularization, providing basic support for subsequent technical implementation.

[0029] For details on the hierarchical relationships, data flow, and key parameters of each core component, please refer to [link / reference]. Figure 1 The diagram presents a hierarchical structure, clearly delineating the connections between core modules at each level and the data flow direction between units within each module. It simultaneously labels the inputs, outputs, and core parameters of each unit, intuitively showcasing the overall system architecture logic. Specifically: In the entity-level annotation module, the long text generation unit output → entity recognition annotation unit input → label generation unit input → noise control unit input form a cascaded connection; the noise control unit output sends high-confidence annotation data to the lightweight probe architecture unit of the probe training and optimization module; the trained probe is loaded into the streaming detection unit of the real-time detection and intervention module; the streaming detection unit output connects to the selective termination unit; and the dynamic threshold adjustment unit configures the threshold parameters of the selective termination unit.

[0030] Long text dataset: using The dataset contains This dataset contains long-form travel information texts covering various entity types in the travel field that are prone to causing hallucinations, including attraction names, transportation routes, ticket information, accommodation standards, safety precautions, local customs, and travel equipment suggestions. All texts have been reviewed and verified by experienced travel professionals and cultural tourism specialists, ensuring they are free of factual errors. This dataset is suitable for training and validating hallucination detection models in the travel field. Figure 1 The core component of the data layer, and authoritative travel information sources, The travel version of the search interface works in tandem to provide high-quality data input for the entire system.

[0031] Target Large Language Model: Selection After being fine-tuned with a travel-specific dataset, the model demonstrates excellent semantic expression capabilities in the task of generating long-form travel science texts. It can accurately cover multiple sub-fields of travel knowledge, such as popular domestic and international attractions, transportation, accommodation booking, and regional customs. It is adaptable to the needs of generating travel science texts with multiple paragraphs and knowledge points, and can generate science content and guidance suggestions that conform to actual travel norms.

[0032] Search Enhancement Annotator: Select This annotation tool enhances the recognition and semantic comparison capabilities of travel-related entities compared to the general version. It can be integrated with authoritative cultural and tourism platforms and official tourist attraction channels to verify the factual accuracy of travel entities. It supports the accurate identification of various travel-specific entities, adapting to the fact-checking needs of the travel sector and improving the accuracy of travel entity annotation and verification. This annotation tool belongs to... Figure 1 The core module of the model layer is mainly used to work with the data layer to complete the annotation and fact-checking of travel entities, providing high-quality labeled data for probe training.

[0033] Probe deployment architecture: employing linear probes and The probe features a switchable architecture, balancing the need for hallucination detection accuracy with lightweight models, and adapting to travel science popularization scenarios of varying complexity: the linear probe is used for routine travel science popularization scenarios, such as basic travel guides for popular attractions, city transportation guides, and general accommodation suggestions. The probe is used in complex travel science popularization scenarios, such as planning itineraries that connect multiple attractions, travel guides for remote areas, travel strategies for special groups, and precautions for cross-border travel.

[0034] The probe architecture is as follows Figure 1 The core content of the probe layer, Figure 3 The text clearly indicates that the probes are deployed in... 76th floor The adapter injection sites are the first 30 layers and the last 20 layers; simultaneously, the linear probe and... For details on the probe's scenario adaptation logic, please refer to [link / reference]. Figure 4The diagram clearly shows the complete logic of scene recognition and probe invocation, as well as the corresponding detection latency.

[0035] Based on the above core component configuration, and targeting the need for illusion detection in long-text travel science popularization articles, the specific implementation steps of this invention are as follows. Each step works in concert to achieve the full-process implementation from data annotation and probe training to real-time detection and intervention. Each step is designed in conjunction with the details of the travel scenario to ensure the feasibility and practicality of the technical solution.

[0036] Step 1: Automated construction of entity-level annotation pipeline in the travel domain.

[0037] For the complete implementation process of this step, please refer to [link / details]. Figure 2 This diagram presents a linear, step-by-step approach, comprising four core steps, branch decisions, and result output nodes. It clearly labels the scenario-specific details and logical connections of each step. The core objective of this step is to construct a fully automated data annotation process specific to the travel industry, integrating the generation capabilities of a finely tuned version of the travel science popularization language model with the objectivity of authoritative travel information sources to obtain high-precision data in the travel field. The system employs factual labeling to provide high-quality travel-specific data support for subsequent probe training. Specific implementation details are as follows: Multi-dimensional entity span recognition in the travel field. This step corresponds to... Figure 2 The "Entity Span Recognition" operation node in the middle.

[0038] The "Entity Span Recognition" operation node clearly identifies four core travel entities and their key sub-items, which correspond one-to-one with the following content. This is based on travel-specific named entity recognition technology and... Annotator, from The model in Generated on the dataset The text is a long article on travel science, in which all potential travel entities are comprehensively extracted.

[0039] Unlike general-domain entity recognition, this embodiment focuses on strengthening the recognition and monitoring of "practical" entities in the travel field. Based on the needs of travel science popularization scenarios, the scope of travel entity recognition is defined as follows: (1) Entities related to the scenic spot, including the scenic spot name, specific address, opening hours, ticket price, ticket reservation method, internal tour route, introduction of core landscape, and flow restriction rules; (2) Traffic-related entities, including intercity traffic routes, intracity traffic modes, locations of traffic hubs, and congestion warnings during travel periods; (3) Accommodation-related entities, including hotel / guesthouse name, specific address, accommodation price, booking channels, supporting facilities, check-in and check-out rules, and surrounding transportation convenience, etc.; (4) Entities related to safety and customs, including travel safety precautions, regional customs and taboos, travel tips for special periods, emergency contact information, local specialties and purchase channels, etc.

[0040] These types of travel entities are most prone to factual bias during the generation of long-form travel science texts, making them a key focus for detecting travel scenario illusions.

[0041] By using the aforementioned method of extracting travel-specific entities, the detection system can ensure that it fully covers the key travel knowledge points in long travel science texts that may cause hallucinations, avoiding missed detections of hallucinations due to the omission of travel entities, thereby preventing users from wasting time, losing property, and facing safety risks when traveling.

[0042] Enhanced fact-checking for travel-specific searches. This step corresponds to... Figure 2 The "Search-enhanced fact-checking" operation node in the middle.

[0043] The "Search-Enhanced Fact Check" operation node embeds three sets of structured query term examples, and clearly marks them. The authoritative travel source matches the following case and source description perfectly.

[0044] For each extracted travel entity span, the search-enhanced fact-checking node automatically deconstructs the travel generation context and the travel sub-sector to which the entity belongs, constructing a set of highly structured travel-specific search query statements. This ensures that the query content can accurately locate the factual truth of the travel entity, adapt to the information retrieval needs of the travel field, and avoid travel evidence bias caused by fuzzy queries.

[0045] Among them, "travel entity span" refers to "entity span in the travel field," which is the text span for travel-related entities (attractions, transportation, tickets, etc.).

[0046] Enhanced fact-checking nodes for search will be automatically generated based on travel scenario requirements. The group of structured query terms corresponds to the verification needs of different travel entities.

[0047] Subsequently, the enhanced fact-checking node calls the travel-specific search interface to prioritize obtaining webpage snapshots and data entries from authoritative travel sources, ensuring the authority of the fact-checking and avoiding verification errors caused by non-authoritative travel sources.

[0048] By introducing the aforementioned authoritative travel information sources, the illusions caused by insufficient timeliness of travel training data in large models, such as adjustments to attraction opening hours, optimization of transportation routes, and changes in ticket prices, can be effectively eliminated. At the same time, the information bias of a single travel information source can be avoided, ensuring the accuracy of fact-checking of travel entities.

[0049] Multi-agent discrimination mechanism in the travel domain. This step corresponds to... Figure 2 The multi-agent discrimination process shown in the diagram ensures the accuracy of travel entity discrimination through multi-angle cross-validation.

[0050] The "Multi-agent discrimination" operation node clearly marks the order of the three semantic comparison logics and the discrimination rule of "all pass = real, any fail = illusion", intuitively showing the branch judgment process.

[0051] To address the practical needs of travel science popularization scenarios, this embodiment introduces a travel-specific multi-agent discrimination mechanism to further improve the accuracy and reliability of travel entity fact verification, avoid omissions and misjudgments of travel illusions due to verification errors, and ensure the practicality and security of travel science popularization content.

[0052] The authoritative travel information obtained through the search and the travel entity information generated by the model are simultaneously input into the preset travel annotation model. This annotation model combines professional knowledge in the travel field to complete a triple travel-specific semantic comparison, ensuring the comprehensiveness and professionalism of the fact-checking. All three comparisons are tailored to the needs of the travel scenario.

[0053] First, the existence of the travel entity is verified to confirm whether the travel entity exists in the authoritative travel information sources found in the search, and to rule out the possibility that the model has fabricated the travel entity. Secondly, the verification of travel entity attributes involves checking whether the relevant attributes of the travel entity conform to the actual situation. The key attributes to be verified include the opening hours of attractions, ticket prices, transportation frequency and duration, accommodation prices, and local customs. These attributes are directly related to the practicality of travel science popularization and are the core of the verification process. Third, verify the logical coherence of the travel context to confirm whether the contextual description of travel entities in long-form travel science popularization texts is consistent with authoritative travel evidence, and avoid the illusion that the entity itself is correct but the contextual logic is contradictory.

[0054] Judgment rule: The travel entity can be judged as "real" only when the authoritative travel evidence found in the search and the travel content generated by the model achieve a high degree of semantic consistency in all three comparisons mentioned above; if any one comparison fails, it is judged as "illusion" and the illusion type is marked.

[0055] This discrimination method based on multi-source authoritative travel evidence can effectively avoid annotation errors caused by incomplete results and information biases from a single travel search engine, significantly improve the accuracy of annotation data in the travel field, and provide reliable travel-specific annotation data for subsequent probe training.

[0056] Level label generation and noise control. This step corresponds to... Figure 2 The "t" in The "Level Label Generation and Noise Control" operation node and the final result output node ( Figure 2 Output in (Level tag)

[0057] “t The "Level Label Generation and Noise Control" operation node clearly indicates the span matching algorithm and confidence filtering threshold. The effective dataset after noise removal, and the final output. Level label data.

[0058] To address the misalignment between travel entity spans and token sequences, this embodiment develops a travel-specific span matching algorithm tailored to the characteristics of travel entities. This algorithm determines the optimal span boundary through dynamic programming based on the alignment relationship between travel entity boundaries and token sequence positions, accurately mapping the identified travel entities to the corresponding token position sequences and ensuring the accuracy of entity-level annotation boundaries.

[0059] If the entity is determined to be a hallucination, the algorithm will backtrack the travel entity to the original... The start and end positions in the stream, covering all sub-streams. All marked as Conversely, if the travel entity is determined to be real, then all its sub-entities will be covered. Marked as ,make sure The level-based labels and travel entity-level factual judgments are accurately matched, avoiding detection errors caused by alignment deviations.

[0060] Considering the timeliness and diversity of travel data, in order to reduce the impact of edge noise on travel training data, this solution sets up a travel-specific minimum confidence filtering mechanism after label generation and before data is stored in the database. Combining the high practicality requirements of travel scenarios, the confidence threshold is set to 0.85 (higher than the general scenario threshold) to filter out labeled data with confidence below the threshold, ensuring the purity of training data and avoiding noise samples from affecting the probe detection accuracy.

[0061] Filtering rules: For search results that are extremely vague, contain obvious information disputes, or have a lower confidence level than the labeler model's decision level. All labeled samples were removed to avoid such ambiguous samples affecting the detection accuracy of the probe, thereby mitigating the risk to users traveling.

[0062] Ultimately, by The tag generation unit inherits the travel entity tag to the corresponding to form a complete Level label data; the noise control unit uses a span-accurate matching mechanism and confidence filtering to... Selected from the initial travel science text A total of 1000 travel science popularization texts were collected to form an effective dataset, which was used for training subsequent travel scenario probes to ensure that the probes could accurately adapt to the detection needs of travel science popularization scenarios.

[0063] The “valid dataset” refers to the “high-precision labeled data in the travel domain” in step 2.

[0064] Step 2: Selection, training, and optimized deployment of probe layers.

[0065] The probe deployment location, training loss logic, and... For probe optimization details, please see Figure 3 The diagram is divided into two sub-modules: "Deployment Location" and "Training Loss," clearly showing the model hierarchy, mathematical formulas, parameter configurations, and feature enhancement logic.

[0066] After acquiring high-precision labeled data in the travel field, this step, tailored to the characteristics of long-form travel science texts, [further details needed]. The model's intermediate layers deploy lightweight probe architecture units, enabling real-time capture of hallucination features in the travel domain through linear decomposition of the model's internal activation states; simultaneously, it introduces... A regularization mechanism is used to avoid probe overfitting and degradation of the original model's travel text generation capabilities, ensuring that after probe deployment, the original model can still accurately generate travel science content that conforms to actual travel regulations. Specific implementation details are as follows: Probe level selection and operating mode. Figure 3 The diagram in the middle is presented as a vertical hierarchy. The model's shallow, deep, and optimal layers clearly identify the probe's "non-invasive bypass" operation mode and the activation vectors. The extraction location corresponds precisely to the content in this section. The probe in this solution uses a "non-invasive bypass" operation mode, without altering the slightly modified version of the travel science popularization. The original weight parameters of the model can effectively avoid interference with the original model's ability to generate travel semantics, ensuring that the quality of the original model's travel text generation is not affected.

[0067] In model decoding each During the process, the system extracts the activation vectors of the deep residual flow of the target model in real time. This activation vector accurately reflects the model's generation of the current trip. The internal semantic state of time is used to determine travel. Are these key characteristics of hallucinations—especially for crucial travel information such as attraction names, transportation routes, and ticket prices? The differences in the distribution of activation vectors within the model can effectively distinguish between real and hallucinatory content.

[0068] Multiple experiments have verified that the slightly modified version of travel science popularization is effective. In large-scale language models, factual features related to travel are most active in deeper regions of the model: shallow activation vectors struggle to effectively distinguish between real and illusory travel. However, excessively deep layers can lead to increased computational costs and feature redundancy, affecting the real-time generation efficiency of travel science texts and failing to meet the real-time needs of online travel consultations and science popularization pushes.

[0069] Therefore, considering the dual requirements of real-time performance and accuracy in travel scenarios, In the model, the preferred extraction is the first... The hidden states of the layer serve as input features for the probe, and the activation values ​​of this layer can effectively reflect the model's current travel state in the output. Beforehand, it is necessary to check whether there is "logical collapse" or "factual deviation" in its internal semantic space, taking into account both the effectiveness of travel features and computational efficiency, and ensuring that the detection delay does not affect the real-time push of travel science popularization content and consultation response.

[0070] The composite loss function for travel scenarios and Regularized design. Figure 3 The composite loss function unit is clearly shown. The mathematical formula for the total loss of the regularization unit is provided, along with annotations for the values ​​of each parameter. The design rationale for each parameter is explained using text boxes, clearly indicating the relevant parameters. The constraint of regularization, which "preserves the original model's ability to generate data," corresponds perfectly to the content in this section.

[0071] The core challenge of probe training lies in preventing the detector from overfitting to the travel entity features in the training set, while avoiding interference with the original model's travel text generation probability distribution, thus ensuring reliability in both detection and generation. To this end, this invention builds upon the traditional binary cross-entropy... Based on the loss function, introduce Divergence constraint terms are used to construct a travel-specific composite loss function to balance travel entities and non-travel entities. The detection weights are adjusted, and the influence of the probes on the original model's travel generation capability is controlled by regularization to adapt to the training requirements of travel scenarios.

[0072] The loss function expression of the composite loss function unit, i.e., the composite task loss function. The definition is as follows:

[0073] Based on the needs of travel scenarios, the parameter configurations and design principles are as follows, all of which are adapted to the detection characteristics of travel entities to ensure training effectiveness: For the weighted parameters of the travel entity, the preferred value in this embodiment is [value to be filled in]. This is used to force probes to prioritize capturing key travel entity features, such as attractions, transportation routes, and ticket information, thereby improving the detection accuracy of travel entity hallucinations—these entities pose the highest risk of hallucinations and are directly related to the practicality of travel science popularization, so they need to be given special attention. for The dynamic annealing coefficient is used to adjust non-travel entities at different training stages. Contribution: Early stage of training The focus is on training the probe to target travel entities. The ability to identify things, such as attractions, transportation routes, and ticket prices; as the training rounds increase, linearly increasing to Gradually upgrade non-travel entities The training weights ensure that the probe can fully cover all Hallucination detection to avoid non-travel entities Hallucinations were missed.

[0074] To further balance the probe detection performance with the original model's travel text generation capabilities, a new approach is introduced. The regularization unit, the total loss expression is as follows: ; in, The regularization parameter is set to a value that suits the needs of the travel scenario. Using the original model's travel text output distribution as a reference distribution .

[0075] Its beneficial effect lies in strengthening the critical travel through a composite loss function. Detection, with the help of Regularizing the difference between the probe output distribution and the original model's travel baseline distribution improves the performance of travel illusion detection while fully preserving the original model's ability to generate travel text, thus avoiding problems such as incorrect travel terminology and traffic route deviations caused by probe training.

[0076] The selection criteria for the core parameter λ in the travel scenario are as follows. In this embodiment, the selected parameter is... As the globally optimal solution for the travel scenario, its selection is based on the gradient balance principle of the travel scenario, ensuring that the contribution of each loss to parameter updates remains balanced during multi-objective optimization, avoiding a single objective dominating the gradient direction, and adapting to the detection characteristics of travel entities. The specific derivation is as follows: Let the detector parameters be... To ensure that the classification loss and regularization loss are consistent with the parameters The updated gradient contribution balance must satisfy the classification gradient. With regular gradient The modulus lengths in the parameter space are on the same order of magnitude.

[0077] For the classification loss term, its gradient magnitude During the training stabilization period, the activation vector is generated by the travel entity. The statistical distribution determines that, because the semantic features of travel entities are more specific and closer to real life, their numerical range is relatively fixed; for regularization terms... Its gradient expression is:

[0078] exist Assuming the activation norm distribution of deep residual flow is stable, and considering the semantic characteristics of travel entities: when At that time, the two gradient terms Norm ratio Internally, it better meets the gradient balancing needs of travel scenarios, ensuring the stability of the training process.

[0079] This gradient weight allocation method can effectively prevent oscillations during the optimization process, ensuring that while the probe extracts factual travel features, it maintains the travel language distribution characteristics of the original model to the greatest extent, avoiding problems such as incorrect travel terminology and deviations in transportation routes caused by probe training, and ensuring the quality of travel text generation of the original model.

[0080] Meanwhile, the probe, acting as a bypass module, must extract facts without distorting the original model's travel language flow. Multiple travel scenario experiments demonstrate that... It is the "critical point" for the stability of the travel system: if it is further reduced... The probe overfits to the travel entity features in the training set, causing the probability distribution of the model to fluctuate drastically when generating non-travel entity connectives, thus affecting the coherence of travel text; if the probe is improved... This can lead to excessively high weights for regularization terms, suppressing the probe's detection capabilities, increasing the risk of missed detections of travel entity illusions, and causing inconvenience and risks for users traveling.

[0081] Optimized deployment of LoRA probes. Figure 3 The injection locations and rank values ​​of the LoRA probes in the model are labeled, demonstrating the feature enhancement logic of low-rank adaptation, which corresponds to the content in this section; specific optimization deployment details can be found in [link / reference]. Figure 4 To understand further.

[0082] at the same time, The switching logic between probes and linear probes can be combined with... Figure 4 To understand further.

[0083] In complex travel science popularization scenarios, the feature capture capability of linear probes is insufficient to meet the requirements of high-precision detection. In such scenarios, the semantic relationships of travel entities are more complex, requiring stronger feature extraction capabilities to accurately identify hallucinations.

[0084] Therefore, this implementation supports the smooth switching of linear probes to... The probe enhances its travel feature extraction capabilities through low-rank adaptation while maintaining the model's lightweight characteristics, ensuring real-time detection performance in complex travel scenarios. Specific optimization deployment details for travel scenarios are as follows: Adapter deployment location, in The front of the model Layers and back Low-rank adapters are injected into each layer to enhance the ability to capture deep semantic features of travel entities and adapt to the needs of complex travel scenarios.

[0085] Adapter rank configuration, adapter rank set to The rank value can achieve an optimal balance between the ability to capture travel features and the computational cost. If the rank value is too low, the feature extraction of complex travel entities will be insufficient, and hallucinations will not be accurately identified. If the rank value is too high, the computational burden will be increased, affecting the real-time generation and detection of travel science popularization texts, and failing to meet the real-time needs of online travel consultation and science popularization push.

[0086] Loss function parameter adjustment, in the composite loss function, non-traveling entities Weighting coefficient The initial value is Each training session wheel lifting Until Gradually strengthen non-travel entities The detection weights are adjusted to avoid ambiguity in travel texts caused by logical errors in non-travel entities, ensuring the consistency and accuracy of travel science texts; travel entities Weighted parameters Keep constant, Regularization parameters Using the original model's travel text output distribution as a reference distribution .

[0087] Training parameter configuration, training epochs set to The system adapts to the complexity and timeliness of travel data; an adaptive learning rate optimization algorithm is used during training, with each... A travel-specific validation set evaluation will be conducted in one round. If it is continuous... Round Validation Set If the value does not improve, the early stopping mechanism is triggered to avoid overfitting and ensure the probe's generalization ability. Parameter saving and recalling, saving after training. The probe parameters are stored separately from the linear probe parameters. The system automatically switches the probe type according to the complexity of the travel science popularization scenario without manual intervention, and is used for real-time hallucination detection of long travel science popularization texts.

[0088] Step 3: Real-time hallucination detection and intervention for long-form travel science texts.

[0089] For details on the real-time detection and intervention process and case studies for this step, please see [link / reference]. Figure 5 This diagram integrates a dynamic process of real-time generation, detection, and intervention, embedding a real-world case study of the Palace Museum's May Day holiday travel, clearly demonstrating the complete logic of the entire detection and intervention process.

[0090] The trained probe layer (linear probe) probe) loaded to The model was used to construct a real-time hallucination detection and intervention system for long-form travel science popularization texts.

[0091] Based on a target large language model loaded with a probe layer, the hallucination detection and intervention system is functionally divided into a streaming detection unit, a selective termination unit, and a dynamic threshold adjustment unit. The streaming detection unit calculates the hallucination probability in real time, the selective termination unit determines whether to trigger termination based on a threshold, and the dynamic threshold adjustment unit dynamically optimizes the threshold parameters according to the scenario. Each unit is adapted to the real-time and practical requirements of travel scenarios.

[0092] The first step is to set a specific probability threshold for hallucinations in travel scenarios. .

[0093] Based on the practical needs of travel scenarios, set thresholds for general travel science popularization scenarios. When the probe calculates Hallucination probability At that time, it was determined to be a hallucination. Immediately trigger the intervention mechanism; when the probability of hallucination... At that time, it was determined to be true. This allows the model to continue generating travel science texts, ensuring the rigor of the detection, minimizing missed detections of travel illusions, and protecting users' travel safety and convenience.

[0094] The second step is scene-adaptive probe switching.

[0095] The system automatically switches probe types based on the complexity of the travel science popularization scenario, requiring no manual intervention and achieving precise adaptation to different scenarios. In typical travel science popularization scenarios, it automatically calls linear probes for rapid detection, with detection latency controlled within a specified range. Within this range, it will not affect the real-time push of travel science content and consultation response, ensuring user experience; in complex travel science scenarios, it will automatically call upon [the appropriate mechanism]. The probe enhances the ability to capture the features of traveling entities, ensuring the detection accuracy of complex traveling entities and controlling the detection latency within a specified range. Within this range, it balances detection accuracy and real-time performance to meet the needs of complex travel science popularization scenarios.

[0096] The third step is real-time detection and intervention, which involves responding to user inquiries and monitoring in real time.

[0097] Each time the model generates one The flow cytometry unit extracts the data in real time. The corresponding model number The layer's hidden state is input to the corresponding probe, which then quickly calculates the value. The probability of hallucination is determined to achieve real-time streaming detection.

[0098] When generating "accommodation price approx. When the probe detected "yuan / night", it was found that... "Yuan / night" corresponds to The probability of hallucination is Exceeding the threshold It was determined to be a hallucination. ; The selective abort unit immediately triggers a generation halt, stopping subsequent text generation to prevent erroneous travel information from continuing to be output and to avoid misleading users; Meanwhile, the system returns a travel-specific hallucination warning mark, which, based on the needs of the travel scenario, marks content with questionable facts, prompting users and relevant staff to verify the facts of that content, while providing clear and authoritative directions for verification, further preventing users from suffering financial losses during their travels; if users continue to ask questions, the system will generate subsequent responses after manual verification to ensure the accuracy and practicality of travel science content.

[0099] Step 4: Dynamic threshold adjustment.

[0100] The dynamic threshold adjustment unit dynamically optimizes the hallucination probability threshold based on the risk level of the travel scenario. This will further improve the detection adaptability to different travel scenarios, meet the differentiated needs of different travel scenarios, and achieve a balance between detection accuracy and generation efficiency.

[0101] High-risk travel scenarios: setting thresholds Adjusted to This will further enhance the stringency of detection, minimize the missed detection of travel hallucinations, and ensure the safety of users' travel. Typical low-risk travel scenarios: thresholds can be set. Adjusted to It balances detection accuracy and generation efficiency, avoids generation lag caused by excessive intervention, and improves user experience.

[0102] Step 4: Experimental verification of travel scenarios.

[0103] For a detailed comparison of the experimental results and verification of generalization ability in this step, please refer to [link / reference]. Figure 6 This chart, combining bar and line graphs, clearly displays the comparative data of multiple indicators and generalization ability data, intuitively demonstrating the advantages of this solution. To verify the detection performance, generalization ability, and real-time performance of this solution in long-text travel science popularization scenarios, multiple sets of comparative experiments were designed based on the needs of travel scenarios. The experimental settings and verification indicators were adapted to the characteristics of the travel domain to ensure the objectivity and reliability of the experimental results. The results are as follows: This invention demonstrates significant technical advantages and practical value in the context of long-form travel science popularization texts.

[0104] In terms of detection accuracy and reliability, this solution... Value reached Compared to traditional semantic entropy baseline methods, it improves Improves upon baseline methods for travel semantic entropy Its false negative rate for travel-related physical illusions is only [percentage missing]. Far lower than the comparison method and This confirms that the detection mechanism based on interlayer activation difference features can accurately identify travel entity illusions, effectively avoiding user travel inconvenience, property loss, and safety hazards caused by information errors. In terms of real-time performance, this solution controls the detection latency in typical scenarios to within [specific parameters]. Within, complex scenarios are below The overall increase in inference latency is insufficient. This fully meets the instantaneous response requirements of streaming interaction scenarios such as online consultation and real-time push notifications. Regarding quality maintenance, after deploying this system, the original model's travel text generation quality score reached [a certain level]. Division, compared to when not deployed The scores are basically consistent, which strongly proves... The regularization mechanism suppresses hallucinations while fully preserving the original model's semantic ability to generate content that conforms to actual travel regulations. Furthermore, this solution demonstrates excellent generalization ability, excelling in cross-model transfer to... hour, The value only decreased The false negative rate only increased This demonstrates that the technology can quickly adapt to various large-scale travel fine-tuning models without repeated training, significantly reducing the cost of technology implementation.

[0105] The above experiments fully validated the detection performance, generalization ability, and real-time advantages of this method in long-text scenarios of travel science popularization. It can effectively identify travel entity-related illusions, significantly reduce the false negative rate, and does not affect the quality of travel text generation of the original model. It fully meets the requirements of high accuracy, high real-time performance, and high practicality in long-text scenarios of travel science popularization. It can be directly applied to various long-text scenarios of travel science popularization, such as online travel consultation, travel science popularization push, and itinerary planning generation, to avoid the travel risks of users caused by travel illusions, improve the reliability of travel science popularization content, and has good practical application value.

Claims

1. A real-time hallucination detection system based on probe technology for large-scale text generation, characterized in that, include: Entity-level annotation module, probe training and optimization module, and real-time detection and intervention module; The entity-level annotation module is connected to the probe training and optimization module via data communication, and the probe training and optimization module interacts with the real-time detection and intervention module via signal communication. The entity-level annotation module is configured to: receive text generated by the target large language model, generate authenticity labels for each token through fact-checking, and construct an annotation dataset containing a comparison between real trajectories and hallucinatory trajectories; The probe training optimization module is configured to: extract the hidden state difference between two adjacent layers of the target large language model during the token generation process as input features, and use the real labels in the labeled dataset as supervision signals to train a lightweight probe to predict the probability that the current token belongs to a hallucination. The real-time detection and intervention module is configured to: extract the hidden state difference in real time when the target large language model performs streaming decoding to generate each token, input it into the trained lightweight probe to calculate the illusion probability, compare the illusion probability with a preset threshold, and perform generation intervention or correction operation according to the comparison result.

2. The real-time hallucination detection system based on probe technology for large-scale text generation according to claim 1, characterized in that, The entity-level annotation module includes a long text generation unit, an entity recognition annotation unit, a label generation unit, and a noise control unit; The output of the long text generation unit is connected to the input of the entity recognition and annotation unit, the output of the entity recognition and annotation unit is connected to the input of the label generation unit, and the output of the label generation unit is connected to the input of the noise control unit. The long text generation unit is configured with a target large language model, the entity recognition and annotation unit has a built-in search-enhanced large language model annotator, and the noise control unit is configured with a span-accurate matching mechanism. The probe training optimization module includes a lightweight probe architecture unit and a composite loss function unit. The regularization unit; the lightweight probe architecture unit is used to output the illusion probability, and the input feature of the lightweight probe is the hidden state difference between adjacent layers of the target large language model; the lightweight probe architecture unit is configured with either a linear probe or a low-rank adapter probe.

3. The real-time hallucination detection system based on probe technology for large-scale text generation according to claim 2, characterized in that, The long text generation unit calls Text generation from long text datasets; The noise control unit, by configuring a span-accurate matching mechanism, eliminates non-aligned samples caused by text truncation or word segmentation deviation, ensuring strong consistency between labels and activation features; The entity recognition and annotation unit uses a large language model annotator to verify the authenticity of entity spans in the generated text and maps the annotation results to the corresponding entities. The sequence is labeled with entity types including names, organizations, and dates. The labeling results are divided into two categories: "support" and "illusion".

4. The real-time hallucination detection system based on probe technology for large-scale text generation according to claim 2, characterized in that, The lightweight probe architecture unit has its input connected to the first... Layer and First The hidden state differencer of the layer has its output connected to a binary classification head. The composite loss function unit employs a function that incorporates entity-level weighted cross-entropy to enhance the probe's ability to capture core entity hallucination signals. The Regularization unit, by introducing Divergence constraints ensure that the lightweight probe preserves the manifold structure of the original semantic representation space of the target large language model when extracting hallucination features.

5. The real-time hallucination detection system based on probe technology for large-scale text generation according to claim 2, characterized in that, The loss function expression of the composite loss function unit is: ; in, The weighted parameter for the travel entity is λt, which is the dynamic annealing coefficient between 0 and 1; The The total loss expression for the regularization unit is: ; Where β is the KL regularization parameter, used to balance the probe detection performance with the original model generation capability.

6. The real-time hallucination detection system based on probe technology for large-scale text generation according to claim 1, characterized in that, The real-time detection and intervention module includes a streaming detection unit, a selective termination unit, a dynamic threshold adjustment unit, and a probability distribution reshaping unit. The streaming detection unit is configured to calculate the hidden state difference in real time and map it to the illusion probability during the decoding cycle of each token, and output the basic threshold. The selective termination unit has a preset threshold. The configuration is set to trigger generation termination and return an "Unverified" flag when the maximum illusion probability within the entity span exceeds the threshold t. The dynamic threshold adjustment unit receives the base threshold output by the streaming detection unit and dynamically scales the soft threshold based on the vertical domain sensitivity of the generated text. With hard threshold ; The probability distribution reshaping unit is configured to, when the hallucination probability exceeds But not exceeding Triggered Orthogonal projection correction of space, when the probability exceeds The process will trigger an abort and return an "unverified" flag.

7. A real-time hallucination detection method for large-model text generation based on probe technology, using the real-time hallucination detection system for large-model text generation based on probe technology as described in any one of claims 1-6, characterized in that, Includes the following steps: Step 1: Using the entity-level annotation module, generate a dataset with comparative features between real fact trajectories and reverse pseudo-fact trajectories. Level dataset; Step 2: Extract the hidden state difference of the target large language model as input, train the lightweight probe through the probe training optimization module, and use the composite loss function and KL regularization to establish the mapping relationship between inter-layer perturbation and factual illusion, so that the trained probe can output token-level illusion probability. Step 3: During the streaming inference process of the target large language model, the hidden state difference is monitored in real time. The token-level illusion probability is calculated through the real-time detection intervention module. The illusion probability is compared with the dynamically set soft threshold and hard threshold. If the illusion probability is lower than the soft threshold, the current token is output normally. If the hallucination probability is between the soft threshold and the hard threshold, then perform Logits space orthogonal projection correction; if the hallucination probability is not lower than the hard threshold, then trigger a halt to generation and return an "unverified" flag.