Large model illusion text detection method based on global-local feature fusion

By constructing local and global text feature extraction modules in a generative large model, text features are fused to detect hallucinatory text, the problem of existing technology relying on external knowledge bases is solved, efficient and reliable hallucinatory text detection is achieved, and user trust is enhanced.

CN120045692APending Publication Date: 2025-05-27QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510097394.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When detecting illusion texts in the output of a generative large model, the prior art relies on external knowledge bases to build and maintain the knowledge bases time-consuming and cannot guarantee timeliness and comprehensiveness, which limits the development of practical applications.

Method used

By constructing a local text feature extraction module and a global text feature extraction module, the local and global features of the text are extracted and fused, and hallucinatory text in the output of a large language model are detected.

Benefits of technology

Effectively detecting illusion text in the output of a generative big model improves the credibility and reliability of the model in the field of low error rates and improves users' trust in the generative big model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045692A_ABST
    Figure CN120045692A_ABST
Patent Text Reader

Abstract

A large model illusion text detection method based on global-local feature fusion relates to the field of text illusion detection in a generative large model, and comprises the following steps: constructing a local text feature extraction module to extract local features of a text, constructing a global text feature extraction module to extract global features of the text, and constructing a local text feature extraction module to extract local features of the text; and then the fusion effect of the local text features and the global text features is controlled by controlling the parameters of feature fusion, so that the model learns richer text feature information, and the illusion text in the output of the generative large model is effectively detected. The credibility and reliability of the large model in the low error rate field are improved, and the credibility of the user on the generative large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detecting text hallucinations in generative large models, and particularly to a method for detecting hallucinated text in large models based on global-local feature fusion. Background Art

[0002] With the booming development of generative large models, a large number of language large models and products have emerged, demonstrating powerful language understanding capabilities and amazing language interaction levels. However, with the powerful capabilities demonstrated by large models in various aspects, an undeniable problem has emerged. There is inevitably fact fabrication in the generation results of large models, and such problems are called hallucination problems. The more powerful the model capabilities, the smoother the generated language, and the more natural the interaction with users, the less likely users are to distinguish the authenticity of the results, and the greater the potential harm of the hallucination problem.

[0003] Currently, the methods for detecting hallucinated text in generative large models mainly include methods based on content consistency, which detect whether the content generated by the large model is consistent with known facts or common sense; detection methods based on external verification, which compare the model output with structured knowledge bases (such as Wikipedia, knowledge graphs, professional databases, etc.) to see if consistent or contradictory information can be found; self-correction based on user feedback, which continuously corrects and improves the model output through user feedback. However, all of these methods require the basis of an external knowledge base. The construction of the knowledge base requires a large amount of time and work, and the content of the knowledge base cannot guarantee timeliness and comprehensiveness, severely restricting the development in practical applications. Therefore, it is crucial to explore a method for detecting large model hallucinations based on the model structure. Summary of the Invention

[0004] In order to overcome the above-mentioned deficiencies in technology, the present invention provides a method for effectively solving the problem of insufficient extraction of text features, which detects hallucinations in the output text of large language models by extracting and fusing global features and local features.

[0005] The technical solution adopted by the present invention to overcome its technical problems is as follows:

[0006] A method for detecting hallucinated text in a large model based on global-local feature fusion, comprising:

[0007] a) Obtain the question text Q input by the user, and input the question text Q into the large language model to output the answer text T;

[0008] b) Preprocess the answer text T to obtain the preprocessed feature vector Hallucination_E 1 ;

[0009] c) Establish a local text feature extraction model composed of a grouped convolution module and a multi-head attention mechanism module;

[0010] d) Input the preprocessed feature vector Hallucination_E 1 into the grouped convolution module of the local text feature extraction model, and output the compressed feature map Hall_Pool_3;

[0011] e) Input the preprocessed feature vector Hallucination_E 1 into the multi-head attention mechanism module of the local text feature extraction model, and output the high-dimensional text feature vector Hallucination_L 2 ;

[0012] f) Establish a global text feature extraction model;

[0013] g) Input the preprocessed feature vector Hallucination_E 1 into the global text feature extraction model, and output the feature map Hallucination_Global;

[0014] h) Input the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L 2 into the local text feature fusion module, and output the feature map Hallucination_Local;

[0015] i) Input the feature map Hallucination_Global and the feature map Hallucination_Local into the local-global text feature fusion module, and output the text fusion feature vector globl_fusion_2;

[0016] j) Determine whether the text T is a real text or a hallucination text according to the text fusion feature vector globl_fusion_2.

[0017] Further, the above large language model is the Qwen2.5-14b-instruct model.

[0018] Further, step b) includes the following steps:

[0019] b-1) Use the Tokenizer method in TensorFlow to convert the answer text T into a sequence T 1 ;

[0020] b-2) Use the pad_sequences function in TensorFlow to pad the sequence T 1Pad to the specified length to obtain the preprocessed text T 2 ;

[0021] b-3) Input the preprocessed text T 2 into the embedding layer, and output the preprocessed feature vector Hallucination_E 1 .

[0022] Preferably, in step b-3), after inputting the preprocessed text T 2 into the embedding layer, map the preprocessed text T 2 to a 128-dimensional space, set the input dimension to 1000, and set the input length to 10. Further, step d) includes the following steps:

[0023] d-1) The grouped convolution module of the local text feature extraction model consists of a first grouped convolution module, a first max pooling layer, a second grouped convolution module, a second max pooling layer, a third grouped convolution module, and a third max pooling layer;

[0024] d-2) The grouped convolution module consists of a first grouped convolution module which is successively composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. Input the preprocessed feature vector Hallucination_E 1 into the first grouped convolution module, and output the feature vector d-3) Input the feature vector into the first max pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_1;

[0025] d-4) The grouped convolution module consists of a second grouped convolution module which is successively composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. Input the preprocessed feature vector Hall_Pool_1 into the second grouped convolution module, and output the feature vector d-5) Input the feature vector into the second max pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_2;

[0026] d-6) The grouped convolution module consists of a third grouped convolution module which is successively composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. Input the preprocessed feature vector Hall_Pool_2 into the third grouped convolution module, and output the feature vector d-7) Input the feature vector into the third max pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_3.

[0027] Further, step e) includes the following steps:

[0028] e-1) The multi-head attention mechanism module of the local text feature extraction model consists of an attention mechanism, a BN layer, a Dropout layer, and a Sigmoid activation function;

[0029] e-2) Input the preprocessed feature vector Hallucination_E 1 into the attention mechanism of the multi-head attention mechanism module, and output the feature map attention_C 1 ;

[0030] e-3) Input the feature map attention_C 1 into the BN layer and Dropout layer of the multi-head attention mechanism module in sequence, and output the feature map attention_C 2 ;

[0031] e-4) Input the feature map attention_C 2 into the Sigmoid activation function of the multi-head attention mechanism module, and output the high-dimensional text feature vector Hallucination_L 2 。

[0032] Further, step g) includes the following steps:

[0033] g-1) The global text feature extraction model consists of a first bidirectional LSTM model, a second bidirectional LSTM model, a BN layer, a Dropout layer, and a ReLU activation function;

[0034] g-2) Input the preprocessed feature vector Hallucination_E 1 into the first bidirectional LSTM model of the global text feature extraction model, and output the feature map g-3) Input the feature map into the second bidirectional LSTM model of the global text feature extraction model, and output the feature map g-4) Input the feature map into the BN layer, Dropout layer, and ReLU activation function of the global text feature extraction model in sequence, and output the feature map Hallucination_Global.

[0035] Further, step h) includes the following steps:

[0036] h-1) The local text feature fusion module consists of a Concat unit, a Dense layer, a ReLU activation function, and a Dropout layer;

[0037] h-2) Input the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L 2 into the Concat unit of the local text feature fusion module and calculate the weighted sum feature map local_fusion_1 of the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L through the formula local_fusion_1 = α 1 Hall_Pool_3 + β 1 Hallucination_L 2 where α 2 and β 1 are both weights, α 1 = 0.4, β 1 = 0.6; 1

[0038] h-3) Input the feature map local_fusion_1 into the Dense layer, ReLU activation function, and Dropout layer of the local text feature fusion module in sequence, and output the feature map Hallucination_Local. Further, step i) includes the following steps:

[0039] i-1) The local-global text feature fusion module consists of a Concat unit, a Dense layer, a ReLU activation function, a global average pooling layer, and a Dropout layer;

[0040] i-2) Input the feature map Hallucination_Global and the feature map Hallucination_Local into the Concat unit of the local-global text feature fusion module and calculate the weighted sum feature map globa_fusion_1 of the feature map Hallucination_Global and the feature map Hallucination_Local through the formula globa_fusion_1 = α 2 Hallucination_Global + β 2 Hallucination_Local, where α 2 and β 2 are both weights, α 2 = 0.4, β 2 = 0.6; i-3) Input the feature map globa_fusion_1 into the Dense layer, ReLU activation function, global average pooling layer, and Dropout layer of the local-global text feature fusion module in sequence, and output the text fusion feature vector globl_fusion_2.

[0041] Further, step j) includes the following steps:

[0042] j-1) Input the text fusion feature vector globl_fusion_2 into the fully connected layer, and output the feature map globl_fusion_2′;

[0043] j-2) Input the feature map globl_fusion_2′ into the Sigmoid function, and output the probability P between (0, 1). When the probability P is greater than 0.5, the text T is determined to be a hallucinated text. When the probability P is less than or equal to 0.5, the text T is determined to be a real text.

[0044] The beneficial effects of the present invention are as follows: By constructing a local text feature extraction module to extract local features of the text, constructing a global text feature extraction module to extract global features of the text, and then controlling the parameters of feature fusion to control the effect of the fusion of local text features and global text features, the model can learn richer text feature information, effectively detect hallucinated texts in the output of the generative large model. Improve the credibility and reliability of the large model in the low error rate field, and enhance the user's trust in the generative large model. Brief Description of the Drawings

[0045] Figure 1 It is the flowchart of the method of the present invention;

[0046] Figure 2 It is the structural diagram of the local text feature extraction model of the present invention;

[0047] Figure 3 It is the structural diagram of the global text feature extraction model of the present invention. Detailed Embodiments

[0048] The following will further describe the present invention in conjunction with the attached Figure 1 、attached Figure 2 、attached Figure 3 Drawings.

[0049] A method for detecting hallucinated texts in a large model based on global-local feature fusion includes:

[0050] a) Obtain the question text Q input by the user, and input the question text Q into the large language model to output the answer text T.

[0051] b) Preprocess the answer text T to obtain the preprocessed feature vector Hallucination_E 1 .

[0052] c) Establish a local text feature extraction model composed of a grouped convolution module and a multi-head attention mechanism module.

[0053] d) Input the preprocessed feature vector Hallucination_E 1 into the grouped convolution module of the local text feature extraction model, and output the compressed feature map Hall_Pool_3.

[0054] e) Input the preprocessed feature vector Hallucination_E 1 into the multi-head attention mechanism module of the local text feature extraction model, and output the high-dimensional text feature vector Hallucination_L 2 .

[0055] f) Establish a global text feature extraction model.

[0056] g) Input the preprocessed feature vector Hallucination_E 1 into the global text feature extraction model, and output the feature map Hallucination_Global.

[0057] h) Input the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L 2 into the local text feature fusion module, and output the feature map Hallucination_Local.

[0058] i) Input the feature map Hallucination_Global and the feature map Hallucination_Local into the local-global text feature fusion module, and output the text fusion feature vector globl_fusion_2.

[0059] j) Determine whether the text T is a real text or a hallucination text according to the text fusion feature vector globl_fusion_2.

[0060] The present invention proposes a method for effectively solving the problem of insufficient text feature extraction, and detecting hallucinations in the output text of large language models by extracting and fusing global features and local features.

[0061] In an embodiment of the present invention, the above-mentioned large language model is the Qwen2.5-14b-instruct model.

[0062] In an embodiment of the present invention, step b) includes the following steps:

[0063] b-1) Use the Tokenizer method in Tensorflow to convert the answer text T into the sequence T 1 .

[0064] b-2) Use the pad_sequences function in TensorFlow to pad the sequence T 1 to the specified length to obtain the preprocessed text T 2 .

[0065] b-3) Input the preprocessed text T 2 into the embedding layer, and output the preprocessed feature vector Hallucination_E 1 .

[0066] In this embodiment, in step b-3), after inputting the preprocessed text T 2 into the embedding layer, the preprocessed text T 2 is mapped to a 128-dimensional space, with the input dimension set to 1000 and the input length set to 10. In an embodiment of the present invention, step d) includes the following steps:

[0067] d-1) The grouped convolution module of the local text feature extraction model consists of a first grouped convolution module, a first max pooling layer, a second grouped convolution module, a second max pooling layer, a third grouped convolution module, and a third max pooling layer.

[0068] d-2) The grouped convolution module consists of the first grouped convolution module, which is sequentially composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. Input the preprocessed feature vector Hallucination_E 1 into the first grouped convolution module, and output the feature vector d-3) Input the feature vector into the first max pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_1.

[0069] d-4) The grouped convolution module consists of the second grouped convolution module, which is sequentially composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. Input the preprocessed feature vector Hall_Pool_1 into the second grouped convolution module, and output the feature vector d-5) Input the feature vector into the second max pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_2.

[0070] d-6) The grouped convolution module consists of the third grouped convolution module, which is sequentially composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. Input the preprocessed feature vector Hall_Pool_2 into the third grouped convolution module, and output the feature vector d-7) Input the feature vector It is input into the third max pooling layer of the grouped convolution module, and the compressed feature map Hall_Pool_3 is output.

[0071] In one embodiment of the present invention, step e) includes the following steps:

[0072] e-1) The multi-head attention mechanism module of the local text feature extraction model is composed of an attention mechanism, a BN layer, a Dropout layer, and a Sigmoid activation function.

[0073] e-2) The preprocessed feature vector Hallucination_E 1 is input into the attention mechanism of the multi-head attention mechanism module, and the feature map attention_C is output 1 .

[0074] e-3) The feature map attention_C 1 is sequentially input into the BN layer and the Dropout layer of the multi-head attention mechanism module, and the feature map attention_C is output 2 .

[0075] e-4) The feature map attention_C 2 is input into the Sigmoid activation function of the multi-head attention mechanism module, and the high-dimensional text feature vector Hallucination_L is output 2 .

[0076] In one embodiment of the present invention, step g) includes the following steps:

[0077] g-1) The global text feature extraction model is composed of a first bidirectional LSTM model, a second bidirectional LSTM model, a BN layer, a Dropout layer, and a ReLU activation function.

[0078] g-2) The preprocessed feature vector Hallucination_E 1 is input into the first bidirectional LSTM model of the global text feature extraction model, and the feature map g-3) The feature map is input into the second bidirectional LSTM model of the global text feature extraction model, and the feature map g-4) The feature map is sequentially input into the BN layer, the Dropout layer, and the ReLU activation function of the global text feature extraction model, and the feature map Hallucination_Global is output. In one embodiment of the present invention, step h) includes the following steps:

[0079] h-1) The local text feature fusion module consists of a Concat unit, a Dense layer, a ReLU activation function, and a Dropout layer.

[0080] h-2) Input the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L 2 into the Concat unit of the local text feature fusion module, and calculate the weighted sum feature map local_fusion_1 of the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L through the formula local_fusion_1 = α 1 Hall_Pool_3 + β 1 Hallucination_L 2 where α 2 and β 1 are both weights, α 1 = 0.4, β 1 = 0.6. 1

[0081] h-3) Input the feature map local_fusion_1 into the Dense layer, ReLU activation function, and Dropout layer of the local text feature fusion module in sequence, and output the feature map Hallucination_Local. In an embodiment of the present invention, step i) includes the following steps:

[0082] i-1) The local-global text feature fusion module consists of a Concat unit, a Dense layer, a ReLU activation function, a global average pooling layer, and a Dropout layer.

[0083] i-2) Input the feature map Hallucination_Global and the feature map Hallucination_Local into the Concat unit of the local-global text feature fusion module, and calculate the weighted sum feature map gobal_fusion_1 of the feature map Hallucination_Global and the feature map Hallucination_Local through the formula gobal_fusion_1 = α 2 Hallucination_Global + β 2 Hallucination_Local, where α 2 and β 2 are both weights, α 2 = 0.4, β 2 = 0.6.

[0084] i-3) Input the feature map global_fusion_1 into the Dense layer, ReLU activation function, global average pooling layer, and Dropout layer of the local-global text feature fusion module in sequence, and output the text fusion feature vector globl_fusion_2.

[0085] In an embodiment of the present invention, step j) includes the following steps:

[0086] j-1) Input the text fusion feature vector globl_fusion_2 into the fully connected layer, and output the feature map globl_fusion_2'.

[0087] j-2) Input the feature map globl_fusion_2' into the Sigmoid function, and output the probability P between (0, 1). When the probability P is greater than 0.5, the text T is determined to be an hallucination text. When the probability P is less than or equal to 0.5, the text T is determined to be a real text.

[0088] To verify the reliability of this method, the model proposed by this method and the existing baseline model SLM were compared using the ROUGE score as shown in Table 1.

[0089] Table 1 compares our model with the baseline model SLM using the Rouge score on the HaluQA dataset.

[0090] Model \ Scoring Criteria Precision Recall F1 Score SLM Baseline Model 0.962 0.418 0.583 The Method of the Present Invention 0.975 0.450 0.616

[0091] In our experimental method, we compared with the current state-of-the-art hallucination text detection model SLM. The ROUGE score calculates precision, recall, and F1 score, and also considers the presence and order of n-grams (continuous word sequences). The higher the ROUGE score, the higher the accuracy of the model in detecting hallucination texts, indicating that the model has a stronger ability to distinguish between hallucination texts and real texts, and a higher reliability in judging the authenticity of texts. It also means that the model can more effectively filter out false information when processing texts, ensuring the quality and credibility of the output content. In terms of the Precision score, our method is better than SLM, with an absolute gain of 5%. This shows that this method has certain advantages, and the advantages are significantly improved in terms of accuracy.

[0092] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A large model hallucination text detection method based on global-local feature fusion, characterized in that: include: a) Obtain the question text Q input by the user, input the question text Q into the large language model, and output the answer text T; b) Preprocess the answer text T to obtain the preprocessed feature vector Hallucination_E1; c) Establish a local text feature extraction model consisting of a grouped convolution module and a multi-head attention mechanism module; d) Input the preprocessed feature vector Hallucination_E1 into the grouped convolution module of the local text feature extraction model, and output the compressed feature map Hall_Pool_3; e) Input the preprocessed feature vector Hallucination_E1 into the multi-head attention mechanism module of the local text feature extraction model, and output the high-dimensional text feature vector Hallucination_L2; f) Establishing a global text feature extraction model; g) Input the preprocessed feature vector Hallucination_E1 into the global text feature extraction model, and output the feature map Hallucination_Global; h) Input the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L2 into the local text feature fusion module, and output the feature map Hallucination_Local; i) Input the feature map Hallucination_Global and the feature map Hallucination_Local into the local-global text feature fusion module, and input the obtained text fusion feature vector globl_fusion_2; j) Determine whether the text T is real text or hallucinated text based on the text fusion feature vector globl_fusion_2.

2. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: The large language model is the Qwen2.5-14b-instruct model.

3. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: Step b) comprises the following steps: b-1) Use the Tokenizer method in Tensorflow to convert the answer text T into a sequence T1; b-2) Use the pad_sequences function in Tensorflow to pad the sequence T1 to the specified length to obtain the preprocessed text T2; b-3) Input the preprocessed text T2 into the embedding layer and output the preprocessed feature vector Hallucination_E1.

4. The large model hallucination text detection method based on global-local feature fusion according to claim 3 is characterized by: In step b-3), the preprocessed text T2 is input into the embedding layer and then mapped to a 128-dimensional space, with the input dimension set to 1000 and the input length set to 10.

5. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: Step d) comprises the following steps: d-1) The grouped convolution module of the local text feature extraction model is composed of a first grouped convolution module, a first maximum pooling layer, a second grouped convolution module, a second maximum pooling layer, a third grouped convolution module, and a third maximum pooling layer; d-2) The grouped convolution module is composed of the first grouped convolution module, which is composed of a grouped convolution layer, a pooling layer, and a ReLU activation function in sequence. The preprocessed feature vector Hallucination_E1 is input into the first grouped convolution module, and the feature vector is output. d-3) The feature vector Input to the first maximum pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_1; d-4) The grouped convolution module is composed of the second grouped convolution module, which is composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. The preprocessed feature vector Hall_Pool_1 is input into the second grouped convolution module, and the feature vector is output. d-5) The feature vector Input to the second maximum pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_2; d-6) The grouped convolution module is composed of the third grouped convolution module, which is composed of a grouped convolution layer, a pooling layer, and a ReLU activation function. The preprocessed feature vector Hall_Pool_2 is input into the third grouped convolution module, and the feature vector is output. d-7) The feature vector Input into the third maximum pooling layer of the grouped convolution module, and output the compressed feature map Hall_Pool_3.

6. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: Step e) comprises the following steps: e-1) The multi-head attention mechanism module of the local text feature extraction model consists of an attention mechanism, a BN layer, a Dropout layer, and a Sigmoid activation function; e-2) Input the preprocessed feature vector Hallucination_E1 into the attention mechanism of the multi-head attention mechanism module, and output the feature map attention_C1; e-3) Input the feature map attention_C1 into the BN layer and Dropout layer of the multi-head attention mechanism module in sequence, and output the feature map attention_C2; e-4) Input the feature map attention_C2 into the Sigmoid activation function of the multi-head attention mechanism module, and output the high-dimensional text feature vector Hallucination_L2.

7. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: Step g) comprises the following steps: g-1) The global text feature extraction model consists of a first bidirectional LSTM model, a second bidirectional LSTM model, a BN layer, a Dropout layer, and a ReLU activation function; g-2) Input the preprocessed feature vector Hallucination_E1 into the first bidirectional LSTM model of the global text feature extraction model, and output the feature map g-3) The feature map Input into the second bidirectional LSTM model of the global text feature extraction model, and output the feature map g-4) The feature map It is input into the BN layer, Dropout layer, and ReLU activation function of the global text feature extraction model in sequence, and the output is the feature map Hallucination_Global.

8. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: Step h) comprises the following steps: h-1) The local text feature fusion module consists of a Concat unit, a Dense layer, a ReLU activation function, and a Dropout layer; h-2) Input the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L2 into the Concat unit of the local text feature fusion module, and calculate the weighted sum feature map local_fusion_1 of the compressed feature map Hall_Pool_3 and the high-dimensional text feature vector Hallucination_L2 by the formula local_fusion_1=α1Hall_Pool_3+β1Hallucination_L2, where α1 and β1 are weights, α1=0.4, β1=0.6; h-3) Input the feature map local_fusion_1 into the Dense layer, ReLU activation function, and Dropout layer of the local text feature fusion module in sequence, and output the feature map Hallucination_Local.

9. The large model hallucination text detection method based on global-local feature fusion according to claim 1 is characterized in that: Step i) comprises the following steps: i-1) The local-global text feature fusion module consists of a Concat unit, a Dense layer, a ReLU activation function, a global average pooling layer, and a Dropout layer; i-2) Input the feature map Hallucination_Global and the feature map Hallucination_Local into the Concat unit of the local-global text feature fusion module, and calculate the weighted sum feature map gobal_fusion_1 of the feature map Hallucination_Global and the feature map Hallucination_Local through the formula gobal_fusion_1=α2Hallucination_Global+β2Hallucination_Local, where α2 and β2 are weights, α2=0.4, β2=0.6; i-3) Input the feature map gobal_fusion_1 into the Dense layer, ReLU activation function, global average pooling layer, and Dropout layer of the local-global text feature fusion module in sequence, and output the text fusion feature vector globl_fusion_2.

10. The large model hallucination text detection method based on global-local feature fusion according to claim 1, characterized in that: Step j) comprises the following steps: j-1) Input the text fusion feature vector globl_fusion_2 into the fully connected layer, and output the feature map globl_fusion_2′; j-2) Input the feature map globl_fusion_2′ into the Sigmoid function, and output a probability P between (0,1). When the probability P is greater than 0.5, the text T is judged to be a hallucinated text. When the probability P is less than or equal to 0.5, the text T is judged to be a real text.