Information authenticity judgment method based on attention mechanism
Through the information authenticity judgment method based on attention mechanism, multiple reasons supporting text content are generated and evaluated, and the problem of difficulty in extracting background knowledge and obtaining social background information in the prior art is solved, and more efficient and accurate judgment of information authenticity is achieved.
Patent Information
- Application Number
- CN202510702897.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing methods for judging information authenticity are difficult to effectively extract background knowledge and potential clues in text content, and obtaining comments, publisher information and social backgrounds during information dissemination requires additional resources and time, which reduces the applicability of the method.
The information authenticity judgment method based on the attention mechanism is adopted, and by generating multiple first reasons that support the text content as real information and multiple second reasons that support the false information, the interactive attention mechanism and the reverse attention mechanism are used to evaluate the importance of each reason, thereby training the information authenticity judgment model.
This method can effectively overcome the problem that potential background information in text content is difficult to obtain, improve the accuracy and applicability of information authenticity judgment, and realize effective detection in an environment where false information is spread rapidly.
Smart Images

Figure CN120234415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information detection, and particularly to a method, device, equipment and medium for judging information authenticity based on an attention mechanism. Background Art
[0002] With the continuous development of digital technology, online communication on social platforms has become increasingly common among people. However, this has also led to a rapid increase in false information, which is widely spread. Since false information is usually deliberately fabricated to attract public attention, it has affected people's daily lives and social stability. Therefore, there is an urgent need to detect false information.
[0003] Existing detection methods judge information authenticity by extracting semantic and emotional features from text content. However, text content usually contains rich background knowledge and potential clues, and it is difficult for a small model to extract these clues. In addition, during the dissemination of text information, although comments, publisher information, and social background, etc. can be used to assist in detecting false information, these knowledge and experiences require additional resources to obtain and time to accumulate, reducing the applicability of such methods. Therefore, there is still no method that can effectively judge information authenticity currently. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, equipment and medium for judging information authenticity based on an attention mechanism, which can solve the problem that effective judgment of information authenticity cannot be achieved currently.
[0005] To solve the above technical problems, an embodiment of the present invention provides a method for judging information authenticity based on an attention mechanism, including the following steps: Obtain sample text content; Generate multiple first reasons supporting the sample text content as true information and multiple second reasons supporting the sample text content as false information; wherein, the number of first reasons and second reasons is the same; Use the sample text content, multiple first reasons and multiple second reasons to train an information authenticity judgment model through the following steps: Encode the sample text content, multiple first reasons, and multiple second reasons respectively to obtain the sample text encoding, multiple first reason encodings, and multiple second reason encodings; adopt an interactive attention mechanism to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively to obtain multiple aggregated encodings; adopt a reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding to evaluate the importance of each aggregated encoding; wherein, the importance of the aggregated encoding is used to represent the contribution degree of the corresponding first reason to supporting the sample text content as true information, or the contribution degree of the second reason to supporting the sample text content as false information; Input the target text content, multiple first reasons corresponding to the target text content, and multiple second reasons into the information authenticity judgment model to detect the authenticity of the target text content.
[0006] Optionally, the information authenticity judgment model is trained using the sample text content, multiple first reasons, and multiple second reasons through the following loss functions: binary classification main loss function, contrast loss function, reason discrimination loss function, and KL divergence loss function; Among them, the binary classification main loss function is constructed based on whether the sample text content is false information, the contrast loss function is constructed based on the similarity between the sample text content and the first reason and the second reason respectively, the reason discrimination loss function is constructed based on the contribution degree of the first reason to supporting the sample text content as true information and the contribution degree of the second reason to supporting the sample text content as false information, and the KL divergence loss function is constructed based on the difference between the first reason and the second reason.
[0007] Optionally, the binary classification main loss function is: ; ; In the formula, represents the cross-entropy loss, is a multi-layer perceptron; y is the label of the sample text content, indicating whether the sample text content is false information; represents the weight of, the sample text encoding enhanced by the self-attention mechanism, is the aggregated encoding formed by aggregating the sample text encoding with the first reason encoding and the second reason encoding respectively; The contrast loss function is: ; In the formula, is an exponential function. When the label of the sample text content is true information, The aggregated code formed by aggregating the sample text code and the first reason code. When the label of the sample text content is false information, The aggregated code formed by aggregating the sample text code and the second reason code, The elements in are opposite to the elements in It represents a random sampling operation, which means randomly taking an instance from the set; It is used to measure the cosine similarity of two vectors, is the temperature parameter, which is used to adjust the kurtosis of the softmax distribution; The reason discrimination loss function is: ; In the formula, CAT means concatenating or ; The aggregated code formed by aggregating the sample text code and the first reason code, The aggregated code formed by aggregating the sample text code and the second reason code; The KL divergence loss function is: ; In the formula, KL means calculating the KL divergence.
[0008] Optionally, there are multiple sample text contents, and the information authenticity judgment model is trained through the following steps: For each sample text content, a base learner is trained by using the sample text content and the corresponding multiple first reasons and multiple second reasons; The base learners of multiple sample text contents are integrated by using a boosting strategy to obtain the information authenticity judgment model.
[0009] Optionally, the integrating the base learners of multiple sample text contents by using a boosting strategy to obtain the information authenticity judgment model includes: S1. Generate the initial base learners of multiple sample text contents, and initialize the weight of each sample text content as , where X is the set of multiple sample text contents; S2. Use the sample text content with weight to train the nth base learner , and calculate the error. The error calculation formula of the nth base learner is: ; In the formula, is the ith sample text content, N is the number of sample text contents, It is an indicator function. When the input is True, the output is 1; when the input is False, the output is 0. is the label of the i-th sample text content, indicating whether the sample text content is false information or true information; If the error is greater than 0.5, terminate the boosting strategy; otherwise, enter S3. S3. Update the weights of each sample text content according to the error. The update formula is as follows: ; In the formula, , is an exponential function; Normalize the updated weights: ; In the formula, is the updated weight; S4. Repeat S2 and S3 N times to obtain N trained base learners; The information authenticity judgment model is: ; In the formula, represents the ensemble of n base learners.
[0010] Optionally, generating multiple first reasons for supporting that the sample text content is true information and multiple second reasons for supporting that the sample text content is false information includes: The system prompt for constructing the large language model is: Please understand the following text content and give multiple reasons for why it is true information and multiple reasons for why it is false information. Additionally, since the text content is scraped from the internet, there may be some noise or defects, please ignore them; Input the sample text content and the system prompt into the large language model to obtain multiple first reasons for supporting that the sample text content is true information and multiple second reasons for supporting that the sample text content is false information.
[0011] Optionally, there are at least three first reasons and three second reasons respectively.
[0012] An embodiment of the present invention also provides an information authenticity judgment device based on an attention mechanism, including: A sample acquisition module for acquiring sample text content; A sample processing module for generating multiple first reasons for supporting that the sample text content is true information and multiple second reasons for supporting that the sample text content is false information; wherein, the number of first reasons and second reasons is the same; A model training module, which is used to train an information authenticity judgment model by using sample text content, multiple first reasons, and multiple second reasons through the following steps: Encode the sample text content, multiple first reasons, and multiple second reasons respectively to obtain a sample text encoding, multiple first reason encodings, and multiple second reason encodings; adopt an interactive attention mechanism to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively to obtain multiple aggregated encodings; adopt a reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding to evaluate the importance of each aggregated encoding; wherein, the importance of the aggregated encoding is used to represent the contribution degree of the corresponding first reason to supporting the sample text content as true information, or the contribution degree of the second reason to supporting the sample text content as false information; A text detection module, which is used to input the target text content, multiple first reasons, and multiple second reasons corresponding to the target text content into the information authenticity judgment model to detect the authenticity of the target text content.
[0013] An embodiment of the present invention further provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned information authenticity judgment method based on the attention mechanism.
[0014] An embodiment of the present invention further provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the above-mentioned information authenticity judgment method based on the attention mechanism is implemented.
[0015] The information authenticity judgment method based on the attention mechanism provided by the present invention has at least the following beneficial effects: For text content, by generating multiple reasons to support it as true information and multiple reasons to support it as false information, more potential background information about the text content is obtained. However, not all the generated reasons contribute equally to detecting whether the text content is false information. Therefore, an interactive attention mechanism and a reverse attention mechanism from the text content to the reasons are adopted to obtain the importance of each reason, and the contribution degree of each reason to supporting the sample text content as true information or false information is obtained. The information authenticity judgment model trained based on this can be used to judge the authenticity of the text content to be detected, which can overcome the problem that it is difficult to obtain potential background information in the text content, and has better applicability on the premise of effectively improving the accuracy of text content information authenticity judgment. Description of the Drawings
[0016] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings, and these exemplary illustrations do not constitute a limitation on the embodiments.
[0017] Figure 1 is a flowchart of a method for judging information authenticity based on an attention mechanism provided according to an embodiment of the present invention; Figure 2 is a comparison chart of the detection results of different language models provided according to an embodiment of the present invention; Figure 3 is a comparison chart of the detection results using different numbers of reasons provided according to an embodiment of the present invention; Figure 4 is a comparison chart of the detection results of different models under different scarcities of false information provided according to an embodiment of the present invention; Figure 5 is a comparison chart of the detection results of the corresponding models using the reasons generated by different large language models provided according to an embodiment of the present invention. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be elaborated in detail below with reference to the drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present invention, many technical details are provided to help readers better understand the present invention. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present invention can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present invention. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.
[0019] An embodiment of the present invention relates to a method for judging information authenticity based on an attention mechanism. The specific process of the method for judging information authenticity based on an attention mechanism in this embodiment can be as Figure 1 shown and includes: Step 101, obtaining the sample text content.
[0020] Step 102, generating a plurality of first reasons supporting the sample text content as true information and a plurality of second reasons supporting the sample text content as false information; wherein, the number of the first reasons and the second reasons is the same.
[0021] Step 103, using the sample text content, a plurality of first reasons, and a plurality of second reasons to train an information authenticity judgment model through the following steps: Encode the sample text content, multiple first reasons, and multiple second reasons respectively to obtain the sample text encoding, multiple first reason encodings, and multiple second reason encodings; use an interactive attention mechanism to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively to obtain multiple aggregated encodings; use a reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding to evaluate the importance of each aggregated encoding; wherein, the importance of the aggregated encoding is used to represent the contribution degree of the corresponding first reason to supporting the sample text content as true information, or the contribution degree of the second reason to supporting the sample text content as false information.
[0022] Step 104, input the target text content, multiple first reasons corresponding to the target text content, and multiple second reasons into the information authenticity judgment model to detect the authenticity of the target text content.
[0023] In this embodiment, for the text content, by generating multiple reasons to support its authenticity and multiple reasons to support its falsehood, more potential background information about the text content is obtained. However, not all the generated reasons contribute equally to detecting whether the text content is false. Therefore, an interactive attention mechanism and a reverse attention mechanism from the text content to the reasons are used to obtain the importance of each reason, and the contribution degree of each reason to supporting the sample text content as true information or false information is obtained. The information authenticity judgment model trained based on this can be used to judge the authenticity of the text content to be detected, which can overcome the problem of difficult access to potential background information in the text content and has better applicability on the premise of effectively improving the accuracy of text content information authenticity judgment.
[0024] The implementation details of the information authenticity judgment method based on the attention mechanism in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0025] In step 102, select a large language model (LLM), such as GPT-4.0, Doubao, etc., and construct the system prompt words for the large language model as: "Please understand the following text content and give multiple reasons to explain why it is true information and multiple reasons to explain why it is false information. In addition, since the text content is crawled from the Internet, there may be some noises or defects, please ignore them". Then for each text content, input it and the system prompt words into the large language model to obtain multiple first reasons to support the text content as true information and multiple second reasons to support the text content as false information. Among them, the number of first reasons and second reasons is the same, and at least three first reasons and second reasons are generated respectively.
[0026] In this embodiment, taking three cases for the first reason and the second reason respectively as examples, the following specific description is given.
[0027] In step 103, first, the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model is used to encode the sample text content, multiple first reasons, and multiple second reasons respectively. Assuming the given sample text content C, and the three generated first reasons , three second reasons , encode them respectively as follows: ; In the formula, , , respectively represent the encoding vectors of the sample text content C, three first reasons supporting the sample text content as true information and three second reasons supporting the sample text content as false information , that is, the sample text encoding, the first reason encoding, and the second reason encoding. At this time, n is the number of the first reasons and the number of the second reasons, n = 1, 2, 3.
[0028] Then, the interactive attention mechanism is adopted to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively as follows: ; In the formula, , and ; , , , are all preset learning parameters; d represents the dimension of the encoding vector (that is, the dimensions of the sample text encoding, the first reason encoding, and the second reason encoding), then is the aggregated encoding formed by aggregating the sample text encoding with the first reason encoding and the second reason encoding respectively.
[0029] Six different aggregated encodings are obtained through the above aggregation operation, which aggregate six reasons respectively.
[0030] Next, the reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding is adopted to evaluate the importance of each aggregated encoding : ; In the formula, the multi-layer perceptron (MLP) converts the attention vector into a weight coefficient , the interpretable weight coefficient can adjust the reason-attention-based text embedding weights, and based on this, evaluate the importance of each aggregated encoding .
[0031] Since not all generated reasons contribute equally to the detection, this embodiment designs an interactive attention module to dynamically focus on the original text and each reason, including a reason aggregation strategy and an interpretable weight adjustment strategy. The weight coefficient reflects the importance of each reason, and can help detect false texts while maintaining interpretability, improving the effectiveness and transparency of the detection process.
[0032] In one example, after performing the above interactive attention module to aggregate the text content and reasons, and dynamically re-weighting the six text embedded encodings (i.e., aggregated encodings) containing different reasons (i.e., evaluating the importance of each aggregated encoding), a contrastive learning strategy is also adopted to improve the encoding quality of the text content.
[0033] Specifically, first use the self-attention mechanism to enhance the text relationship in the text content: ; Then, based on the label of the text content being true information (or false information), minimize the distance between the text content and the reasons supporting the text content being true information (or false information), and maximize the distance between the text content and the reasons supporting the text content being false information (or true information).
[0034] The calculation method of the contrastive loss function is as follows: ; In the formula, is the exponential function. When the label of the sample text content is true information, is the aggregated encoding formed by aggregating the sample text encoding and the first reason encoding, that is, , when the label of the sample text content is false information, is the aggregated encoding formed by aggregating the sample text encoding and the second reason encoding, that is, , and The elements in are opposite to the elements in ; is the random sampling operation, indicating randomly taking an instance from the set; is used to measure the cosine similarity of two vectors, is the temperature parameter, used to adjust the kurtosis of the softmax distribution.
[0035] In a specific implementation, the information authenticity judgment model is trained using sample text content, multiple first reasons, and multiple second reasons through the following loss functions: a binary classification main loss function, the above-mentioned contrast loss function, a reason discrimination loss function, and a Kullback-Leibler Divergence (KL divergence) loss function. Among them, the binary classification main loss function is constructed based on whether the sample text content is false information, the contrast loss function is constructed based on the similarity between the sample text content and the first reason and the second reason respectively, the reason discrimination loss function is constructed based on the contribution degree of the first reason to supporting the sample text content as true information and the contribution degree of the second reason to supporting the sample text content as false information, and the KL divergence loss function is constructed based on the difference between the first reason and the second reason.
[0036] That is, in this embodiment, a comprehensive loss function composed of the above loss functions is adopted. Train the information authenticity judgment model, which includes a binary classification main loss function, a contrast loss function for improving the encoding quality of the original text content, a reason discrimination loss function for aligning reasons with labels, and a KL divergence loss function for distinguishing reasons that support true information and reasons that support false information, as follows: First, calculate the binary classification main loss function. , through a dynamic concatenation-based attention weight calculation module, calculate and normalize the weights of seven main embedded encodings. The weights are calculated as follows: ; In the formula, is a learnable matrix parameter in the attention model, is the concatenated attention coefficient of , is the concatenated attention coefficient of , , is a learnable matrix parameter in the attention mechanism, is a learnable vector parameter, and T represents the transpose operation.
[0037] After obtaining the weighted sum of the embedded encodings, then pass it through a multi-layer perceptron (MLP) to obtain the final prediction result, and calculate the main loss function using cross-entropy loss. The calculation method is as follows: ; In the formula, represents the cross-entropy loss, y is the label of the sample text content, indicating whether the sample text content is false information.
[0038] Then calculate the reason discrimination loss function. , the reasons supporting whether the text content is false can be used as auxiliary evidence for discrimination, and the calculation formula is as follows: ; In the formula, CAT represents concatenating or , is the aggregated code formed by aggregating the sample text code and the first reason code, is the aggregated code formed by aggregating the sample text code and the second reason code.
[0039] Next, to enhance the difference between the two types of reasons supporting whether the text content is false, the difference between them is calculated, that is, the KL divergence loss function of the reasons : ; In the formula, KL represents calculating the KL divergence.
[0040] The contrast loss function is referred to step 103.
[0041] The total loss function is the weighted sum of the above four loss functions, and the calculation formula is as follows: ; In the formula, , , are the weight hyperparameters of each loss function.
[0042] In this embodiment, the pre-trained language model BERT is specifically used to obtain the embedded codes of the text content and its related reasons, and through the content-reason interaction attention module, the text content and each reason are aggregated and re-weighted, so that the model can perceive the potential background information in the text content. Then, the contrast learning module based on sampling is used to improve the encoding quality of the original text content. Finally, a composite loss function is designed to distinguish the embedded codes of the reasons supporting truth and falsehood, and guide the model to be trained, so that the accuracy of the judgment method provided by the present invention is higher.
[0043] In some embodiments, there are multiple obtained sample text contents, and the information authenticity judgment model is trained through the following steps: for each sample text content, a base learner is trained by using the sample text content and the corresponding multiple first reasons and multiple second reasons, and then the base learners of the multiple sample text contents are integrated by using the boosting strategy to obtain the information authenticity judgment model.
[0044] The following details how to integrate multiple base learners by using the boosting strategy, which specifically includes the following steps: S1. Generate initial base learners for multiple sample text contents, and initialize the weight of each sample text content as , where X is the set of multiple sample text contents; S2. Use the sample text content with weight to train the nth base learner , and calculate the error. The error calculation formula for the nth base learner is: ; In the formula, is the ith sample text content, N is the number of sample text contents, is the indicator function. When the input is True, the output is 1; when the input is False, the output is 0. is the label of the ith sample text content, indicating whether the sample text content is false information or true information; If the error is greater than 0.5, terminate the boosting strategy; otherwise, enter S3; S3. Update the weight of each sample text content according to the error. The update formula is as follows: ; In the formula, , is the exponential function; Normalize the updated weights: ; In the formula, is the updated weight; S4. Repeat S2 and S3 for N times to obtain N trained base learners; The information authenticity judgment model is: ; In the formula, represents the ensemble of n base learners.
[0045] At this time, in the specific implementation, if the label dataset is unbalanced or resources permit, the text content to be detected can be input into the trained false detection model , and the category of the text content can be obtained. 1 represents that the text content is false information, and 0 represents that the text content is true information.
[0046] .
[0047] If low overhead is required, the text content to be detected can be directly input into the above base learner to obtain a text representation vector that fuses seven main encodings, and then the predicted text category can be obtained through MLP: = .
[0048] This embodiment provides a flexible framework that can encode text using only base learners and predict the text category (i.e., whether it is false information) through an MLP, which can meet the low overhead. The text category can also be predicted by a model integrated with a boosting strategy to achieve higher accuracy.
[0049] In summary, the method for judging information authenticity based on the attention mechanism of the present invention has the following beneficial effects: (1) It can automatically detect false text information in the information system and can effectively judge even when the data of true information and false information is unbalanced.
[0050] (2) Use the large language model to generate the reasons for the text content to be true information or false information, which includes the background information of the text, and can effectively improve the accuracy of model detection.
[0051] (3) Design an interpretable text detection base learner that can dynamically integrate text content and reason information and improve the encoding quality. This lightweight base learner can directly detect the authenticity of text and outperform the current optimal method in terms of detection accuracy.
[0052] (4) The false detection model designed by the present invention, which is integrated by multiple base learners, hereinafter referred to as the Large Language Model-assisted Fake News Detection method with Adaptive Boosting (LFND-AB), is an adaptive boosting framework that solves the problem of unbalanced dataset labels for text detection. And it can dynamically adjust the sample weights to improve the detection performance, significantly outperforming the existing methods. Both the base learner model and the LFND-AB model have excellent performance and can be flexibly selected in various situations.
[0053] Next, verify the method for judging information authenticity based on the attention mechanism of the present invention: Figure 2 These are the detailed results of the evaluation metrics for detecting false text by the LFND-AB model of the method of the present invention and other comparison methods in Chinese and English datasets. Figure 2It can be seen that the proposed lightweight model (i.e., the interpretable base learner LFND) has the best detection performance. Compared with the optimal baseline, the Adaptive Rationale Guidance network for fake news detection (ARG), LFND has an average improvement of 1.2% on the Chinese dataset and 1.1% on the English dataset. There are three reasons for the improvement: (1) The reasons for the true and false supported text information generated by the LLM provide more detailed background knowledge from a comprehensive perspective. In addition, by requiring three reasons for each aspect, the bias of the given reasons can be minimized and the robustness can be enhanced. (2) The sampling-based contrast learning not only introduces randomness to avoid overfitting, but also improves the quality of the text content embedded coding by narrowing the distance between positive pairs and pushing away the distance between negative pairs. (3) The joint loss function comprehensively considers the roles of various embeddings in the fake news detection task, ensures the quality of all embeddings, and thus optimizes the model training.
[0054] Figure 3 It is a comparison chart of the ablation experiment results of different reason settings and the removal of each key module. LFND_1, LFND_2, LFND_R, and LFND_F represent using one pair of reasons, two pairs of reasons, only reasons supporting true information, and only reasons supporting false information, respectively. LFND_W is the model that uses a matrix of all 1s to replace the interpretable weight matrix, LFND- 、LFND- 、LFND- 、LFND- represent the models with the main loss function, contrast loss function, reason discriminant loss function, and KL divergence loss function removed, respectively. LFND-AB_AVG represents the model that keeps the sample weights of each base learner unchanged. From Figure 3It can be seen that the model performance is poor when only a pair of reasons are used. After the reasons are increased from one pair to two pairs, the performance of the model improves slightly, indicating that combining multiple reasons can enhance the model's ability to effectively detect false information. In addition, using only the reasons supporting false information leads to a slightly lower F1 score, suggesting that focusing only on the reasons supporting false information does not provide sufficient information to effectively distinguish text information. Similarly, using only the reasons supporting true information also results in a further performance decline, indicating that the amount of information using only the reasons for truthfulness is insufficient to accurately detect false information. Moreover, the results also show that removing key components from the model architecture leads to varying degrees of performance degradation. In particular, LFND_W, which uses a matrix of all 1s to replace the interpretable weight matrix, experiences a significant performance decline. This demonstrates the importance of the interpretable weight matrix in guiding the model's decision-making process. Deleting different loss terms also causes a slight performance decline, indicating that the loss terms contribute to model training. LFND-AB_AVG, which keeps the weight of each sample unchanged in the new base learner, has slightly lower performance than the complete LFND-AB, suggesting that the adaptive weight adjustment in LFND-AB helps improve performance by focusing on the samples with the most information. In summary, each component, including the interpretable weight matrix, reason discrimination, contrast learning, and adaptive weight adjustment, plays an important role in improving model performance.
[0055] Figure 4 Shows the performance of different methods under different scarcities of false information. Even when the number of false information is reduced to only 1 / 3 of the original dataset, the performance of our proposed method drops very little, indicating its robustness in the case of label imbalance. Several factors contribute to these results. For the base learner LFND, the contrast learning module improves the discriminative ability of the embedded encoding by ensuring that the representations of false and true information remain separable even when the training samples are reduced. For the complete LFND-AB, the boosting strategy dynamically reweights the misclassified samples during training, thus paying more attention to the correctly classified instances. This procedure allows the model to focus on the classes of a few instances, thereby avoiding the performance decline caused by imbalanced data. Even when the available number of false information is reduced, LFND-AB can ensure that the model can learn to effectively classify the boundaries.
[0056] To study the impact of using different large language models to generate reasons, experiments were conducted on the Chinese dataset using ChatGPT-3.5, ChatGPT-4.0, and Doubao. Each LLM can generate reasons supporting both true and false text content. Figure 5 The results shown compare the performance of LFND and LFND-AB when using the reasons generated by these different LLMs. Figure 5The experimental results in show that the performance of ChatGPT-4.0 is slightly higher than that of ChatGPT-3.5 and Doubao. This improvement may be due to the fact that ChatGPT-4.0 can generate more detailed and context-rich justifications, which in turn can provide a more fine-grained representation for the model. However, this has little impact on detecting the final performance of the model. This may be because existing modules (such as contrastive learning and attention mechanisms) can effectively and selectively integrate and refine the information provided by the LLM, thus minimizing the performance differences caused by choosing the LLM.
[0057] In summary, the method for detecting false texts assisted by an interpretable large language model provided by the present invention utilizes the large language model to fully mine the background information of the text, improving the detection accuracy and interpretability of the model. And an adaptive boosting strategy is adopted to integrate the base learners, enhancing the robustness of the model in unbalanced data scenarios. Compared with existing methods, a higher detection accuracy is achieved, which makes the method for judging the authenticity of information based on large model prompts of the present invention have practical guiding significance.
[0058] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of the present invention; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are all within the protection scope of this invention.
[0059] Another embodiment of the present invention relates to an information authenticity judgment device based on the attention mechanism. The implementation details of the information authenticity judgment device based on the attention mechanism in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution. The information authenticity judgment device based on the attention mechanism in this embodiment includes: A sample acquisition module for acquiring the content of the sample text; A sample processing module for generating multiple first justifications supporting the sample text content as true information and multiple second justifications supporting the sample text content as false information; wherein, the number of the first justifications and the second justifications is the same; A model training module for training an information authenticity judgment model through the following steps by using the sample text content, multiple first justifications, and multiple second justifications: Encode the sample text content, multiple first reasons, and multiple second reasons respectively to obtain the sample text encoding, multiple first reason encodings, and multiple second reason encodings; adopt an interactive attention mechanism to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively to obtain multiple aggregated encodings; adopt a reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding to evaluate the importance of each aggregated encoding; wherein, the importance of the aggregated encoding is used to represent the contribution degree of the corresponding first reason to supporting the sample text content as true information, or the contribution degree of the second reason to supporting the sample text content as false information. A text detection module, configured to input the target text content, multiple first reasons, and multiple second reasons corresponding to the target text content into the information authenticity judgment model to detect the authenticity of the target text content.
[0060] It is not difficult to find that this embodiment is a device embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiment are still valid in this embodiment. For the sake of reducing repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0061] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or can be implemented by a combination of multiple physical units. In addition, in order to highlight the innovative part of the present invention, units not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0062] Another embodiment of the present invention relates to a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the attention mechanism-based information authenticity judgment method in the above embodiments.
[0063] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices over a transmission medium. The data processed by the processor is transmitted over a wireless medium via an antenna. Further, the antenna also receives data and transmits the data to the processor.
[0064] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0065] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0066] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0067] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.
Claims
1. An information authenticity judgment method based on the attention mechanism, characterized in that Including: Obtain the content of the sample text; Generate multiple first reasons to support that the content of the sample text is true information and multiple second reasons to support that the content of the sample text is false information; wherein, the number of the first reasons and the second reasons is the same; Use the content of the sample text, multiple first reasons and multiple second reasons to train an information authenticity judgment model through the following steps: Encode the content of the sample text, multiple first reasons and multiple second reasons respectively to obtain a sample text encoding, multiple first reason encodings and multiple second reason encodings; adopt an interactive attention mechanism to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively to obtain multiple aggregated encodings; adopt a reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding to evaluate the importance of each aggregated encoding; wherein, the importance of the aggregated encoding is used to represent the contribution degree of the corresponding first reason to supporting that the content of the sample text is true information, or the contribution degree of the second reason to supporting that the content of the sample text is false information; Input the content of the target text and multiple first reasons and multiple second reasons corresponding to the content of the target text into the information authenticity judgment model to detect the authenticity of the content of the target text.
2. The method for judging information authenticity based on the attention mechanism according to claim 1, wherein The information authenticity judgment model is trained through the following loss functions by using the content of the sample text, multiple first reasons and multiple second reasons: a binary classification main loss function, a contrast loss function, a reason discrimination loss function and a KL divergence loss function; Wherein, the binary classification main loss function is constructed based on whether the content of the sample text is false information, the contrast loss function is constructed based on the similarity between the content of the sample text and the first reasons and the second reasons respectively, the reason discrimination loss function is constructed based on the contribution degree of the first reasons to supporting that the content of the sample text is true information and the contribution degree of the second reasons to supporting that the content of the sample text is false information, and the KL divergence loss function is constructed based on the difference between the first reasons and the second reasons.
3. The information authenticity judgment method based on the attention mechanism according to claim 2, wherein The binary classification main loss function is: ; In the formula, represents the cross-entropy loss, is a multi-layer perceptron; y is the label of the sample text content, indicating whether the sample text content is false information; represents the weight of, , is the sample text encoding enhanced by the self-attention mechanism, is the aggregated encoding formed by aggregating the sample text encoding with the first reason encoding and the second reason encoding respectively; The contrast loss function is: ; In the formula, is an exponential function. When the label of the sample text content is real information, is the aggregated code formed by aggregating the sample text code and the first reason code. When the label of the sample text content is false information, is the aggregated code formed by aggregating the sample text code and the second reason code, The elements in are opposite to the elements in; is a random sampling operation, indicating randomly taking an instance from the set; is used to measure the cosine similarity of two vectors, is the temperature parameter, used to adjust the kurtosis of the softmax distribution; The reason discrimination loss function is: ; Wherein, CAT represents concatenating or to perform concatenation, is the aggregated code formed by aggregating the sample text code and the first reason code, is the aggregated code formed by aggregating the sample text code and the second reason code; The KL divergence loss function is: ; In the formula, KL represents calculating the KL divergence.
4. The method for judging information authenticity based on the attention mechanism according to claim 1, wherein There are multiple pieces of the content of the sample text, and the information authenticity judgment model is trained through the following steps: For each piece of the content of the sample text, use the content of the sample text and the corresponding multiple first reasons and multiple second reasons to train a base learner; Adopt a boosting strategy to integrate the base learners of multiple pieces of the content of the sample text to obtain an information authenticity judgment model.
5. The method for judging information authenticity based on the attention mechanism according to claim 4, wherein The step of adopting a boosting strategy to integrate the base learners of multiple pieces of the content of the sample text to obtain an information authenticity judgment model includes: S1. Generate initial base learners for multiple sample text contents, and initialize the weight of each sample text content to , where X is a set of multiple sample text contents; S2. Use the sample text content with weights to train the nth base learner , and calculate the error. The error calculation formula for the nth base learner is as follows: ; Wherein, is the content of the i-th sample text, N is the number of sample text contents, is an indicator function. When the input is True, the output is 1; when the input is False, the output is 0. is the label of the i-th sample text content, indicating whether the sample text content is false information or true information; If the error is greater than 0.5, terminate the boosting strategy, otherwise enter S3; S3. Update the weight of each piece of the content of the sample text according to the error, and the update formula is as follows: ; In the formula, , is an exponential function; Normalize the updated weights: ; Wherein, is the updated weight; S4. Repeat S2 and S3 for N times to obtain N trained base learners; The information authenticity judgment model is: ; In the formula, represents the ensemble of n base learners.
6. The method for judging information authenticity based on the attention mechanism according to claim 1, wherein Generating multiple first reasons supporting that the sample text content is true information and multiple second reasons supporting that the sample text content is false information, including: The system prompt for building the large language model is: Please understand the following text content and give multiple reasons for why it is true information and multiple reasons for why it is false information; in addition, since the text content is scraped from the Internet, there may be some noise or defects, please ignore them; Input the sample text content and the system prompt into the large language model to obtain multiple first reasons supporting that the sample text content is true information and multiple second reasons supporting that the sample text content is false information.
7. The method for judging information authenticity based on the attention mechanism according to any one of claims 1 to 6, characterized in that There are at least three of the first reasons and the second reasons respectively.
8. An information authenticity judgment device based on an attention mechanism, characterized in that Including: A sample acquisition module for acquiring sample text content; A sample processing module for generating multiple first reasons supporting that the sample text content is true information and multiple second reasons supporting that the sample text content is false information; where the number of the first reasons and the second reasons is the same; A model training module for training an information authenticity judgment model by using the sample text content, multiple first reasons, and multiple second reasons through the following steps: Encoding the sample text content, multiple first reasons, and multiple second reasons respectively to obtain a sample text encoding, multiple first reason encodings, and multiple second reason encodings; using an interactive attention mechanism to aggregate the sample text encoding with each first reason encoding and each second reason encoding respectively to obtain multiple aggregated encodings; using a reverse attention mechanism from the sample text encoding to each first reason encoding and each second reason encoding to evaluate the importance of each aggregated encoding; where the importance of the aggregated encoding is used to represent the contribution degree of the corresponding first reason to supporting that the sample text content is true information, or the contribution degree of the second reason to supporting that the sample text content is false information; A text detection module for inputting the target text content and multiple first reasons and multiple second reasons corresponding to the target text content into the information authenticity judgment model to detect the authenticity of the target text content.
9. A computer device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the attention mechanism-based information authenticity judgment method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the attention mechanism-based information authenticity judgment method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for determining text tags
CN111324738A
Text abstract method fusing double attention and generative adversarial network
CN115526149A
Cross-domain face anti-counterfeiting detection method and device based on multi-modal text enhancement
CN119441939A
Semi-supervised text classification method and system based on multi-encoder generative adversarial learning
CN119475126A
Text processing method and apparatus, and electronic device, computer-readable storage medium and computer program product
WO2025066553A1