Method for detecting false news based on improved CNN (Convolutional Neural Network) of self-attention mechanism

Through the CNN method improved by the self-attention mechanism, the accuracy and efficiency of Chinese short text false news detection is solved, and efficient recognition of Chinese short text and attention to key semantics is achieved, which is suitable for the false information recognition task of social media and news portals.

CN120493939APending Publication Date: 2025-08-15XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510565482.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing fake news detection methods are inaccurate and processing efficiency in short Chinese text scenarios, making it difficult to effectively identify misleading content and rumor keywords, and the training of complex models is expensive and the reasoning speed is slow.

Method used

The improved CNN method based on the self-attention mechanism is adopted to extract local semantic features through character-level embedding and multi-scale convolution kernel, and the global semantic relationship modeling is enhanced with the self-attention mechanism, and a full-connection layer is used for classification detection.

Benefits of technology

It has improved the processing efficiency and semantic recognition capabilities of fake news detection, enhanced the ability to pay attention to key semantic words, and is suitable for social platform content review and news public opinion monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493939A_ABST
    Figure CN120493939A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting false news based on an improved CNN (Convolutional Neural Network) of a self-attention mechanism, which comprises the following steps of: performing coding representation on a news text by taking characters as basic units, and obtaining a character-level embedding vector through a character embedding layer; constructing a convolutional neural network structure, and extracting local semantic features of the character-level embedded vector by using a multi-scale convolution kernel; a self-attention mechanism is introduced on the basis of convolutional feature output, and features of different positions in the text are endowed with different weights; and splicing the extracted features, performing classification detection through a full connection layer, and outputting a detection result. According to the method, the processing efficiency and semantic recognition capability of false news detection are improved through character-level modeling and a lightweight convolution structure, a self-attention mechanism is introduced to endow the model with the attention capability on key semantic words, the recognition capability on misleading contents and rumor keywords is enhanced, and the recognition efficiency is improved. The method can be widely applied to false information rapid identification tasks of scenes such as social platform content review and news portal public opinion monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of false news detection and Chinese short text information recognition methods, and specifically relates to a false news detection method based on an improved CNN with a self-attention mechanism. Background Art

[0002] Fake news refers to fabricated, altered, or exaggerated information disseminated online with the intent to mislead the public. It is widely available on platforms such as social media and news portals. With the rapid development of online social platforms, fake news is spreading at an ever-increasing speed and scope, becoming a major threat to the security of online public opinion and the stability of public perception. Fake news not only undermines the quality of information users receive but also negatively impacts social stability, public sentiment, and public opinion. There is an urgent need to develop fast and accurate detection technologies to identify fake content and curb the spread of misinformation.

[0003] Fake news detection methods can be categorized into two types: traditional machine learning-based methods and deep learning-based methods. Traditional methods, such as support vector machines, random forests, and logistic regression, primarily rely on manually designed text features, such as word frequency, TF-IDF, and sentiment word proportions, to perform classification modeling. While computationally efficient, these methods suffer from limited feature dimensionality and poor context modeling, leading to unstable performance and weak generalization when dealing with new, variant, or semantically complex fake news. Traditional methods often face challenges with sparse features and limited expressiveness when dealing with large social media corpora with complex semantic features.

[0004] The rise of deep learning has brought new developments to fake news detection. In particular, convolutional neural networks, recurrent neural networks, and their variants have demonstrated powerful capabilities for automatic feature learning and context modeling in natural language processing tasks. For example, researchers have used CNNs to extract local n-gram semantic features, which are suitable for modeling fake news in short text structures. Others have proposed integrating models such as BERT and LSTM to enhance semantic understanding. However, existing deep learning methods still have certain limitations. First, in the context of short Chinese texts, word-level modeling tends to lose fine-grained information, while character-level modeling suffers from weak contextual dependencies. Second, while complex models can improve accuracy, they are expensive to train and slow to infer, making them unsuitable for real-time detection. Third, some models fail to adequately focus on key semantic features in the text, resulting in insufficient ability to extract core signals such as misleading vocabulary and emotional language. The self-attention mechanism provides models with the ability to model global semantic dependencies, making it particularly suitable for capturing long-range dependencies and important information from text. In the fake news detection task, the introduction of the self-attention mechanism can dynamically adjust the importance of each word, improving the model's ability to focus on deceptive information and key sentences. The improved lightweight convolutional architecture can quickly extract local features and enhance the processing efficiency of short Chinese texts. Therefore, the modeling method that integrates CNN and self-attention mechanism has become a hot research direction in recent years. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for false news detection based on an improved CNN with a self-attention mechanism, which solves the problems of low accuracy and processing efficiency of existing false news detection methods.

[0006] The technical solution adopted by the present invention is: a method for detecting fake news based on an improved CNN with a self-attention mechanism, comprising the following steps:

[0007] Step 1: Obtain labeled Chinese news text data and perform preprocessing;

[0008] Step 2: Encode the preprocessed Chinese news text using characters as basic units and obtain character-level embedding vectors through the character embedding layer;

[0009] Step 3: Build an improved convolutional neural network structure and use multi-scale convolution kernels to extract local semantic features of character-level embedding vectors;

[0010] Step 4: Based on the convolutional feature output, a self-attention mechanism is introduced to assign different weights to features at different positions in the text;

[0011] Step 5: Concatenate the extracted features, perform classification detection through the fully connected layer, and output the detection results of true and false news.

[0012] The present invention is also characterized in that:

[0013] The preprocessing in step 1 is to remove noise information.

[0014] Step 2 specifically includes the following steps:

[0015] Step 2.1: Assume that the preprocessed data contains N news texts, each news text x n Each consists of M characters x n,m Composition, then the news text x n Expressed as:

[0016] x n =[x n,1 ,x n,2 ,...,x n,M ]

[0017] Step 2.2, each character x n,m Mapped into a high-dimensional sparse vector v through one-hot encoding n,m :

[0018]

[0019] Where B is the size of the character table, that is, the dimension of the one-hot encoding vector; represents B-dimensional real number space;

[0020] Step 2.3: embed the high-dimensional sparse vector v through the embedding layer n,m Mapped to a low-dimensional continuous vector V n,m :

[0021] V n,m =W V v n,m +b V

[0022] Where W V is the embedding matrix with size D×B, where D is the target dimension of the embedding vector; b V is the bias term.

[0023] Step 2 also includes:

[0024] Step 2.4, through the character-level embedding vector V n,m x n The embedding vector V n Expressed as:

[0025] V n =[V n,1 ,V n,2 ,...,V n,M ].

[0026] Step 3 specifically includes the following steps:

[0027] Step 3.1, select each character x n,m The d characters before and after each construct a context window of length 2d+1, and the context representation vector C is formed based on the embedding vector of each character in the window n,m :

[0028] C n,m =[V n,m-d ,...,V n,m ,...,V n,m+d ]

[0029] Step 3.2: Represent the context vector C n,m As the input of the convolution kernel, the convolution operation is used to extract local context features to obtain the local semantic feature vector S n,m :

[0030] S n,m =σ(W C *C n,m +b C )

[0031] Among them, σ(·) represents the activation function, W C is the convolution kernel weight matrix, * represents the convolution operation, b C is the bias term corresponding to the convolution operation;

[0032] Step 3.3: All local semantic feature vectors S n,m According to the order of characters in the text, a local semantic feature vector sequence S is formed. n :

[0033] S n =[S n,1 ,S n,2 ,...,S n,M ].

[0034] Step 4 specifically includes the following steps:

[0035] Step 4.1: local semantic feature vector sequence S n Perform linear transformation to generate query matrix Q, key matrix K and value matrix V' respectively:

[0036] Q=W Q S n ,K=W K S n ,V'=W V 'S n

[0037] Step 4.2. Calculate the attention score matrix A based on the query matrix Q and the key matrix K:

[0038]

[0039] Where T is the matrix transpose operation, d k is the scaling factor;

[0040] Step 4.3: Perform weighted summation based on the attention score matrix A and the value matrix V' to obtain the attention enhancement feature matrix H:

[0041] H=AV'

[0042] Step 4.4: Substitute the local semantic feature vector sequence S n Perform residual connection with the attention enhancement feature matrix H to obtain the enhanced feature vector sequence

[0043]

[0044] Step 4.5: Enhanced feature vector sequence Perform the maximum pooling operation to obtain the final feature vector

[0045]

[0046] Step 5 specifically includes the following steps:

[0047] Step 5.1: Pool the feature vector The concatenated input is input into the fully connected layer, and the binary classification probability is output through the Softmax activation function. The Softmax activation function μ(·) is expressed as:

[0048]

[0049] Among them, ζ and ζ' are the two category scores output by the neural network;

[0050] Step 5.2 uses the binary cross entropy loss function combined with the Adam optimizer to train the model. The binary cross entropy loss function is expressed as:

[0051]

[0052] in, is the true label, O n The probability of fake news detected by the model;

[0053] Step 5.3: The probability of fake news detected by the model O according to the decision threshold τ n Classify, according to the final classification results The classification rules for determining whether news is true or false are:

[0054]

[0055] The decision threshold τ in step 5.3 is set to 0.5.

[0056] The beneficial effects of the present invention are as follows: the method for detecting fake news based on the improved CNN of the present invention based on the self-attention mechanism is aimed at the Chinese short text scenario in the social media environment. Through character-level modeling and lightweight convolutional structure, the processing efficiency and semantic recognition ability of fake news detection are improved. The introduction of the self-attention mechanism gives the model the ability to pay attention to key semantic words, and enhances the ability to recognize misleading content and rumor keywords. It can be widely used in the task of quickly identifying false information in scenarios such as content review of social platforms and public opinion monitoring of news portals. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 1 is a flow chart of a method for detecting fake news using an improved CNN based on a self-attention mechanism;

[0058] Figure 2 Schematic diagram of the structure of the method for detecting fake news using an improved CNN based on the self-attention mechanism of the present invention;

[0059] Figure 3 is an example illustration of a single word embedding operation;

[0060] Figure 4 It is a schematic diagram of the convolution sliding window operation;

[0061] Figure 5 It is the structural diagram of the CNN model;

[0062] Figure 6 It is a calculation flow chart of the self-attention mechanism module of the present invention;

[0063] Figure 7 This is the running speed graph of the n=6 Rumor dataset;

[0064] Figure 8 This is the running speed graph of the n=7 Rumor dataset;

[0065] Figure 9 This is the running speed graph of the n=6 CHEF dataset;

[0066] Figure 10 This is the running speed graph of the n=7 CHEF dataset;

[0067] Figure 11 It is the accuracy graph on the Rumor dataset;

[0068] Figure 12 It is the accuracy graph on the CHEF dataset. DETAILED DESCRIPTION

[0069] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0070] Example 1

[0071] The present invention provides a method for detecting fake news using an improved CNN based on the self-attention mechanism. Figure 1 As shown, first, a fake news dataset with true and false labels is obtained from a Chinese social platform, and the Chinese text is preprocessed, including removing stop words and special characters; the processed text is embedded at the character level to generate a fixed-length vector representation; a convolutional neural network is used to extract local semantic features, and a self-attention mechanism is introduced to enhance the model's ability to model global semantic relationships; the extracted features are spliced and the true and false news classification is completed through a fully connected layer; the model is trained using a cross-entropy loss function, and the model parameters are updated through multiple rounds of iterations; in the testing phase, the new input text is classified and tested, and the fake news identification result is output. The method of the present invention gives the model the ability to pay attention to key semantic words through the self-attention mechanism, combined with the efficient feature extraction capability of the convolutional structure, effectively improving the detection accuracy and model generalization ability in the Chinese short text scenario, and is suitable for false information identification tasks such as social media content review and fact checking.

[0072] The overall network architecture of the model of the present invention is shown in FIG. Figure 2 As shown in the figure, it includes an embedding layer, a convolutional layer, a self-attention mechanism module, and a classification output layer. The specific implementation steps are as follows:

[0073] Step 1: Obtain labeled Chinese news text data from social platforms or fake news-related datasets, and clean the obtained Chinese news text, including removing noise information such as special symbols, HTML tags, stop words, web links, and non-language characters to ensure the standardization and consistency of text input.

[0074] Step 2: Encode the pre-processed Chinese news text using characters as the basic unit and obtain character-level embedding vectors through the character embedding layer. Using character-level embedding, the input text is converted into a continuous vector representation that can be processed by the neural network. An example of a single word embedding operation is shown below: Figure 3 shown.

[0075] Step 3: Figure 4 and Figure 5 As shown in the figure, an improved convolutional neural network structure is constructed, and multi-scale convolution kernels are used to extract local n-gram features of news text to form basic semantic representation.

[0076] Step 4: Figure 6As shown in the figure, a self-attention mechanism is introduced based on the convolutional feature output to assign different weights to features at different positions in the text, thereby enhancing the model's ability to recognize key semantic words or fragments.

[0077] Step 5: The fused features are classified and detected through the fully connected layer, and the detection results of true and false news are output.

[0078] Step 6: Use real social platform datasets for model training and verification. After model training is completed, the weights can be saved for migration to other data or platforms to achieve rapid fake news detection for new input text.

[0079] Example 2

[0080] The present invention provides a method for detecting fake news using an improved CNN based on a self-attention mechanism. Based on Example 1, step 2 is preferably:

[0081] Step 2.1: Assume that the preprocessed data contains N news texts, each news text x n Each consists of M characters x n,m Composition, then the news text x n Expressed as:

[0082] x n =[x n,1 ,x n,2 ,...,x n,M ]

[0083] Step 2.2: The original characters are discrete symbols and cannot be directly used for neural network calculations. They need to be mapped into high-dimensional vectors. n,m Mapped into a high-dimensional sparse vector v through one-hot encoding n,m :

[0084]

[0085] Where B is the size of the character table, that is, the dimension of the one-hot encoding vector; represents the B-dimensional real space, v n,m A vector has one position that is 1 and all other positions are 0.

[0086] Step 2.3: The dimension of one-hot encoding is high, the computational overhead is high, and it cannot express the semantic relationship between characters. The one-hot vector is further mapped to a low-dimensional continuous representation through the embedding layer. The high-dimensional sparse one-hot encoding vector v generated in the previous step is embedded in the mapping process. n,m Mapped into a low-dimensional continuous dense vector V through embedding transformation n,m , the specific formula is as follows:

[0087] V n,m =WV v n,m +b V

[0088] In the formula, an embedding matrix W is used V For sparse vector v n,m Perform a linear transformation and add the bias term b V , where the embedding matrix W V The size of is D×B, where B is the size of the character table and D is the target dimension of the embedding vector. The bias term is used to further optimize the mapping result and enhance the expressive power.

[0089] Step 2.4: Construct text embedding representation based on character-level embedding vectors. In this step, the embedding vector V of each character obtained in step 2.3 is used. n,m , the same news x n The embedding vectors of all characters in the text are sequentially concatenated to obtain the text-level embedding vector V n , specifically expressed as:

[0090] V n =[V n,1 ,V n,2 ,...,V n,M ]

[0091] Through this operation, the entire news text can be encoded into a sequence of character embedding vectors, which can be used as the input for subsequent feature extraction and modeling. The text-level embedding vector V n It is used to organize character embedding vectors sequentially, making it easier for convolution operations to be processed in text order. Each character is mapped to a low-dimensional vector space, enabling the neural network to learn the semantic relationship between characters and improve the model's ability to capture fake news text patterns.

[0092] Example 3

[0093] The present invention provides a method for detecting fake news using an improved CNN based on a self-attention mechanism. Based on Example 1, step 3 is preferably:

[0094] Step 3.1: Use CNN for feature modeling to extract local n-gram-level patterns and combine it with the self-attention mechanism to further enhance the global feature expression capability. n,m , with it as the center, and d characters before and after, a context window of length 2d+1 is constructed, and the context representation vector C is formed based on the embedding vector of each character in the window n,m , specifically expressed as:

[0095] C n,m =[V n,m-d ,...,Vn,m ,...,V n,m+d ]

[0096] Where C n,m Indicates the character x n,m V is the center, containing the embedding vector sequence of its neighboring context characters; n,m Represents character x n,m The embedding vector of ; d represents the window radius, i.e. d characters are taken forward and backward respectively.

[0097] Step 3.2: In the convolution calculation phase, a sliding convolution window mechanism is used to extract local context features. The entire text sequence is scanned by the convolution kernel. Each window is regarded as a convolution receptive field, and features are extracted through the convolution operation. n,m As the input of the convolution kernel, the local context features are extracted by convolution operation. The result of the convolution operation is recorded as the local feature vector S n,m , the calculation formula is as follows:

[0098] S n,m =σ(W C *C n,m +b C )

[0099] Where W C is the convolution kernel weight matrix; b C is the bias term corresponding to the convolution operation; * represents the convolution operation; σ(·) represents the activation function, which is used to increase the nonlinear expression ability of the feature.

[0100] Step 3.3, by setting all S n,m According to the order of characters in the text, the local feature sequence S is formed. n , used to characterize local patterns in text:

[0101] S n =[S n,1 ,S n,2 ,...,S n,M ]

[0102] Where M represents the number of characters in the text xn.

[0103] Example 4

[0104] The present invention provides a method for detecting fake news using an improved CNN based on a self-attention mechanism. Based on Example 1, step 4 is preferably:

[0105] Step 4.1: The local features captured by CNN alone cannot model long-distance dependencies, so the self-attention mechanism is introduced to further enhance the global feature representation capability.n Perform linear transformation to generate query matrix Q, key matrix K and value matrix V' respectively. The formula is:

[0106] Q=W Q S n ,K=W K S n ,V'=W V' S n

[0107] Step 4.2, attention weight calculation, calculate the attention score matrix A based on the query matrix and the key matrix, the formula is:

[0108]

[0109] Where T represents the matrix transpose operation; d k is a scaling factor to prevent the gradient from vanishing or exploding.

[0110] Step 4.3: Use the attention score matrix A and the value matrix V' for weighted summation to obtain the attention enhancement feature matrix H, which is:

[0111] H=AV'

[0112] Among them, H is the enhanced feature representation, which can integrate the global information in the text, so that the model can focus on cross-sentence dependencies when detecting fake news, rather than relying solely on local patterns.

[0113] Step 4.4: Substitute the local semantic feature vector sequence S n Perform residual connection with the attention enhancement feature matrix H to obtain the enhanced feature vector sequence The formula is:

[0114]

[0115] Step 4.5: Enhanced feature vector sequence Perform the maximum pooling operation to extract the most important key information and obtain the final feature vector The formula is:

[0116]

[0117] Where, Represents the overall feature vector of the text after pooling, which is used for subsequent classification tasks.

[0118] Example 5

[0119] The present invention provides a method for detecting fake news using an improved CNN based on a self-attention mechanism. Based on Example 1, step 5 is preferably:

[0120] Step 5.1: After completing character embedding and feature modeling, the extracted text features need to be classified to determine the authenticity of the news. The goal of the classification detection stage is to calculate the probability of the news text belonging to "fake news" or "true news" based on the feature representation learned by the model, and optimize it using an appropriate loss function. The concatenated input is input into the fully connected layer, and the binary classification probability is output through the Softmax activation function. The Softmax activation function μ(·) is defined as:

[0121]

[0122] Where ζ and ζ' are the two category scores output by the neural network.

[0123] Step 5.2: In order to optimize the model and make the detection results closer to the true label, the binary cross entropy loss function is used. The cross entropy loss is suitable for binary classification problems and can measure the difference between the predicted distribution and the true distribution. It is defined as follows:

[0124]

[0125] in, is the true label, O n is the probability of fake news detected by the model. Adam is used to update the gradient of the loss function to accelerate convergence and avoid local optimal solutions.

[0126] Step 5.3: In the test phase, given a news text xn, the model calculates its final probability O n , and the probability of fake news detected by the model is O according to the set threshold n To classify:

[0127]

[0128] in, is the final classification result; τ is the decision threshold, which is usually set to 0.5.

[0129] Example 6

[0130] The trained detection model is used to test the test data. After obtaining the output results, the results are evaluated using accuracy, precision, recall, and F1 score as evaluation criteria. The pseudo code of the improved convolutional neural network algorithm based on the self-attention mechanism is shown in Table 1.

[0131] Table 1 Pseudocode of the improved convolutional neural network algorithm based on the self-attention mechanism

[0132]

[0133] To validate the effectiveness of our model in fake news detection, we conducted experiments using two representative Chinese social media datasets: the Rumor dataset and the CHEF dataset. These datasets, sourced from a wide range of Chinese social media platforms, including Sina Weibo and Tencent's Jiuzhen platform, as well as rumor-debunking organizations, provide fake news annotations in real social contexts, enabling a comprehensive evaluation of the model's effectiveness in real-world Chinese short text detection tasks.

[0134] Among them, the Rumor dataset is a two-category dataset, containing 3,387 social comment texts with true and false labels, of which 1,538 are marked as fake news and 1,849 are marked as true news. This dataset has the characteristics of short text length, emotional language, and obscure expression, and is suitable for building a short text recognition model for rumor detection. The CHEF dataset focuses on Chinese fact-checking tasks, comes from multiple mainstream Chinese rumor-busting platforms, and contains a total of more than 10,000 labeled news corpora. The present invention selects 3,543 supporting statement samples and 5,065 opposing statement samples as experimental data to achieve in-depth modeling of complex factual texts.

[0135] Considering the significant differences in text length distribution across datasets, to ensure consistency in the model's input feature dimensions, this paper standardized all sample text lengths to 150 characters: texts shorter than 150 characters were padded with special placeholder characters, and texts longer than 150 characters were truncated. This preprocessing strategy improves the model's adaptability to diverse data distributions while avoiding modeling errors caused by inconsistent input dimensions.

[0136] In the experiment, the improved CNN-SAN model proposed in this invention was applied to the above two datasets respectively. Combining the global modeling capabilities of convolutional extraction and self-attention mechanism, a horizontal comparative analysis was conducted with multiple mainstream models, systematically verifying the accuracy and robustness of the invention in the scenario of false news detection on Chinese social media.

[0137] Figure 7 and Figure 8 The changes in the running speed curve of the present invention under different sample number settings (n=6 and n=7) of the Rumor dataset are demonstrated, indicating that the model has high computational efficiency in short text data processing tasks and runs stably in multiple groups of experiments. Figure 9 and Figure 10 This is a graph of the running speed of the present invention under different settings on the CHEF dataset, which further verifies the efficiency of the model in processing longer texts and fact-checking tasks. Figure 11 and Figure 12The model's accuracy comparison results on the Rumor and CHEF datasets are presented. As can be seen, the proposed method outperforms other mainstream models in classification accuracy on both datasets, demonstrating strong classification and discrimination capabilities, demonstrating its application value in Chinese fake news detection. To verify the effectiveness of the proposed improved CNN method based on the self-attention mechanism for fake news detection, five comparison algorithms were used to compare with the proposed classification algorithm. The specific experimental results are shown in Tables 2 and 3.

[0138] Table 2 Algorithm comparison results

[0139]

[0140] Table 3 Algorithm comparison results

[0141]

[0142] On the Rumor dataset, as shown in the table, the CNN-SAN model achieved the best accuracy in both tasks (n=6 and n=7), achieving 85.24% and 86.85% respectively, with F1 scores of 85.19% and 86.84% respectively, demonstrating stable and outstanding performance. The BERT model achieved an accuracy of 85.16% when n=6, close to that of CNN-SAN, but dropped to 78.59% when n=7, showing significant fluctuations, indicating that its performance is sensitive to task granularity. The MS-LSTM and LSTM+GRU performed relatively steadily in both tasks, maintaining accuracy around 82%. While the ML-Transformer and Transformer+GRU showed some advantages when n=6, their accuracy dropped slightly when n=7, indicating that the fusion model still faces generalization bottlenecks when tasks are expanded.

[0143] On the CHEF dataset, CNN-SAN also achieved the best accuracy, reaching 85.16% when n=6 and further improving to 86.38% when n=7. Similar to the Rumor dataset, the BERT model's accuracy decreased slightly on this dataset, but remained around 82% overall, demonstrating its strong semantic modeling capabilities for standard text. The MS-LSTM and LSTM+GRU models performed reliably, achieving an accuracy of around 84% in both tasks and achieving relatively good F1 scores. However, the ML-Transformer and Transformer+GRU models exhibited relatively low accuracy and significant fluctuations on this dataset, indicating that their feature extraction capabilities for this type of Chinese data still require improvement.

[0144] Through this approach, the present invention demonstrates high detection accuracy and model robustness in Chinese social media and fact-checking scenarios. By introducing a self-attention mechanism, the model dynamically focuses on key semantic information in the text, enhancing its ability to identify misleading content and rumor keywords. This structural design effectively improves the model's ability to model and generalize complex semantic relationships, and is widely applicable in scenarios such as content review on social platforms and public opinion monitoring on news portals.

Claims

1. A method for detecting fake news using an improved CNN based on the self-attention mechanism, characterized in that: The following steps are involved: Step 1: Obtain labeled Chinese news text data and perform preprocessing; Step 2: Encode the preprocessed Chinese news text using characters as basic units and obtain character-level embedding vectors through the character embedding layer; Step 3: Build an improved convolutional neural network structure and use multi-scale convolution kernels to extract local semantic features of character-level embedding vectors; Step 4: Based on the convolutional feature output, a self-attention mechanism is introduced to assign different weights to features at different positions in the text; Step 5: Concatenate the extracted features, perform classification detection through the fully connected layer, and output the detection results of true and false news.

2. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 1, wherein: The preprocessing in step 1 is to remove noise information.

3. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 1, wherein: The step 2 specifically includes the following steps: Step 2.1: Assume that the preprocessed data contains N news texts, each news text x n Each consists of M characters x n,m Composition, then the news text x n Expressed as: x n =[x n,1 ,x n,2 ,...,x n,M ] Step 2.2, each character x n,m Mapped into a high-dimensional sparse vector v through one-hot encoding n,m : Where B is the size of the character table, that is, the dimension of the one-hot encoding vector; represents B-dimensional real number space; Step 2.3: embed the high-dimensional sparse vector v through the embedding layer n,m Mapped to a low-dimensional continuous vector V n,m : V n,m =W V v n,m +b V Where W V is the embedding matrix with size D×B, where D is the target dimension of the embedding vector; b V is the bias term.

4. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 3, wherein: The step 2 further comprises: Step 2.4, through the character-level embedding vector V n,m Embedding vector V of news text xn n Expressed as: V n =[V n,1 ,V n,2 ,...,V n,M ]。 5. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 3, wherein: The step 3 specifically includes the following steps: Step 3.1, select each character x n,m The d characters before and after each construct a context window of length 2d+1, and the context representation vector C is formed based on the embedding vector of each character in the window n,m : C n,m =[V n,m-d ,...,V n,m ,...,V n,m+d ] Step 3.2: Represent the context vector C n,m As the input of the convolution kernel, the local context features are extracted using the convolution operation to obtain the local semantic feature vector S n,m : S n,m =σ(W C *C n,m +b C ) Among them, σ(·) represents the activation function, W C is the convolution kernel weight matrix, * represents the convolution operation, b C is the bias term corresponding to the convolution operation; Step 3.3: All local semantic feature vectors S n,m According to the order of characters in the text, a local semantic feature vector sequence S is formed. n : S n =[S n,1 ,S n,2 ,...,S n,M ]。 6. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 5, wherein: The step 4 specifically includes the following steps: Step 4.1: local semantic feature vector sequence S n Perform linear transformation to generate query matrix Q, key matrix K and value matrix V' respectively: Q=W Q S n ,K=W K S n ,V'=W V 'S n Step 4.

2. Calculate the attention score matrix A based on the query matrix Q and the key matrix K: Where T is the matrix transpose operation, d k is the scaling factor; Step 4.3: Perform weighted summation based on the attention score matrix A and the value matrix V' to obtain the attention enhancement feature matrix H: H=AV' Step 4.4: Substitute the local semantic feature vector sequence S n Perform residual connection with the attention enhancement feature matrix H to obtain the enhanced feature vector sequence Step 4.5: Enhanced feature vector sequence Perform the maximum pooling operation to obtain the final feature vector 7. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 6, wherein: The step 5 specifically includes the following steps: Step 5.1: Pool the feature vector The concatenated input is input into the fully connected layer, and the binary classification probability is output through the Softmax activation function. The Softmax activation function μ(·) is expressed as: Among them, ζ and ζ' are the two category scores output by the neural network; Step 5.2 uses the binary cross entropy loss function combined with the Adam optimizer to train the model. The binary cross entropy loss function is expressed as: in, is the true label, O n The probability of fake news detected by the model; Step 5.3: The probability of fake news detected by the model O according to the decision threshold τ n Classify, according to the final classification results The classification rules for determining whether news is true or false are:

8. The method for detecting fake news using an improved CNN based on a self-attention mechanism as claimed in claim 7, wherein: The decision threshold τ in step 5.3 is set to 0.5.