False News Recognition Method Based on Heterogeneous Graph Convolutional Network
By building a heterogeneous graph convolution network, combining topological smoothing and hierarchical graph attention mechanisms, the topological imbalance problem in fake news detection is solved, and more efficient false news recognition and early detection is achieved.
Patent Information
- Application Number
- CN202210911726.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-07-28
AI Technical Summary
The existing false news detection methods ignore the authenticity and topological imbalance of edges in the news dissemination map, resulting in limited learning effects of news features.
Build a heterogeneous graph convolution network, obtain text features through natural language processing, design topological smoothing strategies and hierarchical graph attention mechanisms, and integrate text and structural features for false news recognition.
Effectively alleviate the problem of topological imbalance, improve the accuracy and early detection capabilities of false news recognition, and is better than existing methods.
Smart Images

Figure CN115438274B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology in the field of application of graph neural networks, and specifically to a method for identifying fake news based on a heterogeneous graph convolutional network. Background Art
[0002] Fake news refers to messages deliberately posted on social media and can be verified as false. The wide application of social media makes the spread of fake news faster and wider, and the spread of fake news will not only affect network security and social economy, but also damage the credibility of the government and the media. Therefore, identifying fake news as early as possible has become a crucial task. Current fake news detection methods can be divided into two categories: text-content-based methods and social-network-interaction-information-based methods.
[0003] Text-content-based methods focus on extracting lexical features, grammatical features, and writing style features from news texts, and making judgments on fake news through feature classification methods. However, this method usually analyzes news texts independently, ignoring the deep structural relationships between news and news, and between news and users during news dissemination.
[0004] To make up for the above problems, social-network-interaction-information-based methods, on the basis of texts, integrate the relationships between users and news, news and news, and users and comments in social networks, and improve the performance of fake news identification through these deeper relationships. Bian and Ma et al. formalize the relationship between the source news and comments as a tree-shaped propagation graph, and then further classify it through graph representation methods. Yuan and Yang et al. model users, source news, and comments together as a news dissemination heterogeneous graph, and then learn node features through a graph representation learning model and classify them. Although such methods have achieved excellent results in fake news detection, the authenticity of the edges in the news dissemination graph and the topological imbalance existing in the graph itself are ignored during the graph learning process, which limits the news feature learning effect of such methods. Summary of the Invention
[0005] Technical Problems to be Solved
[0006] To avoid the deficiencies of the prior art, the present invention provides a method for identifying fake news based on a heterogeneous graph convolutional network.
[0007] Technical Solution
[0008] A method for identifying fake news based on a heterogeneous graph convolutional network, characterized by the following steps:
[0009] Step 1: Obtain news data from social platforms. The news data includes source news m, relevant comments c, and corresponding users u, and construct a heterogeneous news propagation graph HNG based on the relationships among the three.
[0010] Step 2: Use a natural language processing model to obtain text feature information from the source news content and comment content.
[0011] Step 2.1: Use a natural language processing model to obtain initial features of the text.
[0012] Step 2.2: To further obtain the context semantic features between the source news and comments, use a multi-head self-attention model to obtain the relevance between comments and the source news, so as to obtain new context-semantic features for the news and comments; and use this feature as the initial feature vector of the source news node and comment node in heterogeneous graph learning.
[0013] Step 3: Design a hierarchical graph convolutional model to learn the HNG structure and obtain the structural features of the nodes.
[0014] Step 3.1: Design a topological smoothing strategy to obtain the topological position weights for each node in the news propagation network.
[0015] Step 3.2: Design a hierarchical graph attention mechanism to train the constructed HNG and perform feature learning on each node in the network.
[0016] Step 4: Integrate the network structure features obtained in Step 3 with the text information features obtained in Step 2, and then generate new vectors for further classification operations to achieve the purpose of fake news detection.
[0017] A further technical solution of the present invention: In Step 1, the social platforms are Weibo and Twitter, and three datasets, namely weibo, Twitter15, and Twitter16, are obtained from them.
[0018] A further technical solution of the present invention: The specific construction method of the heterogeneous news propagation graph HNG in Step 1 is as follows:
[0019] ① If there is a following relationship between users, or if they both comment on or forward the same news, then connect the two users.
[0020] ② If a user comments on or posts a news, then connect the user to the comment node and connect the user to the news node.
[0021] ③ If news is published in the same time period or has common users, then connect the news to the news.
[0022] ④If one comment is a reply to another comment, then connect these two comments.
[0023] A further technical solution of the present invention: The natural language processing model used in step 2.1 is a CNN model, and the purpose is to learn a feature vector representing this sentence for each piece of news and each piece of comment information.
[0024] A further technical solution of the present invention: The input of the multi-head self-attention model used in step 2.2 is the feature vectors of each piece of news and each piece of comment obtained in step 2.1. Through the multi-head self-attention mechanism, the semantic relationship between the sentences of the news and the comments is cross-learned, and finally a context semantic feature vector is obtained for each piece of news and each piece of comment.
[0025] A further technical solution of the present invention: The calculation of the topological weight of each node in the topological smoothing strategy in step 3.1 is specifically as follows:
[0026] First, the personalized PageRank algorithm is used to measure the node influence distribution of each labeled node, and finally the probability matrix P is obtained. The calculation formula is shown in (1), where a ∈ (0, 1] is the random walk probability;
[0027] P = a(I - (1 - a)A′) -1 ⑴
[0028] Secondly, assume a labeled news node m i When it is strongly influenced by the neighbor nodes of other labels, the node m i encounters a greater influence in message passing and is close to the topological class boundary; based on this assumption, the present invention designs a topological imbalance quantization index T m to capture the topological imbalance degree of the graph, while reducing the training weights of the nodes close to the class boundary and increasing the training weights of the nodes close to the class center, to re-weight the target nodes; the weight calculation formula is as follows:
[0029]
[0030] In the formula, w min , w min are hyperparameters, T m represents the topological value, Rank(T m ) represents sorting the topological value T m in ascending order, Y represents the labeled news node; finally, the corresponding topological weight value is obtained for each node in the network, and only the weight value w m of the news node is taken for subsequent calculations.
[0031] Further technical solution of the present invention: In step 3.2, the learning of the feature vectors of each type of node in the hierarchical graph attention mechanism is specifically as follows:
[0032] First, capture the importance of other types of neighbor nodes of the target node through node-level attention; then obtain the weights of neighbor nodes of the same type as the target node through type-level attention, and the formulas are shown in (3) and (4);
[0033]
[0034]
[0035] In the formula, σ(·) represents the LeakyReLU function; τ represents the node types, which are news, comment, and user, respectively.
[0036] Further technical solution of the present invention: In step 4, the feature fusion and classification module is specifically as follows:
[0037] First, for any news node m i , obtain its text feature through step 2.2 Obtain its structural feature through step 3.2 To process the features more effectively, the present invention will Fuse them to obtain the final feature, and then use cross-entropy to train the node weights of the last layer for fake news classification. The calculation formula is as follows:
[0038]
[0039]
[0040] In the formula, W is the parameter matrix, b is the error parameter, and l represents the number of categories.
[0041] Beneficial effects
[0042] A method for identifying fake news based on a heterogeneous graph convolutional network provided by the present invention. First, design a new topological smoothing strategy to measure the topological weights of each node, and obtain the topological weights of each node by increasing the weights of nodes close to the class center and reducing the weights of nodes far from the class center. Secondly, adopt a hierarchical attention mechanism to adaptively learn the weights of each edge in the news propagation network to measure the importance of each edge and alleviate the negative impact brought by untrue edges.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] 1. The present invention designs a topological smoothing strategy to measure the topological weights of labeled nodes to alleviate the problem of topological imbalance.
[0045] 2. On this basis, the present invention proposes a hierarchical attention mechanism to learn the features of HNG, and by appropriately measuring the weights of each relationship, to identify the authenticity of the relationship, thereby effectively reducing the impact of non-authentic relationships on HNG.
[0046] 3. The experimental results on the standard dataset prove that the technical model involved in the present invention has achieved better performance than the existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings are only for the purpose of illustrating specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0048] Figure 1 It is the overall model framework diagram of the method described in the embodiments of the present invention.
[0049] Figure 2 It is a schematic diagram of the heterogeneous news propagation graph (HNG) involved in the embodiments of the present invention.
[0050] Figure 3 It is the algorithm framework diagram of the multi-head self-attention mechanism in the method described in the embodiments of the present invention.
[0051] Figure 4 It is a comparison diagram of the early news detection effects between the method described in the embodiments of the present invention and the existing methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0053] The present invention proposes a false news recognition method based on a heterogeneous graph convolutional network, which consists of four sub-modules: a text data acquisition and heterogeneous news propagation graph construction module, a text feature acquisition module, a hierarchical graph convolutional module, and a node classification task training module. The overall model framework is as Figure 1 shown and is described in detail as follows:
[0054] 1. Text data acquisition and heterogeneous news propagation graph construction
[0055] 1.1 Text data acquisition
[0056] The data used in the present invention is obtained from Weibo and Twitter social platforms. The finally obtained Weibo dataset, Twitter15 dataset, and Twitter16 dataset are publicly available data that have been proven to be used. The data contains news M = [m1, m2,..., m n , and the comments R = [r1, r2,..., r j corresponding to each news, and users U = [u1, u2,..., u r
[0057] 1.2 Construction of Heterogeneous News Propagation Graph (HNG)
[0058] Based on three types of nodes: news text, comments, and news users, the present invention models five relationships: <user - post - news>, <source news - similar time / similar - source news>, <comment - comment on - source news>, <comment opinion - agree / dispute - news>, <user - follow - user>, and constructs a heterogeneous fake news network HNG to enrich the information of fake news. The finally constructed HNG is as Figure 2 shown. For more convenient description of the method, the present invention denotes HNG as G = (V, E), A represents the adjacency matrix, A' = A + I represents the adjacency matrix with self - loops added, and D represents the degree matrix.
[0059] 2. Acquisition of Text Features
[0060] 2.1 Acquisition of Initial Text Features
[0061] For a source news m i and its comments R = [r1, r2,..., r j . First, use CNN to obtain the initial sequence features i of news m The CNN feature acquisition formula is:
[0062]
[0063] In the formula, W represents the convolution kernel parameter matrix, and σ(·) represents the non - linear activation function. Similarly, extract the features j of each reply r
[0064] 2.2 Acquisition of Text Context Semantic Features
[0065] To further refine the semantic representation between comments and source news, a multi - head self - attention mechanism is used to capture the correlation between news content and comments. Specifically, the attention mechanism is used to cross - check all sentences to capture the coherence between them. After the above - mentioned semantic consistency encoding process, the text features of each news and the features of comments The multi-head self-attention model is as follows Figure 3 shown.
[0066] 3. Hierarchical graph convolutional model
[0067] 3.1 Topological smoothing strategy
[0068] In the graph structure HNG, training samples of different categories not only have differences in quantity but also in position structure. Specifically, in the node classification task, the distribution of labeled (training) nodes on the graph is also uneven, resulting in a topological imbalance problem. To alleviate the problem of poor model training ability caused by topological imbalance, first, the personalized PageRank algorithm is used to measure the node influence distribution of each labeled node, and finally, the probability matrix P is obtained. The calculation formula is as shown in (8), where a ∈ (0, 1] is the random walk probability.
[0069] P = a(I - (1 - a)A′) -1 ⑻
[0070] Secondly, assume a labeled news node m i When it is strongly influenced by neighbor nodes with other labels, node m i encounters a greater impact in message passing and is close to the topological class boundary. Based on this assumption, the present invention designs a topological imbalance quantification index T m based on node information conflict detection to capture the topological imbalance degree of the graph. While reducing the training weights of nodes close to the class boundary and increasing the training weights of nodes close to the class center, the target nodes are re-weighted. The weight calculation formula is as follows:
[0071]
[0072] In the formula, w max , w min are hyperparameters, T m represents the topological value, Rank(T m ) represents sorting the topological value T m in ascending order, and Y represents the labeled news node. Finally, the corresponding topological weight value is obtained for each node in the network, and only the weight value w m of the news node is used for subsequent calculations.
[0073] 3.2 Hierarchical graph attention mechanism
[0074] In the heterogeneous news propagation structure HNG, given a specific node, adjacent nodes of different types may have different impacts on it, and adjacent nodes of the same type may also have different importance. Therefore, in order to capture different importance at both the node level and the type level simultaneously, a two-layer attention mechanism is adopted to identify false news. Specifically, the importance of neighbor nodes of other types of the target node is captured through node-level attention; then the weights of neighbor nodes of the same type as the target node are obtained through type-level attention, and the formulas are shown in (10) and (11). In the formulas, σ(·) represents the LeakyReLU function; τ represents node types, namely news, comment, and user, three categories in total.
[0075]
[0076]
[0077] 4. False news classification
[0078] The present invention regards false news detection as a classification problem. For any news node m i , its structural features in HNG are combined with text features. Finally, cross-entropy is used to train the node weights of the last layer for false news classification, and the calculation formula is as follows:
[0079]
[0080]
[0081] In the formula, W is the parameter matrix, b is the error parameter, and l represents the number of categories. For example, the weibo dataset has only two categories (true news, false news), while the Twitter15 and Twitter16 datasets have four categories.
[0082] 5. Experiments and results
[0083] 5.1 Classification effect
[0084] Table 1 shows the classification effect of the present invention on the Twitter15 and Twitter16 datasets. The results show that the performance of the present invention on all datasets is better than that of the state-of-the-art graph-based GLAN. Specifically, on all metrics of the Twitter15 and Twitter16 datasets, the accuracy of TRHAN is 2.5% and 1.7% higher than that of the best model respectively. This is mainly attributed to two reasons. First, TRHAN takes into account the unreliable relationships and rich structural features inherent in the news dissemination graph. Second, different from CGAT and GLAN, TRHAN pays more attention to the node topology imbalance problem on the news graph, which helps to improve the model effect.
[0085] Table 1 Detection performance of the TRHAN method on the Twitter15 and Twitter16 datasets
[0086]
[0087]
[0088] 5.2 Early detection performance
[0089] Detecting fake news in the early stage is particularly important for restricting the spread of fake news. The earlier the detection period, the less dissemination information such as comments and users can be obtained. To evaluate the performance of early fake news detection, the present invention sets a series of detection periods [0h, 2h, 4h, 6h, 8h, 12h, 24h). Figure 4 Shows the performance of early fake news detection. It can be seen from the figure that the TRHAN method reaches a high accuracy very early. Specifically, the accuracy of TRHAN on the Weibo dataset is as high as 94% within 2 hours, and the accuracies on the Twitter15 dataset and the Twitter16 dataset reach 87.2% and 84.9% respectively, which are much higher than the results of other methods.
[0090] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A method for identifying fake news based on heterogeneous graph convolutional network, characterized in that The steps are as follows: Step 1: Obtain news data from social platforms. The news data includes source news m, relevant comments c, and corresponding users u, and construct a heterogeneous news propagation graph HNG according to the relationships among the three; Step 2: Use a natural language processing model to obtain text feature information for the source news content and comment content; Step 2.1: Use a natural language processing model to obtain initial features for the text; Step 2.2: To further obtain the context semantic features between the source news and comments, obtain the relevance between comments and source news through a multi-head self-attention model, so as to obtain new context-semantic features for news and comments; and use this feature as the initial feature vector of the source news node and comment node in heterogeneous graph learning; Step 3: Design a hierarchical graph convolutional model to learn the HNG structure and obtain the structural features of the nodes; Step 3.1: Design a topological smoothing strategy to obtain the topological position weights for each node in the news propagation network; Step 3.2: Design a hierarchical graph attention mechanism to train the constructed HNG and perform feature learning on each node in the network; Step 4: Fuse the network structure features obtained in Step 3 with the text information features obtained in Step 2, and then generate new vectors for further classification operations to achieve the purpose of fake news detection.
2. The method for identifying false news based on heterogeneous graph convolutional network according to claim 1, wherein In Step 1, the social platforms are Weibo and Twitter, and three datasets, namely weibo, Twitter15, and Twitter16, are obtained from them.
3. The method for identifying fake news based on heterogeneous graph convolutional network according to claim 2, wherein The specific construction method of the heterogeneous news propagation graph HNG in Step 1 is as follows: ① If there is a follow relationship between users, or both comment on or forward the same news, then connect the two users; ② If a user comments on or publishes a news, then connect the user to the comment node and connect the user to the news node; ③ If news is published in the same time period or has common users, then connect the news to the news; ④ If one comment is a reply to another comment, then connect the two comments.
4. The method for identifying fake news based on heterogeneous graph convolutional network according to claim 3, wherein The natural language processing model used in Step 2.1 is a CNN model, and the purpose is to learn a feature vector representing this sentence for each news and each comment information.
5. The method for identifying fake news based on a heterogeneous graph convolutional network according to claim 4, characterized in that The input of the multi-head self-attention model used in Step 2.2 is the feature vectors of each news and each comment obtained in Step 2.
1. Through the multi-head self-attention mechanism, cross-learn the semantic relationships of sentences between news and comments, and finally obtain a context semantic feature vector representing each news and each comment.
6. The false news recognition method based on heterogeneous graph convolutional network according to claim 5, characterized in that The calculation of the topological weight of each node in the topological smoothing strategy in Step 3.1 is specifically as follows: First, use the personalized PageRank algorithm to measure the node influence distribution of each marked node, and finally obtain the probability matrix P. The calculation formula is as shown in (1), where a ∈ (0, 1] is the random walk probability; P = a(I - (1 - a)A′) -1 (1) Secondly, assume a labeled news node m i When it is strongly influenced by neighbor nodes with other labels, node m i encounters a large influence in message passing and is close to the topological class boundary; based on this assumption, the present invention designs a topological imbalance quantization index T m to capture the topological imbalance degree of the graph, and while reducing the training weights of nodes close to the class boundary and increasing the training weights of nodes close to the class center, re-weight the target nodes; the weight calculation formula is as follows: where w min , w max are hyperparameters, T m represents the topological value, Rank(T m ) represents ascending sorting of the topological value T m , Y represents the labeled news nodes; finally, the corresponding topological weight values are obtained for each node in the network, and only the weight value w m of the news nodes is used for subsequent calculations.
7. The false news recognition method based on heterogeneous graph convolutional network according to claim 6, characterized in that The feature vector learning of each type of node in the hierarchical graph attention mechanism in Step 3.2 is specifically as follows: First, capture the importance of other types of neighbor nodes of the target node through node-level attention; then obtain the weights of neighbor nodes of the same type as the target node through type-level attention, and the formulas are shown in (3) and (4). In the formula, σ(·) represents the LeakyReLU function; τ represents node types, which are news, comment, and user, respectively.
8. The method for identifying fake news based on heterogeneous graph convolutional network according to claim 7, wherein In step 4, the feature fusion and classification module is specifically as follows: First, for any news node m i , its text features are obtained through step 2.2 Its structural features are obtained through step 3.2 To process the features more effectively, the present invention will fuse them to obtain the final features, and then use cross-entropy to train the node weights of the last layer for fake news classification. The calculation formula is as follows: In the formula, W is the parameter matrix, b is the error parameter, and l represents the number of categories.
Citation Information
Patent Citations
False news identification method based on heterogeneous graph contrast learning
CN114020928A
Social user depression tendency detection method based on heterogeneous graph attention network
CN114628008A
Cited By
False news detection method based on news transmission process and related device
CN117194806A