Unreal information detection method based on user-post interaction network

By constructing a user-post interactive network model with false information data sets containing forwarded text and users and multiple entities and multiple relationships, the problems of incomplete social background information and insufficient fusion of heterogeneous node characteristics in the existing methods are solved, and the accuracy of false information detection is improved.

CN120011508APending Publication Date: 2025-05-16SHIHEZI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090003.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing false information detection method based on social backgrounds has the problem of insufficient social background information and insufficient fusion of heterogeneous node characteristics, especially when processing Chinese data sets, there is a lack of data sets containing specified forward user information.

Method used

A false information detection method based on user-post interactive network is proposed. By constructing a false information data set containing forwarded text and users, a user-post interactive network model with multiple entities and multiple relationships is established, and information exchange and deep fusion is used for blocked heterogeneous graph attention network and gate mechanism to improve the accuracy of false information detection.

Benefits of technology

By improving social background information and deeply fusion of heterogeneous node characteristics, the accuracy of false information detection is significantly improved, and the problems of incomplete content and insufficient fusion of feature in the existing methods are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011508A_ABST
    Figure CN120011508A_ABST
Patent Text Reader

Abstract

The invention discloses an unreal information detection method based on a user-post interactive network. The method comprises the following steps: firstly, carrying out initialization coding on different types of node features; then modeling a user, a text node and a relationship between the user and the text node by using a block heterogeneous graph attention network, and carrying out adjacent node feature aggregation operation mainly in three parts at the same time: U-GAT is used for modeling a user forwarding network and aggregating features between the user nodes; the P-GAT is used for modeling a post forwarding network and aggregating features among text nodes; uP-GAT serves as a bridge between a user and a text node to aggregate heterogeneous type nodes, and plays a role in information exchange and feature fusion. And finally, fusing the user features, the text features and the root node features through a gating mechanism, and using the fused features for unreal information detection. The unreal information detection method for the multi-entity and multi-relation user-post interaction network can be applied to social media, provides reliable and accurate information for the user, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of false information detection, and in particular to but not limited to a false information detection method based on a user-post interaction network. Background Art

[0002] In recent years, the proliferation of false information on the Internet in social media has seriously affected people's production and life. Compared with traditional false information, the openness and inclusiveness of social media have led to a faster spread of false information on the Internet, a wider range of influence, and a deeper degree of harm. Therefore, the detection of false information on the Internet on social media has attracted the attention of many scholars. However, the publishers of false information often fabricate false information by imitating the writing characteristics of real information or modifying some fragments of real information to deliberately deceive readers, making it difficult to detect the truth of the message based solely on the content of false information. The method based on social context detects the truth of the message by capturing the differences in the propagation structure between false information and non-false information. This method reduces the dependence of the detection results on the content of false information and has stronger generalization performance, thus attracting much attention from scholars.

[0003] At present, the false information detection methods based on social context still have many limitations:

[0004] First, social context information is not perfect. Complex social context information contains multiple entities and relationships between entities, which brings great challenges to modeling and utilizing social context information. Previous studies only considered ordinary forwarding users and ordinary forwarding relationships, ignoring the designated forwarding users and designated forwarding relationships (such as @ a certain user) hidden in the forwarding text, which limits the further improvement of false information detection performance.

[0005] Secondly, the integration of social context information is not sufficient. The main body of social context information is the interaction between users and posts. Since users and posts (texts) are different types of entities, most previous studies only use one of users or posts in combination with the propagation structure to detect false information. The lack of comprehensive utilization and deep integration of user and post content limits the further extraction of social context information by false information detection models.

[0006] In addition, there is a lack of Chinese false information datasets that contain social context information. Research on false information detection methods based on social context requires datasets with rich social context information. However, the existing Chinese datasets containing social context information are outdated and cannot meet the rapidly changing message propagation patterns in social media. In addition, the current datasets containing social context information do not contain information about designated forwarding users, which limits scholars' research on false information detection methods based on social context on Chinese datasets. Summary of the invention

[0007] In view of this, an embodiment of the present invention provides a method for detecting false information based on a user-post interaction network, which at least solves the problems of incomplete dissemination content and insufficient fusion of heterogeneous node features in existing research.

[0008] The technical solutions of the embodiments of the present invention are as follows:

[0009] An embodiment of the present invention provides a method for detecting false information based on a user-post interaction network, the method comprising:

[0010] Construct a false information dataset containing forwarded texts and users, wherein each sample in the false information dataset includes a post set, a post forwarding graph, a user set, a user forwarding graph, and a bipartite graph representing the subordinate relationship between the post set and the user set; wherein the user set includes original users, ordinary forwarding users, and designated forwarding users; the post forwarding graph is constructed based on the forwarding relationship between all posts; the user forwarding graph is constructed based on the forwarding relationship between all users; construct a user-post interaction network model and train it using the false information dataset; wherein the user-post interaction network model includes: a feature extraction module for initializing and encoding text information, users, and forwarding relationships; a block heterogeneous graph attention network module for simultaneously modeling users and text information and the relationship between the two; a gating mechanism fusion module for root node enhancement; and a false information classifier for predicting whether a sample is false information; and predict the false information data to be tested using the trained user-post interaction network model.

[0011] In some embodiments of the present invention, the block heterogeneous graph attention network includes three parallel parts: a user forwarding network U-GAT, a post forwarding network P-GAT, and a fusion exchange network UP-GAT: U-GAT is used to model the user forwarding graph and aggregate features between user nodes; P-GAT is used to model the post forwarding graph and aggregate features between text nodes; UP-GAT acts as a bridge between user nodes and text nodes to aggregate heterogeneous types of nodes, playing a role in information exchange and feature fusion.

[0012] In some embodiments of the present invention, after the U-GAT and P-GAT aggregation operations of the current layer are completed, the user node features and the text node features are updated, and the information exchange between the user nodes and the text nodes is realized through the UP-GAT of the next layer; after multiple iterations, the updated user features and text node features are respectively input into the fully connected layer, the average pooling layer and the maximum pooling layer for node feature convergence operations to obtain the corresponding user feature vector representation and text feature vector representation.

[0013] In some embodiments of the present invention, in U-GAT, if there are multiple forwardings between two users, the weights of the edges representing the user forwarding relationship will be accumulated according to the following formula:

[0014]

[0015] in, is the weight of the edge corresponding to the forwarding relationship between user nodes i and j, r1 and r2 are the number of common forwarding and specified forwarding between user nodes i and j, respectively, β∈(0,2] represents the weight of the specified forwarding edge, and the weight of the common forwarding edge is set to 1 by default.

[0016] In some embodiments of the present invention, in the feature extraction module, for three types of user features, namely numerical, Boolean and discrete, corresponding feature transformation methods are respectively used to encode the feature representation of the user node from multiple angles; all user features of each sample are normalized using heterogeneous layer normalization; all samples in the same batch are stacked to obtain the final user features; and the final user features are input into the fully connected and activated layers to obtain the user node encoding representation of a specific dimension.

[0017] In some embodiments of the present invention, in the feature extraction module, for the text content, a pre-trained BERT is used as a fixed text feature extractor to extract all the original texts and forwarded texts after segmentation, and the output vector of the classification mark is used as the feature representation of each text node.

[0018] In some embodiments of the present invention, the three types of user features, namely numerical, Boolean and discrete, are respectively transformed using corresponding feature transformation methods, and the feature representation of the user node is obtained by encoding from multiple angles, including: for numerical user features, a variety of normalization methods are used for processing: the user's relative registration days and forwarding delay time are normalized to zero mean; the number of friends, the number of speeches and the number of fans are normalized to zero mean after taking the natural logarithm; the number of characters in the authentication reason and the number of characters in the personal profile are normalized using the maximum and minimum values; for Boolean type features, binarization is performed; for discrete type features, the corresponding feature representation is obtained by conversion through one-hot encoding or embedding; wherein the user's relative registration days are determined based on the difference between the user's forwarding time and the user's registration time; the forwarding delay time is determined based on the difference between the user's forwarding time and the post publishing time.

[0019] In some embodiments of the present invention, the gating mechanism fusion module performs root node feature enhancement through the following process: for a given k-layer graph attention network output and Represent the feature matrices of users and text nodes respectively; and Input the fully connected layer for dimension conversion and Use maximum pooling and average pooling to transform the feature matrix of all user nodes respectively Aggregate and concatenate to get the user feature vector c u , the transformation feature matrix of all text nodes Aggregate and concatenate to get the text feature vector c p ; Concatenate the original user features and original text features of the root node to obtain the root node joint feature representation h root ; Through the gating mechanism, c u 、c p and h root Fusion is performed and the joint feature representation h is finally output out .

[0020] In some embodiments of the present invention, the false information classifier inputs the joint feature representation into a feedforward neural network, obtains the probability that the sample belongs to a certain category through a normalization function, calculates the classifier learning loss using a cross entropy loss function, and trains the user-post interaction network model through back propagation.

[0021] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0022] In the embodiment of the present invention, a false information dataset containing social context information is first collected, including designated forwarding relationships and forwarding user information, to further improve the social context information. Then, a multi-entity and multi-relation user-post interaction network is proposed to achieve information exchange and deep fusion between heterogeneous nodes. In addition, the original user and original text are fused again through the gating mechanism to enhance the role of the original content in false information detection and further improve the accuracy of false information detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work, among which:

[0024] Figure 1 A flow chart of a method for detecting false information based on a user-post interaction network provided by an embodiment of the present invention;

[0025] Figure 2 A schematic diagram of a false information propagation network provided by an embodiment of the present invention;

[0026] Figure 3 A schematic diagram of the architecture of a user-post interaction network model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0028] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0029] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.

[0030] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.

[0031] Figure 1 A flow chart of a method for detecting false information based on a user-post interaction network provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the method comprises at least the following steps:

[0032] Step S110, constructing a false information dataset including forwarded texts and users, wherein each sample in the false information dataset includes a post set, a post forwarding graph, a user set, a user forwarding graph, and a bipartite graph representing the subordinate relationship between the post set and the user set.

[0033] Here, the user set includes original users, common forwarding users and designated forwarding users; the post forwarding graph is constructed based on the forwarding relationships between all posts; and the user forwarding graph is constructed based on the forwarding relationships between all users.

[0034] The false information dataset containing forwarded texts and users is defined as: D = {T1,T2,···,T |D|}, where T i is the i-th sample in the data set, P i A collection of posts. This is the original post. It is T i The jth forwarded post in |P i | indicates the number of posts; According to T i The post forwarding graph constructed by the forwarding relationship between all posts in is defined as P i , Represent the node set and edge set of the post forwarding graph respectively, where s represents the starting point, t represents the end point, represents the edge between the post starting point s and the post ending point t, that is, the forwarding relationship between posts; U i is a collection of users. Is the original user, It is T i The jth forwarding user in the , which includes ordinary forwarding users (users who actively participate in forwarding) and designated forwarding users (@ users); According to T i The user forwarding graph is constructed from the forwarding relationships between all users in the , which includes common forwarding and specified forwarding relationships, and is defined as U i , Represent the node set and edge set of the user forwarding graph respectively, where represents the edge between the user starting point s and the user end point t, that is, the forwarding relationship between users. It should be noted that and The topological structures of the two networks are different, mainly due to the following two reasons: (1) Contains the specified forwarding user and specified forwarding relationship. Since the forwarded text does not involve the specified forwarding relationship, It does not include nodes and edges related to the specified forwarding; (2) There may be multiple forwarding posts between users. If the content of the forwarded posts is different, Multiple forwarding edges and multiple post nodes will be constructed to record multiple forwardings of the same post by the user. The content of the forwarded post is not considered, and multiple forwardings between users are abstracted into one edge. Indicates T i Posts Collection P i and user set U i The bipartite graph between Represents the edge between the post starting point s and the user end point t, that is, the subordinate relationship between the post and the user.

[0035] like Figure 2 As shown, a piece of false information T i It can be viewed as a heterogeneous graph, which contains two types of nodes (users, posts) and three types of edges (user-user, post-post, user-post). i There are corresponding category labels y i ∈{0,1} (0 represents non-false information, 1 represents false information), Y={y1,y2,···,y i ,y |D|} represents the set of sample labels, and |D| represents the number of samples. So far, the identification of false information based on the dissemination content can be defined as a supervised classification task: given a training set D train ={T train ,Y train}, T train represents the training sample, Y train Represents the labels corresponding to the training samples and the test set T test , how to use the training set D train Learn a false information classifier based on T test Predict the authenticity label Y of false information test .

[0036] Step S120, constructing a user-post interaction network model and using the false information data set for training; wherein the user-post interaction network model includes: a feature extraction module for initializing and encoding text information, users, and forwarding relationships; a block heterogeneous graph attention network module for simultaneously modeling user and text information and the relationship between the two; a gating mechanism fusion module for root node enhancement; and a false information classifier for predicting whether a sample is false information.

[0037] Here, the framework of the user-post interaction network model is as follows Figure 3As shown in the figure, node feature encoding is performed in the feature extraction module, and the block heterogeneous graph attention network module is divided into three parts to simultaneously perform adjacent node feature aggregation operations, and the user features, text features and root node features are fused through the gating mechanism, and the fusion results are input into the false information classifier for false information identification.

[0038] Step S130, predicting the false information data to be tested by using the trained user-post interaction network model.

[0039] Here, the trained user-post interaction network model is used on the false information data to be tested, i.e., the test set, to predict the authenticity label of the false information.

[0040] In the embodiment of the present invention, a false information dataset containing social context information is first collected, including designated forwarding relationships and forwarding user information, to further improve the social context information. Then, a multi-entity and multi-relation user-post interaction network is proposed to achieve information exchange and deep fusion between heterogeneous nodes. In addition, the original user and original text are fused again through the gating mechanism to enhance the role of the original content in false information detection and further improve the accuracy of false information detection.

[0041] In some embodiments, the block heterogeneous graph attention network includes three parallel parts: a user forwarding network U-GAT, a post forwarding network P-GAT, and a fusion exchange network UP-GAT: U-GAT is used to model the user forwarding graph and aggregate features between user nodes; P-GAT is used to model the post forwarding graph and aggregate features between text nodes; UP-GAT acts as a bridge between user nodes and text nodes to aggregate heterogeneous types of nodes, playing a role in information exchange and feature fusion.

[0042] Here, in order to fine-grainedly fuse the feature representations of user and text nodes, according to the characteristics of the user-post interaction network, a heterogeneous graph attention network is used to simultaneously model user and text information, and it is divided into three parts: U-GAT, UP-GAT, and P-GAT. When aggregating neighbor node features, U-GAT and P-GAT aggregate homogeneous neighbor node features, while UP-GAT aggregates heterogeneous neighbor node features. The three adaptively adjust the influence of neighbor nodes on the central node through the attention mechanism.

[0043] In some embodiments, after the U-GAT and P-GAT aggregation operations of the current layer are completed, the user node features and the text node features are updated, and the information exchange between the user nodes and the text nodes is realized through the UP-GAT of the next layer; after multiple iterations, the updated user features and text node features are respectively input into the fully connected layer, the average pooling layer and the maximum pooling layer for node feature convergence operations to obtain the corresponding user feature vector representation and text feature vector representation.

[0044] Here, taking P-GAT as an example, the specific process of aggregating neighbor node features in the GAT network is introduced.

[0045] The edges of the post forwarding network P-GAT are extracted from the post forwarding relationship. The node feature matrix of P-GAT is: M+1 is the total number of text nodes. The central node features of the text nodes and 1-hop neighbor node characteristics Calculate the attention score α between the two ij , the formula is as follows:

[0046]

[0047] in, is a learnable row vector, W c and W n are the parameter matrices used by the center node and neighbor node to transform the feature dimension, || represents the concatenation operation, LeakyReLU(·) is the activation function, and N(i) is the set of neighbor nodes. The attention scores of all neighbor nodes are e ij After softmax normalization, we get a set of attention score sequences [α i1 ,α i2 ,···,α i|N(i)| ] is used to indicate the influence of each neighbor node on the central node. Finally, the feature representation of the central node is updated by weighted aggregation of neighbor nodes:

[0048]

[0049] where σ is the ReLU(·) activation function, It is the updated central node feature of the text node. After the first layer of aggregation, P-GAT obtains the updated text node feature matrix

[0050] The user node feature matrix can be expressed as N represents the number of user nodes. U-GAT also contains the specified forwarding relationship. Therefore, U-GAT weights and aggregates its own node features. and neighbor node features It can be expressed as:

[0051]

[0052] in, is the weight of the edge corresponding to the forwarding relationship between user nodes i and j, It is the central node feature after the user node is updated.

[0053] After the aggregation operation of U-GAT, the updated user node feature matrix can be obtained Similarly, after the aggregation operation of UP-GAT, the updated node feature matrix can be obtained Among them, each node feature and The node features of heterogeneous types are weighted and aggregated respectively, thus realizing the information exchange and sharing between users and texts. The updated node feature matrix is ​​spliced ​​and fused as shown in the following formula:

[0054]

[0055] Thus, the user node feature matrix containing the forwarded text information is obtained and the text node feature matrix containing the forwarding user information M and N represent the number of text and user nodes respectively. and They are used as the input of the second layer of the block heterogeneous graph attention network. Finally, after k layers of operations, the fine-grained fusion of user-user, text-text and user-text information is achieved.

[0056] In some embodiments, in U-GAT, if there are multiple forwardings between two users, the weights of the edges representing the user forwarding relationship will be accumulated according to the following formula:

[0057]

[0058] in, is the weight of the edge corresponding to the forwarding relationship between user nodes i and j, r1 and r2 are the number of common forwarding and specified forwarding between user nodes i and j, respectively, β∈(0,2] represents the weight of the specified forwarding edge, and the weight of the common forwarding edge is set to 1 by default.

[0059] Here, two types of user forwarding, ordinary forwarding and designated forwarding, are considered, corresponding weights are assigned to different types of forwarding edges, and deep fusion of heterogeneous node features is achieved through user-post interaction.

[0060] In some embodiments, in the feature extraction module, corresponding feature transformation methods are used for three types of user features, namely numerical, Boolean and discrete, to encode feature representations of user nodes from multiple angles; heterogeneous layer normalization is used to normalize all user features of each sample; all samples in the same batch are stacked to obtain final user features; and the final user features are input into the fully connected and activated layers to obtain a user node encoding representation of a specific dimension.

[0061] Here, the embodiment of the present invention selects the forwarding user features listed in Table 1 as the feature representation of the user node, which can be roughly divided into three types: numerical type, Boolean type and discrete type, and adopts corresponding feature transformation methods to enrich the feature representation of the user node, thereby improving the robustness of the model.

[0062] Table 1 User node feature transformation

[0063]

[0064] In the model training stage, the heterogeneous layer normalization (HeteroLayerNorm) is first used to normalize all forwarding user features of each sample, as shown in the following formula:

[0065]

[0066] in, It's U i The feature matrix of all users in , μ and σ 2 are the vectors composed of the feature mean and feature variance of the user, are learnable parameters and biases, d is the dimension of the user feature, and · is the Hadamard product. This normalization operation can reduce the size gap between features, thereby preserving the size relationship between different users, and adaptively adjust the size relationship between features through learnable parameters.

[0067] Then, all samples in the same batch are stacked to obtain user features And input the full connection and activation layer to get the user node encoding representation X of a specific dimension u :

[0068]

[0069] Among them, ReLU(·) is the activation function, W u and b u is a learnable parameter matrix and bias, where W u and b u is a shared parameter for samples in the same batch.

[0070] In some embodiments, the three types of user features, namely numerical, Boolean and discrete, are respectively transformed using corresponding feature transformation methods, and the feature representation of the user node is encoded from multiple angles, including: for numerical user features, multiple normalization methods are used for processing: the user's relative registration days and forwarding delay time are normalized to zero mean; the number of friends, number of speeches and number of fans are normalized to zero mean after taking the natural logarithm; the number of characters in the authentication reason and the number of characters in the personal profile are normalized using the maximum and minimum values; for Boolean type features, binarization is performed; for discrete type features, the corresponding feature representation is obtained by conversion through one-hot encoding or embedding; wherein, the user's relative registration days are determined based on the difference between the user's forwarding time and the user's registration time; the forwarding delay time is determined based on the difference between the user's forwarding time and the post publishing time.

[0071] Here, for the features of numerical type, multiple normalization methods are used for processing. Through experiments, it is found that normalizing the relative registration days (regist_day) and forwarding delay time (delay_second) features to 0 mean; taking the natural logarithm of the number of followers, fans and Weibo features and then normalizing them to 0 mean; and normalizing the number of characters in the authentication reason (reason_len) and the number of characters in the personal profile (description_len) using the maximum and minimum values ​​(Max-Min) can achieve more stable training results and better performance.

[0072] In particular, due to the large time span of the Weibo23 data set, the present invention found in the study that using the user's relative registration days is more in line with the actual situation than the absolute registration days, and also achieves better experimental results. The calculation method is as follows: user relative registration days = user forwarding (publishing) time - user registration time.

[0073] Statistics show that the relative registration days (regist_days) of users who forward false information are generally smaller than those who forward non-false information, indicating that there are more "new" users who forward false information, and more "old" users who forward non-false information. In addition, a large number of studies have proven the role of forwarding timing in detecting false information. The present invention incorporates dynamic user forwarding timing features (forwarding delay time) into the static propagation structure, so that the model can capture the dynamic changes of the propagation structure, where forwarding delay time = user forwarding time - post publishing time; the unit of forwarding delay time is seconds, and segmented feature transformation is performed to construct a more discriminative forwarding timing feature.

[0074] In addition, in order to construct forwarding timing features with better recognition and generalization performance, on the one hand, the forwarding time is discretized into segments, and on the other hand, is_day and is_week features are constructed to indicate whether the user forwards during the day and on weekdays, thereby fully mining the forwarding timing information.

[0075] In some embodiments, in the feature extraction module, for the text content, a pre-trained BERT is used as a fixed text feature extractor to extract all the original text and forwarded text after segmentation, and the output vector of the classification mark is used as the feature representation of each text node.

[0076] Here, the text content is a powerful basis for distinguishing the truth from false information. The original text of false information is more likely to contain misleading and exaggerated expressions to attract readers' attention and incite public sentiment. The forwarded text of false information is more likely to contain questioning and opposing content. The embodiment of the present invention uses pre-trained BERT to initialize the text representation. In order to improve the training and reasoning efficiency of the model, the present invention no longer fine-tunes BERT but uses the text feature extractor as a frozen parameter to pre-process all original and forwarded texts and input them into BERT in advance, and uses the output vector of the classification token [CLS] as the feature representation of the text node, as shown in the following formula:

[0077]

[0078] Where S = {w1,w2,···,w n} is a token sequence after a text is preprocessed. is the jth text node in the i-th sample The initialization feature representation of . For the convenience of representation, this chapter omits the superscript i in the subsequent symbolic representation. So far, the text node feature matrix of a sample can be expressed as Where M represents the number of text nodes.

[0079] In some embodiments, the gating mechanism fusion module performs root node feature enhancement through the following process: for a given k-layer graph attention network output and Represent the feature matrices of users and text nodes respectively; and Input the fully connected layer for dimension conversion and Use maximum pooling and average pooling to transform the feature matrix of all user nodes respectively Aggregate and concatenate to get the user feature vector c u , the transformation feature matrix of all text nodes Aggregate and concatenate to get the text feature vector c p ; Concatenate the original user features and original text features of the root node to obtain the root node joint feature representation h root ; Through the gating mechanism, c u 、c p and h root Fusion is performed and the final output is the joint feature representation h out .

[0080] Here, the original content has a greater impact on the spread of false information. For example, false information with rich text content and a large number of followers is more likely to be widely forwarded. In order to make full use of the original content, the root node is enhanced through the gating mechanism after aggregating the node features. u and the text feature vector c p The calculation formula is as follows:

[0081]

[0082] Among them, W lu 、b lu , W lp and b lp are all learnable parameters; secondly, the original user and original text features of the root node are concatenated to obtain the root node joint feature representation h root ; Finally, c u 、c p and h root Fusion is performed to obtain the final output result h out , the calculation formula is as follows:

[0083]

[0084] g=σ(W g [h root ||c u ]) ;

[0085] c up =g*c u +(1-g)*c p ;

[0086] h out =[c up ||h root ];

[0087] in, and Represent the feature representations of user root nodes and post root nodes respectively, W gis the learnable parameter matrix in the gating mechanism, and σ represents the sigmoid function, which is responsible for scaling each component in the gating parameter g to the range of 0 to 1.

[0088] In some embodiments, the false information classifier inputs the joint feature representation into a feedforward neural network, obtains the probability that the sample belongs to a certain category through a normalization function, calculates the classifier learning loss using a cross entropy loss function, and trains the user-post interaction network model through back propagation.

[0089] Here, sample T i The probability of belonging to a certain category,

[0090] p i =softmax(W f h out +b f );

[0091] Among them, W f and b f is a learnable parameter in a feedforward neural network, p i Represents the prediction sample T i is the probability value of false information. This chapter uses the cross entropy loss function to calculate the error and trains the model through back propagation.

[0092]

[0093] Among them, y i is the sample T i The true label of , θ represents the set of all learnable parameters in the model.

[0094] The embodiment of the present invention constructs a false information identification method based on a multi-entity and multi-relationship user-post interaction network, while considering two types of user forwarding, ordinary forwarding and designated forwarding, assigning corresponding weights to different types of forwarding edges, and realizing deep fusion of heterogeneous node features through user-post interaction.

[0095] The present invention collects Chinese data sets containing social background information from 2011 to 2023 on the Weibo platform, including designated forwarding relationships and forwarding user information, to further improve social background information. At the same time, the present invention controls the subject content of false information and non-false information samples to be as close as possible, so as to alleviate the phenomenon of falsely high accuracy in previous studies on Chinese data sets, to a certain extent solve the problem of data bias, and let the false information detection method based on social background play its advantage of weak dependence on false information content. In order to make full use of social background information, the present invention adopts a rich text and post node initialization representation method, and proposes a multi-entity multi-relation user-post interaction network (MUPI) that considers both user and forwarded text entities, and considers both ordinary forwarding and designated forwarding social communication methods, assigns different weights to different types of forwarding edges, and uses a heterogeneous attention network to realize information exchange and deep fusion between user nodes and text nodes.

[0096] In addition, the original user and the original text are integrated again through the gating mechanism to enhance the role of the original content in false information detection, and further improve the accuracy of false information detection. A large number of experiments were conducted on the Weibo23-propagation content sub-dataset, and the results showed that the MUPI model can effectively learn the propagation characteristics of false information and obtain better recognition results than the existing baseline model. The method provided by the present invention significantly improves the performance of short video fake news detection and provides strong technical support for combating the spread of false information.

[0097] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the size of the serial number of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0098] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0099] In the several embodiments provided by the present invention, it should be understood that the disclosed methods can be implemented in other ways. The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0100] The above is only an embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for detecting false information based on a user-post interaction network, characterized in that: include: Construct a false information dataset containing forwarded texts and users, wherein each sample in the false information dataset includes a post set, a post forwarding graph, a user set, a user forwarding graph, and a bipartite graph representing the subordinate relationship between the post set and the user set; wherein the user set includes original users, ordinary forwarding users, and designated forwarding users; the post forwarding graph is constructed based on the forwarding relationship between all posts; and the user forwarding graph is constructed based on the forwarding relationship between all users; Constructing a user-post interaction network model and using the false information dataset for training; wherein the user-post interaction network model includes: a feature extraction module for initializing and encoding text information, users, and forwarding relationships; a block heterogeneous graph attention network module for simultaneously modeling user and text information and the relationship between the two; a gating mechanism fusion module for root node enhancement; and a false information classifier for predicting whether a sample is false information; The false information data to be tested is predicted by the trained user-post interaction network model.

2. The false information detection method according to claim 1, characterized in that: The block heterogeneous graph attention network includes three parts: a parallel user forwarding network U-GAT, a post forwarding network P-GAT and a fusion exchange network UP-GAT: U-GAT is used to model the user forwarding graph and aggregate the features between user nodes; P-GAT is used to model the post forwarding graph and aggregate the features between text nodes; UP-GAT acts as a bridge between user nodes and text nodes to aggregate heterogeneous types of nodes, playing a role in information exchange and feature fusion.

3. The false information detection method according to claim 2, characterized in that: After the U-GAT and P-GAT aggregation operations of the current layer are completed, the user node features and text node features are updated, and the information exchange between the user node and the text node is realized through the UP-GAT of the next layer; After multiple iterations, the updated user features and text node features are respectively input into the fully connected layer, the average pooling layer and the maximum pooling layer for node feature aggregation operation to obtain the corresponding user feature vector representation and text feature vector representation.

4. The false information prediction method according to claim 2, characterized in that: In U-GAT, if there are multiple forwardings between two users, the weights of the edges representing the user forwarding relationship will be accumulated according to the following formula: in, is the weight of the edge corresponding to the forwarding relationship between user nodes i and j, r1 and r2 are the number of common forwarding and specified forwarding between user nodes i and j, respectively, β∈(0,2] represents the weight of the specified forwarding edge, and the weight of the common forwarding edge is set to 1 by default.

5. According to any one of claims 1 to 4, in the feature extraction module, for three types of user features, namely numerical, Boolean and discrete, corresponding feature transformation methods are respectively used to encode the feature representation of the user node from multiple angles; Use heterogeneous layer normalization to normalize all user features of each sample; All samples in the same batch are stacked to obtain the final user features; The final user features are input into the fully connected and activated layers to obtain a user node encoding representation of a specific dimension.

6. The false information prediction method according to any one of claims 1 to 4, characterized in that: In the feature extraction module, for the text content, the pre-trained BERT is used as a fixed text feature extractor to extract all the original texts and forwarded texts after segmentation, and the output vector of the classification mark is used as the feature representation of each text node.

7. The false information prediction method according to claim 5, characterized in that: The three types of user features, namely numerical, Boolean and discrete, are respectively transformed using corresponding feature transformation methods to encode the feature representation of the user node from multiple angles, including: For numerical user features, we use a variety of normalization methods to process them: the user's relative registration days and forwarding delay time are normalized to zero mean; the number of friends, number of speeches, and number of fans are normalized to zero mean after taking the natural logarithm; the number of characters in the authentication reason and the number of characters in the personal profile are normalized using the maximum and minimum values; For Boolean type features, binary processing is performed; For discrete type features, the corresponding feature representation is obtained by conversion through one-hot encoding or embedding; Among them, the user's relative registration days are determined based on the difference between the user's forwarding time and the user's registration time; the forwarding delay time is determined based on the difference between the user's forwarding time and the post publishing time.

8. The false information prediction method according to any one of claims 1 to 4, characterized in that: In the gating mechanism fusion module, root node feature enhancement is performed through the following process: Output of a given k-layer graph attention network and The feature matrices representing users and text nodes respectively; Will and Input the fully connected layer for dimension conversion and Use maximum pooling and average pooling to transform the feature matrix of all user nodes respectively Aggregate and concatenate to get the user feature vector c u , the transformation feature matrix of all text nodes Aggregate and concatenate to get the text feature vector c p ; The original user features and original text features of the root node are concatenated to obtain the root node joint feature representation h root ; Through the gating mechanism, c u 、c p and h root Fusion is performed and the final output is the joint feature representation h out .

9. The false information prediction method according to claim 8, characterized in that: The false information classifier inputs the joint feature representation into a feedforward neural network, obtains the probability that the sample belongs to a certain category through a normalization function, calculates the classifier learning loss using a cross entropy loss function, and trains the user-post interaction network model through back propagation.