Cross-domain false information detection method
By adopting the text feature extraction method of domain-specific word extraction and feature debias in cross-domain false information detection, combining the propagation path map feature extraction of propagation trees and diffusion trees, and optimizing using comparison learning, the problem of poor cross-domain detection performance in the existing technology is solved, and better generalization ability and detection accuracy are achieved.
Patent Information
- Application Number
- CN202510270080.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-26
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing cross-domain false information detection methods perform poorly when facing unseen fields, and mostly ignore the comprehensive impact of text content and communication structure, resulting in poor generalization capabilities of the model.
Text feature extraction methods based on domain-specific word extraction and feature debias are adopted, and information propagation tree and diffusion tree are combined to extract information propagation path map feature. After the text features and structural features are fused, comparison learning is used to optimize to narrow the domain distribution differences between the source domain and the target domain.
By extracting domain-independent shared features and capturing more comprehensive propagation path map structural features, the detection performance of the model in the target domain is improved, the generalization ability is enhanced, and the instability of decision boundaries is reduced.
Smart Images

Figure CN120104796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data mining technology, natural language processing technology and graph neural network technology, and specifically to a cross-domain false information detection method. Background Art
[0002] In recent years, with the development of Internet technology, online interactive platforms represented by social media have rapidly become popular. Social media provides users with convenient channels for obtaining information and communicating, but it has also become a hotbed for the spread of false information. The rapid spread of false information may cause social panic, economic losses and political influence, and even threaten public safety. Especially in breaking news events, false information spreads quickly and has a wide range of impact, and effective detection methods are urgently needed for real-time supervision and response.
[0003] Early false information detection methods mainly rely on traditional machine learning and manually designed features. Such methods generally include the following steps: (1) Screening and extracting features with strong data representation capabilities from training samples. (2) Based on the screened and extracted features, the classification model is trained on the training data set: (3) Using the established model, data from outside the training data set is predicted, and the authenticity of the data is determined through continuous evaluation and optimization. Researchers manually design explicit features (such as message length, user information) or implicit features (such as emotional features, credibility) to characterize false information. Although rule-based or template-based detection methods are more intuitive, they are highly dependent on domain knowledge and are difficult to generalize to cross-domain false information detection scenarios.
[0004] With the rise of deep learning technology, false information detection has gradually introduced neural network models, which has improved detection efficiency and accuracy. It can be divided into detection methods based on tweet content and detection methods based on social context.
[0005] Content-based false information detection methods mainly construct the entire event tweets into time series segments, and then detect them by extracting the latent semantic features of the news event. For example, Chen et al. and Ma et al. divided the tweets of each event into fixed-length and variable-length time series, respectively, and input them into recurrent neural networks (RNNs) and their variants to capture time-language features for false information detection. In addition, Yu et al. stated that RNN-based methods may be biased against the latest input tweets and are not sufficient to detect false information in advance. Therefore, researchers proposed a CNN-based detection model CAMI, which can effectively extract key features in posts and use convolution technology to achieve more refined feature interactions. Social context-based detection methods mainly embed interactions in social networks (such as user comments / likes / forwards) or information diffusion structures into dense vectors through neural networks for subsequent detection. For example, Liu et al. observed that real news and fake news have different propagation patterns, so they used gated recurrent units (GRUs) and CNNs to extract global and local features of forwarding sequences for detecting fake news. On this basis, Lu et al. introduced graph convolutional neural networks (GCNs) to learn more accurate structural information of news propagation paths. Different from previous studies, Bian et al. emphasized the depth of fake news propagation while also emphasizing the wide spread of fake news. Therefore, they proposed the Bi-GCN model to comprehensively describe the spread of fake news in social networks from two directions. In order to systematically integrate various features and make the model more robust, interpretable and accurate, Lao et al. proposed a rumor detection model based on linear and nonlinear propagation fields (RDLNP), which automatically detects rumors propagated linearly and nonlinearly using tweet content, social context and time information. RDLNP consists of four parts: rumor hybrid feature learning (RHFL), nonlinear structure learning (NLSL), linear sequence learning (LSL) and shared feature learning (SFL).
[0006] However, existing methods tend to improve false information detection metrics on specific datasets, are heavily dependent on training datasets, and are not feasible to apply to unseen domains due to differences in event distributions. There is still little research on how to use knowledge learned from historical events to detect fake news in upcoming events.
[0007] In order to solve the problem of cross-domain false information detection, researchers have proposed detection methods based on machine learning and adversarial learning. Tolosi et al. first proposed the problem of cross-domain false information detection in 2016. Their study analyzed false and true information tweets in the PHEME dataset and found that most feature distributions are domain-specific. Even features that are expected to be domain-independent show a distribution in a specific domain, which causes the performance of the model in practical applications to be inferior to that of the experiment. Later, Castelo et al. demonstrated the existence of cross-domain shared features in their 2019 study and proposed a topic-independent classification strategy to identify false information by analyzing network and language features. This method can maintain a high detection accuracy in the face of the dynamics and topic changes of false information in social networks.
[0008] With the development of deep learning technology, deep neural network architecture has brought impressive progress in a variety of machine learning tasks and applications. The efficiency and accuracy of false information detection have also been greatly improved, and the potential high-level feature representation of tweets can be automatically extracted. However, most of these methods ignore the problem of domain distribution differences. The experimental results of Khoo et al. show that when using the PHEME (Twitter platform dataset) dataset for experiments, when the datasets in the training set and the test set do not overlap, the detection effect will drop significantly. After the dataset is re-divided, the experimental results will return to normal levels. This also proves from another perspective that most of the existing methods extract domain-specific features. When faced with unseen domains that do not exist in the training set, the generalization ability of the model is poor and the expected detection effect cannot be achieved. Inspired by the cross-domain adversarial network structure system, Wang et al. proposed the event adversarial neural network (EANN) in 2018, which is an end-to-end framework that can learn domain-invariant features to detect false information in new domains. The EANN model includes a multimodal feature extractor, a false information detector, and a domain discriminator. The domain discriminator removes event-specific features and retains shared features, achieving good generalization ability. However, EANN does not consider the difference in the importance of tweets to the target domain. In response to this, Ding et al. proposed MetaDetector in 2021, which evaluates the impact of source domain tweets on the target domain through a pseudo-domain discriminator, thereby improving detection efficiency. At the same time, Yuan et al. also proposed DAGA-NN in 2021, which uses graph attention neural networks and domain sharing features for cross-domain false information detection.
[0009] In practical applications, the content of social media is domain-diverse, and data in new domains usually lacks labels. Traditional detection methods perform poorly when facing unseen domains. Adversarial learning-based methods align the feature distributions of the source domain and the target domain through domain discriminators, effectively improving the generalization ability of the model. However, existing methods mostly focus on extracting domain-shared features, ignoring the combined influence of text content and communication structure. In addition, the category inconsistency problem in the feature alignment process may lead to instability of the decision boundary.
[0010] Based on the above challenges, the present invention combines the ideas of contrastive learning and feature debiasing, integrates text sharing features and propagation path structure features, and through the semantic alignment of source domain and target domain data, can effectively alleviate the impact of unlabeled data and improve the detection performance of the model in the target field.
[0011] Compared with the prior art, the differences are as follows:
[0012] Technical comparison with the paper "Cross-Domain Rumor Detection Based on Dual-Modal Domain Alignment"
[0013] The paper "Cross-Domain Rumor Detection Based on Dual-Modal Domain Alignment" discloses a cross-domain false information detection method based on dual-modality, which includes the following steps: using the TextRank method to identify domain-specific keywords in the source domain and target domain texts, and then performing data enhancement based on keyword replacement to obtain modified sentence pairs, and then using Bi-LSTM network encoding to obtain the text features of the sentences; using the Bi-GCN network to obtain the propagation structure features of the tweets, and then using contrastive learning to optimize the structural features; and finally performing classification.
[0014] The present invention proposes a cross-domain false information detection method, which is innovative and practical. By extracting domain features and debiasing domain features of domain texts, universal text features that are independent of the domain are obtained, which greatly reduces the interference caused by low-quality data and inflexible data enhancement rules, and is more robust; by constructing a bidirectional propagation tree and a diffusion tree to capture the interactive relationship of tweets, the propagation path graph feature is extracted; after fusing the text features and the propagation path graph features, contrastive learning is used to optimize them before classification.
[0015] Different from the method disclosed in the paper "Cross-Domain Rumor Detection Based on Dual-Modal Domain Alignment", the present invention trains a domain-text fusion model and a domain model on the basis of the FLAN-T5-large model to perform feature extraction and domain debiasing. It does not rely on the quality of enhanced data and enhances the generalization ability of the model. After fusing the text features and structural features, contrastive learning is used for optimization, which helps the model learn the key features to distinguish between false information and real information and improve the accuracy of detection. Summary of the invention
[0016] In view of the above problems, the present invention proposes a cross-domain false information detection method, which extracts text features based on domain-specific words to obtain text features without domain bias, extracts features of the information propagation path graph based on propagation trees and diffusion trees, and fuses the features and uses contrastive learning to optimize the features.
[0017] To achieve the above object, the technical solution adopted by the present invention is:
[0018] A cross-domain false information detection method comprises the following steps:
[0019] (1) Text feature extraction method based on domain-specific word extraction and feature debiasing;
[0020] A text feature extraction method based on domain-specific word extraction is used to extract domain text from source and target domain tweets. During the training process, the domain-text fusion model is trained using the original text content, and the domain model is trained using the domain text. In the inference phase, the domain features are subtracted from the original text features to obtain text features without domain information.
[0021] (2) Propagation structure feature extraction module;
[0022] This module combines the tweet content representation obtained in step (1) to learn the propagation characteristics of tweets from the perspective of tweet propagation, based on the interactive relationship between tweets. During the training process, the propagation tree and diffusion tree of information are constructed from two different directions, top-down and bottom-up. The propagation characteristics and diffusion characteristics of information are then extracted from the two directions through the GAT model. Finally, the two features are fused to obtain the complete structural representation of the tweets.
[0023] (3) Feature fusion and label prediction module;
[0024] The shared features of the tweet texts in the source domain and the target domain obtained in step (1) and the structural features of the tweet propagation path graphs in the source domain and the target domain obtained in step (2) are concatenated to obtain a more comprehensive final vector representation of the tweets. The spherical K-means algorithm is used to perform unsupervised clustering of the target domain tweets to generate predicted labels for the target domain tweets.
[0025] (4) Cross-domain feature alignment module based on contrastive learning;
[0026] Using the predicted label information and the final vector representation of the tweet obtained in step (3), select data of the same category across the source and target domains as positive pairs, and data of different categories as negative pairs. Re-edit the distance between features of different categories in the feature space through comparative learning, thereby narrowing the domain distribution difference between the source and target domains and obtaining the final optimized features.
[0027] (5) False information detection module;
[0028] Based on the optimized features obtained in step (4), a multi-layer perceptron is used to classify tweets by calculating the authenticity score vector of the tweets, thereby realizing false information detection in the target domain.
[0029] (6) System function demonstration;
[0030] According to the cross-domain false information detection model proposed above, a prototype model for real-time monitoring of cross-domain false information is designed based on the datasets PHEME and Misinfdect collected by the platform. The prototype system consists of two modules, namely the online module and the offline module. The function of the offline module is data preprocessing and model training, and the function of the online module is to give prediction results based on user input, visual analysis and display, etc.
[0031] As a further improvement of the present invention, step (1) comprises the following specific steps:
[0032] (11) extracting special words from the corresponding sentences in the source domain and the target domain as described in step (1), specifically, using the TF-IDF method without labeled data and pre-training to identify and mark the domain-related special words in the tweet texts of the source domain and the target domain;
[0033] (12) The original text content is used to train the domain-text fusion model described in step (1), and the domain text is used to train the domain model. Specifically, the FLAN-T5-large model is trained as a domain model based on the special words of the domain text extracted in (12) to capture the interference caused by the domain information on text classification, and then another FLAN-T5-large model is trained as a domain-text fusion model based on the original text. In the inference stage, the original text and the special words corresponding to the original text are formed into a text pair, and the original text vector features encoded by the domain-text fusion model are subtracted from the special word vector features encoded by the domain model to obtain text information based only on shared features.
[0034] As a further improvement of the present invention, step (2) specifically comprises the steps of:
[0035] (21) The text content vector representation z of each tweet obtained in step (1) i , assuming that the propagation tree structure of the tweet is G i , constructing the information propagation tree in the top-down GAT network Constructing a diffusion tree of information in a bottom-up GAT network Use adjacency matrix to represent the interaction between different tweets;
[0036] (22) LeakyReLU is used as the activation function of the activation layer of the GAT network, and a multi-head attention mechanism is adopted to capture the representation of node i from different angles. Then, the Attention mechanism is used to learn the influence of different tweet nodes on the final result. Finally, the features in the two directions are concatenated to obtain the final structural representation.
[0037] As a further improvement of the present invention, in step (3), the vector features of the shared features of the tweet texts and the propagation path graph structure in the source domain and the target domain are obtained from step (1) and step (2), and the fusion feature O is obtained after splicing. i ,To generate predicted labels for tweets in the target domain, the spherical K-means algorithm is used, the number of clusters is set to the number of classes M, and the class prototypes from the source domain are used as the initial clusters. The cosine similarity is used to measure the target features. The distance between the center of the m-th cluster is used to classify the tweets into two categories.
[0038] As a further improvement of the present invention, step (4) obtains the fusion feature O from step (3) i , we select data of the same category from the source domain and the target domain as positive pairs, and data of different categories as negative pairs, and use contrastive learning to align the features of tweets from the two domains to obtain the final representation of the tweets.
[0039] Beneficial effects: Compared with the prior art, the present invention adopts the above technical solution and has the following advantages:
[0040] (1) In order to better extract cross-domain features from text content, we first detect domain-related special words in all tweet text content, and then use the extracted domain-specific words and the original text content to form text pairs. Then, we use the domain model and the domain-text fusion model to extract features, and then perform feature debiasing, so that the model can capture shared features that are not related to the domain.
[0041] (2) While taking into account the extraction of shared features of tweet text content, the structural features of the propagation path graph of tweets throughout the entire propagation process are also considered, capturing more shared features of false information in different fields and obtaining a more comprehensive final representation of tweets;
[0042] (3) The contrastive learning method is used to align the fused features across domains. The distance between features of different categories is re-edited in the feature space, thereby narrowing the domain distribution difference between the source domain and the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is the principle diagram of the present invention;
[0044] Figure 2 It is a model diagram of a propagation path graph structural feature extraction module of the present invention;
[0045] Figure 3 This is a model diagram of the feature fusion module of the present invention;
[0046] Figure 4 This is a model diagram of the label prediction module of the present invention;
[0047] Figure 5 Generate schematic diagrams for cross-domain positive and negative examples of the present invention;
[0048] Figure 6 It is a schematic diagram of cross-domain feature alignment of the present invention;
[0049] Figure 7 It is the overall system framework diagram of the present invention;
[0050] Figure 8 This is the PHEME dataset comparison experiment result page of the prototype system;
[0051] Fig. 9 This is the Misinfdect dataset comparison experiment result page of the prototype system;
[0052] Fig.10 This is the real-time monitoring page diagram of the prototype system;
[0053] Fig.11Page diagram showing the feature alignment process for the prototype system. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.
[0055] The principle diagram of the present invention is as follows Figure 1 As shown in the figure, the propagation path graph structure feature extraction module model diagram is as follows Figure 2 As shown in the figure, the feature fusion module model is as follows Figure 3 As shown, the label prediction module model diagram is as follows Figure 4 As shown, the schematic diagram of cross-domain positive and negative example generation is as follows Figure 5 As shown, the schematic diagram of cross-domain feature alignment is as follows Figure 6 shown.
[0056] like Figure 7 As shown, a cross-domain false information detection method described in the present invention includes the following steps:
[0057] (1) Text feature extraction method based on domain-specific word extraction and feature debiasing
[0058] The main function of this module is to perform initial characterization and shared feature extraction on the tweet text content. Before entering the tweet text shared feature extraction module, it is necessary to use the TF-IDF method without labeled data and pre-training to identify and mark the domain-related special words in the source domain and target domain tweet texts in the text data preprocessing stage.
[0059] Assume that the text content of a tweet in the domain is Where i represents the text of the i-th tweet in the source domain, n represents the number of words in the tweet, and the words marked with overlines Indicates the domain-related special words in the tweet. After using TF-IDF to extract domain features, the domain features of the tweet will be obtained. The original tweet text and tweet domain features form text pairs FLAN-T5-large is trained as a domain-text fusion model using the original tweet text to capture the classification of text under the combined effect of domain information and domain-independent information, including the prediction bias caused by domain information, as shown in formula (2). FLAN-T5-large is trained as a domain model using tweet domain features to capture the classification of text under the effect of domain information alone, as shown in formula (3).
[0060]
[0061] Among them, Odomain is the domain feature, O domain,general is the original text feature containing domain features, It is denoted as cross entropy loss.
[0062] By using the hyperparameter α, the domain features are subtracted from the original text features to obtain irrelevant features without interference from domain information. Through this module, the representation of tweets is learned, and the information of each tweet is represented by Learn
[0063] (2) Propagation structure feature extraction module
[0064] The text content vector z of each tweet obtained in step (1) is represented as i , assuming that the propagation tree structure of the tweet is G i , constructing the information propagation tree in the top-down GAT network Constructing a diffusion tree of information in a bottom-up GAT network Use the adjacency matrix to represent the interaction between different tweets:
[0065]
[0066] in, represents the embedding representation of each node in the l-th layer of the graph neural network, and The specific calculation formula of the attention coefficient is as follows:
[0067]
[0068] in, represents the tweet node i and its neighbor node set, h l,i and h l,j represents the embedding representation of node i and node j in the l-th layer graph attention network, It is used to calculate the attention between nodes, and LeakyReLU is the activation function used in the activation layer. After aggregating the relevant tweet information, each tweet node uses the attention mechanism to calculate the representation of the node i in the l+1th layer:
[0069]
[0070] In order to ensure that GAT is more stable, a multi-head attention mechanism is used, and the final representation of node i is:
[0071]
[0072] In the formula, K is the number of multi-head attention heads, and || is the connection operator. Finally, the Attention mechanism is used to learn the influence of different tweet nodes on the final result. The calculation formula is similar to formulas (6) and (7). After obtaining the representation of each tweet node in the last layer of the network, the representation of the propagation tree structure H can be obtained through the attention mechanism. TD The calculation process of diffusion features is similar to that of propagation features. Finally, the features in the two directions are concatenated to obtain the final structural representation.
[0073]
[0074]
[0075] H=H TD ||H BU (12)
[0076] (3) Feature fusion and label prediction module
[0077] After obtaining the vector representation of the shared features of tweet texts and the propagation path graph structure in the source domain and the target domain through steps (1) and (2), the two features can be concatenated to obtain a more comprehensive final vector representation of tweets O i ,|| is the connection operation operator.
[0078] O i =V i ||H i (13)
[0079] When generating positive and negative examples across domains, it is necessary to input the categories of source domain tweets and target domain tweets. During the training process, the ground-truth labels of the target domain are not available, so this module needs to generate predicted labels for the target domain tweets. To generate predicted labels for the target domain tweets, an unsupervised classification algorithm is required, and spherical K-means is a widely used K-means clustering algorithm. This module uses cosine similarity between texts instead of Euclidean distance to improve computational efficiency, where the calculation of cosine similarity is shown in formula (14):
[0080]
[0081] Since the spherical K-means clustering algorithm is sensitive to initialization, using randomly generated clusters cannot guarantee the relevant semantics of predefined categories. To solve this problem, this method sets the number of clusters to the number of classes M and uses the class prototypes from the source domain as the initial clusters. Specifically, this method first calculates the centroid of the source domain tweets in each category as the corresponding class prototype, and the initial cluster center C of the mth class is m Defined as:
[0082]
[0083] Given the features from the target domain, spherical K-means clustering is then performed using centers initialized from source domain tweets. When determining the label of each target sample, the cosine similarity is used to measure the target features. The distance between the center of the mth cluster and the center of the mth cluster. After clustering, each sample in the target domain There will be a corresponding prediction label In order to ensure the validity of the predicted labels, this method removes low-quality samples from their designated cluster centers and selects samples with higher confidence.
[0084] The algorithm description is shown in Table 1:
[0085] Table 1 Feature fusion and label prediction algorithm
[0086]
[0087] (4) Cross-domain feature alignment module based on contrastive learning
[0088] This module aims to form positive and negative pairs through contrastive learning to align the features of the source domain and the target domain and learn domain-shared features. Features After l 2 After regularization, it is used as an anchor point to form a positive pair with the same type of samples in the source domain, and its characteristics are recorded as This method expresses the cross-domain contrast loss as:
[0089]
[0090] in, The source domain has the same label The number of samples, Indicates the selection of samples. sim is the cosine similarity calculation function, and τ represents temperature. When extracting text shared features, the tweet texts in the source domain and the target domain are projected into the same semantic space. However, when extracting propagation structure features, if there is no constraint of contrast loss, the features extracted by the network tend to be domain-specific features and cannot be applied to the target domain. For the data in the target domain, the contrast classification loss is also calculated in the same way as formula (16).
[0091]
[0092] Among them, N t is the total number of tweets in the target domain in a Batch, The source domain has the same label Therefore, our method projects source and target tweets belonging to the same class closer than source and target tweets belonging to different classes.
[0093] The cross-domain contrastive loss improves the performance by using tweets from both domains to align features in a bidirectional manner. Finally, the cross-domain contrastive loss is combined with the standard cross entropy loss L ce Combined, this method obtains the final training objective function:
[0094]
[0095] (5) False Information Detection Module
[0096] The main function of this module is to evaluate the authenticity of the original tweets in the target domain based on the final representation of the tweets obtained above. Its main component is a multi-layer perceptron, which performs classification prediction by calculating the authenticity score vector of the tweets.
[0097] (6) System function display
[0098] The system functions are mainly divided into offline modules and online modules. The online module is mainly used for online false information detection, including result display module, user interaction module and visual analysis module. The offline module performs model training after data collection and data preprocessing.
[0099] The data sets used in this method include the English data set PHEME and the Chinese data set Misinfdect, and the MySQL database is used to store the false information data to be detected.
[0100] The data preprocessing module preprocesses the text and structure information in the dataset, restores the English abbreviations in the PHEME dataset, converts the text content to lowercase, and uses the stop word list to filter the stop words in it; the stop words in the Misinfdect dataset are filtered, and then the jieba word segmentation tool is used for word segmentation, and then the network elements in the text content are uniformly converted into a consistent representation form, such as links are converted into <url>, @user converted to <user>, words starting with # are uniformly converted to <hashtag>, emojis are converted to <emoji>For the processing of the propagation structure, a top-down propagation tree and a bottom-up diffusion tree are constructed respectively, with the original tweet as the root node, and the processing of the propagation structure is completed according to the forwarding and commenting relationship.
[0101] In the model training module, for each detection task of the PHEME dataset, this method selects 800 source domain tweets with true and false labels and 800 target domain tweets without true and false labels for training, 100 target tweets for verification, and 200 target tweets for testing. For each detection task of the Misinfdect dataset, this method randomly selects 1300 source domain tweets with true and false labels and 1300 target domain tweets without true and false labels for training, 300 target domain tweets for verification, and 400 target domain tweets for testing.
[0102] The online module mainly provides users with real-time detection functions and performs real-time visual analysis of the content input by users. Through this module, users can obtain experimental results on two benchmark data sets, so as to have a deeper understanding of the detection capabilities of the model proposed in this invention. The results are shown as follows: Figure 8 , Fig. 9 More importantly, users can input the relevant Weibo IDs to be detected through the interface provided by the system. After receiving the tweet ID, the system will crawl the specific tweets on the corresponding website and return the detection results of the model. In addition, the system will detect keywords based on the specific content of the tweets and collect some tweets in related fields. Based on the collected tweets, the system will perform tweet length analysis, sentiment polarity analysis and high-frequency word cloud analysis. The results are shown below. Fig.10 In order to allow users to more intuitively experience the process of cross-domain false information detection, the system will also dynamically display the process of feature migration. The effect is shown as follows: Fig.11 .
[0103] The above description is only a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent change made based on the technical essence of the present invention still falls within the scope of protection required by the present invention.< / emoji> < / hashtag> < / user> < / url>
Claims
1. A cross-domain false information detection method, characterized by: The following steps are involved: (1) Text feature extraction method based on domain-specific word extraction and feature debiasing; A text feature extraction method based on domain-specific word extraction is used to extract domain-shared features from tweet texts. The domain model is trained using the domain-shared features, and the domain-text fusion model is trained using the original text. In the inference phase, the domain features are subtracted from the domain-text features to obtain a text feature vector without domain information. (2) Propagation structure feature extraction module; This module combines the tweet content representation obtained in step (1) to learn the propagation path graph features of tweets based on the interactive relationship between tweets at the tweet propagation level. During the training process, the propagation tree and diffusion tree of information are constructed from two different directions, top-down and bottom-up. Then, the propagation features and diffusion features of information are extracted from the two directions respectively through the GAT model. Finally, the two features are fused to obtain the complete structural representation of the tweet. (3) Feature fusion and label prediction module; The shared features of the tweet texts in the source domain and the target domain obtained in step (1) and the propagation structure features of the tweets in the source domain and the target domain obtained in step (2) are concatenated to obtain a more comprehensive final vector representation of the tweets. The spherical K-means algorithm is used to perform unsupervised clustering of the tweets in the target domain to generate predicted labels for the tweets in the target domain. (4) Cross-domain feature alignment module based on contrastive learning; Using the predicted label information and the final vector representation of tweets obtained in step (3), select data of the same category as positive pairs in the source domain and the target domain, and data of different categories as negative pairs. Re-edit the distances of different categories in the feature space through contrastive learning, thereby narrowing the domain distribution difference between the source domain and the target domain, and obtaining the final optimized features. (5) False information detection module; Based on the optimized features obtained in step (4), a multi-layer perceptron is used to classify tweets by calculating the authenticity score vector of the tweets, thereby realizing false information detection in the target domain. (6) System function demonstration; According to the cross-domain false information detection model based on contrastive learning proposed above, a prototype model for real-time monitoring of cross-domain false information is designed based on the datasets PHEME and Misinfdect collected from the platform. The prototype system consists of two modules, namely the online module and the offline module. The function of the offline module is data preprocessing and model training, and the function of the online module is to give prediction results based on user input, visual analysis and display, etc.
2. A cross-domain false information detection method according to claim 1, characterized in that: Step (1) includes the following specific steps: (11) extracting special words from the corresponding sentences in the source domain and the target domain as described in step (1), specifically, using the TF-IDF method without labeled data and pre-training to identify and mark the domain-related special words in the tweet texts of the source domain and the target domain; (12) In step (1), the domain model is trained using domain shared features, and the domain-text fusion model is trained using the original text. Specifically, two T5-large models are trained separately to perform text classification tasks on content containing only domain text information and original text information.
3. A cross-domain false information detection method according to claim 1, characterized in that: Step (2) specifically includes the steps of: (21) The text content vector representation z of each tweet obtained in step (1) i , assuming that the propagation tree structure of the tweet is G i , constructing the information propagation tree in the top-down GAT network Constructing a diffusion tree of information in a bottom-up GAT network Use adjacency matrix to represent the interaction between different tweets; (22) LeakyReLU is used as the activation function of the activation layer of the GAT network, and a multi-head attention mechanism is adopted to capture the representation of node i from different angles. Then, the Attention mechanism is used to learn the influence of different tweet nodes on the final result. Finally, the features in the two directions are concatenated to obtain the final structural representation.
4. A cross-domain false information detection method according to claim 1, characterized in that: In step (3), the vector features of the shared features and propagation structures of the tweet texts in the source domain and the target domain are obtained from step (1) and step (2), and the fusion features O are obtained after concatenation. i ,To generate predicted labels for tweets in the target domain, the spherical K-means algorithm is used, the number of clusters is set to the number of classes M, and the class prototypes from the source domain are used as the initial clusters. The cosine similarity is used to measure the target features. The distance between the center of the m-th cluster is used to classify the tweets into two categories.
5. A cross-domain false information detection method according to claim 4, characterized in that: The step (4) obtains the fusion feature O from step (3) i , we select data of the same category from the source domain and the target domain as positive pairs, and data of different categories as negative pairs, and use contrastive learning to align the features of tweets from the two domains to obtain the final representation of the tweets.
Citation Information
Patent Citations
Time hypergraph neural network rumor detection method and model based on event fusion
CN118194085A
Multi-modal false news detection method based on dynamic propagation social graph
CN118568261A
Technologies for flexible tree-based lookups for network devices
US20190318022A1
Decentralized Cybersecure Privacy Network For Cloud Communication, Computing And Global e-Commerce
US20190386969A1
Extracting features from screen images for task mining
US20240211836A1