Multi-task Rumor Detection Model and Method Based on Parallel Information Sharing Module
By introducing a parallel information sharing module and a hierarchical attention mechanism in the multi-task rumor detection model, combining text content, writing style and social background information, the problem of inefficient data sharing and feature extraction in the existing technology is solved, and more efficient information utilization and detection effects are achieved.
Patent Information
- Application Number
- CN202211312019.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-10-25
AI Technical Summary
Existing multitasking rumor detection methods have problems of inefficiency and underutilization of information in data sharing and feature extraction.
A multi-task rumor detection model based on parallel information sharing module is proposed. Through a hierarchical attention mechanism and parallel information sharing structure, combining text content, writing style and social background information, a multi-task rumor detection model with parallel information sharing structure is formed.
This model can make more efficient use of data information, reduce information interference, improve model convergence speed and detection accuracy, and overcome the problem of inefficient data sharing and feature extraction in existing methods.
Smart Images

Figure CN115687795B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a multi-task rumor detection model and method based on a parallel information sharing module. Background Art
[0002] The main technologies used in the rumor detection method based on a parallel information sharing module are Deep Learning and Multi-Task Learning (MTL). Its main purpose is to utilize the writing style features, text content features, and social background information of the source tweet and its related tweets (here referring to comment or retweet tweets). The social background information consists of the number of retweets, likes of the tweet, and some relevant information of the user.
[0003] To alleviate the impact of insufficient data volume on the deep learning model, the MTL model was designed. It can help users reuse existing data for learning and reduce the cost of manually labeled data. Since the MTL model utilizes data from different tasks, it can learn a more robust, general, and powerful multi-task model, thereby better realizing knowledge sharing between tasks, improving model performance, and reducing the risk of overfitting. Therefore, the performance of the MTL model is often better than that of the single-task model. Inspired by MTL, multi-task rumor detection methods emerged. These multi-task rumor detection methods often combine the stance detection task and the rumor detection task to form an MTL model because the comments on the information confirmed as rumors often contain more skeptical or negative content than real information.
[0004] The detection method disclosed in the Chinese patent "A Multi-Task Rumor Detection Method Based on a Bidirectional Propagation Graph" (Application No.: 202110454550.0) is used to detect whether a social network post is a rumor and the stance of the comment information. The patent generates a text feature matrix, a user feature matrix, and a text statistical feature matrix based on the post content, then constructs a bidirectional propagation graph of the rumor, extracts the propagation features of the rumor by calculating the bidirectional graph convolution and performing root node feature enhancement, and finally performs average pooling of the propagation features and feature integration to train a softmax classifier to obtain the stance detection and rumor detection results.
[0005] The Chinese patent "A Multi-Task Joint Rumor Detection Method Incorporating Comments" (Application No.: 202110337896.2) uses a self-attention mechanism to respectively obtain rich context information of the Weibo text and user comments, then uses a gated unit with a filtering mechanism and an attention mechanism to effectively screen the user comments, and finally the output layer adopts a linear transformation and a softmax function to predict the relevant labels of the user comments and the event labels of the Weibo.
[0006] The existing multi-task rumor detection methods mainly have the following defects:
[0007] 1) Both of the two tasks that make up the MTL model are stance detection tasks and rumor detection tasks. However, since the two tasks use different data sets and the data structures used in the two tasks are quite different, although the model can share some information during the training process, the improvement brought about is not high;
[0008] 2) Existing multi-task rumor detection models all rely on other data sets or information helpful for rumor detection in other tasks to strengthen the feature representation of rumors, but they have not been able to fully exploit the information in the rumor detection data set itself.
[0009] 3) The multi-task rumor detection method based on the propagation structure requires a large amount of time to extract propagation structure features, that is, similar to the rumor detection method based on traditional machine learning algorithms, a large amount of time is required for feature engineering. Summary of the Invention
[0010] The object of the present invention is to propose a multi-task rumor detection model and method based on a parallel information sharing module that can both make full use of data information and improve the detection effect. This method combines the rumor detection method based on writing style features, the rumor detection method based on content, and the rumor detection method based on social background information with the content-based detection method as the center through a multi-task form to form a multi-task rumor detection model with a parallel information sharing structure; and uses the hierarchical attention mechanism (Hierarchical attention) for document classification in the conference paper "Hierarchical Attention Networks for Document Classification, in the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies. 2016. p. 1480-1489" by Yang Z et al. to pay attention to the words and tweets that are more important for rumor detection, which speeds up the convergence speed of the model, reduces the impact of noise, and overcomes the defects that the data results of the two tasks in the existing multi-task rumor detection methods are quite different and the data itself has not been fully utilized.
[0011] The technical solution adopted by the present invention is as follows:
[0012] A multi-task rumor detection model based on a parallel information sharing module, the model comprising:
[0013] A multi-task module based on text content and social background information, the module comprising:
[0014] Feature vector D0, which is obtained after encoding the user information of the source tweet poster through the Sentence Encoder module;
[0015] Feature vectors D1 - D n , which are feature vectors of the user information of related tweets;
[0016] Feature vector I0, which is the feature vector of the discrete social background information combination of the source tweet;
[0017] The social background feature vector of the i-th tweet of the detection object, which is composed of the concatenation of D i and I i ;
[0018] Feature vectors X0 - X n , which are obtained after encoding the text content of the source tweet and related tweets through the Sentence Encoder module;
[0019] A multi-task module based on text content and writing style, the module comprising:
[0020] Feature vector S0, which is the writing style feature vector of the source tweet;
[0021] Feature vectors S1 - S n , which are the writing style feature vectors of the related context;
[0022] Feature vector X i , which is the text content feature vector of the corresponding tweet.
[0023] When detecting based on text content,
[0024] S1: First, input the text content feature vectors X0 - X of the source tweet and related tweets n into the text content and social background information sharing module to obtain shared information containing content and social background information; at the same time, also input them into the text content and writing style sharing module to obtain shared information containing writing style information and text content;
[0025] S3: Combine the two sets of obtained shared information with the original in-text content information, and input them into the layer unique to the text content detection task, and finally obtain the final feature vector: H content ;
[0026] S3: Obtain the classification result through the fully connected layer.
[0027] The specific calculation formula is as follows:
[0028] X t = SentenceEncoder(W1,..., W n )
[0029] Where: X t represents the feature vector obtained after encoding the encoding matrix of the t-th by SentenceEncoder, and W i is the Word2Vec word vector representation of the i-th word of the t-th tweet, i = 1, 2,..., n,
[0030] h st1 = SharedModel1(X t )
[0031] h st2 = SharedModel2(X t )
[0032] Where: h st1 , h st2 are the intermediate states of the text content and writing style sharing module and the text content and social background information sharing module at the t-th moment respectively;
[0033]
[0034] Where: h t is the tweet vector representation at the t-th moment X t concatenated with the result after being calculated by the two sharing modules;
[0035] h ct = ContentBiLSTM(h t )
[0036] Where:
[0037] H content = SentenceAttention(h c0 ,..., h cn )
[0038] Where: h ci is the intermediate state of the unique layer of the text content rumor detection task at the i-th moment; i = 0, 1, 2,..., n;
[0039] Y c = Sigmoid(H content W c + b c)
[0040] Where: W c , b c are the trainable parameters of the fully connected layer. Y c represents the final output result of the rumor detection task based on the text content.
[0041] When detecting based on social background information,
[0042] Step1: Obtain the feature vector D of the user description through the Sentence Encoder i ;
[0043] Step2: Concatenate the feature vector D of the user description i with the corresponding discrete social background information feature I i to jointly form the social background information feature vector;
[0044] Step3: First, input the social background information feature vector into the text content and social background information sharing module to obtain the feature containing both text content information and social background information, and then combine it with the original social background information feature and input it into the layer unique to this task;
[0045] Step4: Obtain the social background feature vector H through Sentence Attention social ;
[0046] Step5: Obtain the final detection result Y through a Sigmoid layer so .
[0047] The specific calculation formula is as follows:
[0048] D t = SentenceEncoder(W1,..., W n )
[0049] Where: D t represents the feature vector obtained by the SentenceEncoder for the personal description text of the publisher corresponding to the t-th tweet.
[0050]
[0051] Where: I t is the corresponding other discrete information, and together with D t it forms the feature vector S of the social background information of the t-th tweet ot ;
[0052] h st2 = SharedModel2(St )
[0053]
[0054] h sot = SocialContextBiLSTM(h t )
[0055] H social = SentenceAttention(h so0 ,..., h son )
[0056] where: h soi is the intermediate state of the unique layer for the rumor detection task based on social context information at time i, i = 0, 1, 2,..., n;
[0057] Y so = Sigmoid(H social W c + b c )
[0058] When performing rumor detection based on writing style features,
[0059] A1: Input S0 - S n into the alignment layer Align Layer for preliminary feature extraction;
[0060] A2: Input the features extracted in A1 into the text content and writing style sharing module to obtain intermediate features that contain both text content information and writing style information;
[0061] A3: Combine the intermediate features obtained in A1 with the features passing through the alignment layer and input them into the writing style unique layer,
[0062] A4: Use the output of the last Bi - LSTM cell as the writing style feature H style , and obtain the final detection result Y st .
[0063] The specific calculation formula is as follows:
[0064] h at = AlignBiLSTM(S t )
[0065] where: h at is the intermediate state obtained by passing the writing style feature S t of the t - th tweet through the alignment layer;
[0066] h st= SharedBiLSTM(h at )
[0067] Where: h st is the corresponding intermediate state obtained by inputting h at into the text content and writing style sharing module;
[0068]
[0069] Where: h t is the intermediate feature vector obtained by concatenating h at and h st ;
[0070] H style = StyleBiLSTM(h n )
[0071] Where: h n is the intermediate vector corresponding to the last tweet;
[0072] Y st = Sigmoid(H style W s + b s )
[0073] Where: W s , b s represent the parameters to be trained of the Sigmoid layer function.
[0074] When detecting rumors based on text content,
[0075] B1: Encode the word vector matrix of each tweet into a feature vector X i through the Sentence Encoder;
[0076] B2: Input X0 - X n into the text content and writing style sharing module and the text content and social background information sharing module respectively to obtain intermediate features containing both types of information;
[0077] B3: Combine the obtained intermediate features with the text content features of itself and input them into the content layer;
[0078] B4: Obtain the text content feature vector Hcontent containing three types of information through a Sentence Attention;
[0079] B5: Classify through the fully connected layer.
[0080] The specific calculation formula is as follows:
[0081] u ct = tanh(h ct W c + b c )
[0082]
[0083]
[0084] where: u ct is the parameter obtained by the attention mechanism based on h ct ; α ct is the weight of the t-th tweet; u c , W c , b c are parameters to be trained;
[0085] H c is the weighted sum of the feature representations of each tweet in the detection object, and its weight represents the importance of each tweet for detection. By assigning different weights to different tweet features, tweets useful for rumor detection are focused on while ignoring the noise, thus accelerating the convergence speed of the model and improving the accuracy of the model.
[0086] A multi-task rumor detection method based on a parallel information sharing module includes the following steps:
[0087] Step 1: Define the social background information feature representation So, including the user self-description information D and the discrete social background information I;
[0088] In the said Step 1:
[0089] The user self-description information D is the feature vector obtained by encoding the word vector matrix obtained by segmenting and word embedding the user information description through a Sentence Encoder.
[0090] The discrete social background information I = {user_frc, user_folc, user_stac, user_ver, user_fac, t_attc, t_rec, time_gap}, where: user_frc, user_folc, user_stac, user_ver, user_fac respectively represent the number of friends, the number of fans, the number of tweets sent, whether authenticated, and the number of follows of the user; time_gap represents the time difference between the registration time of the user account and the publication time of the tweet; t_attc and t_rec respectively represent the number of likes and the number of retweets of the tweet.
[0091] Step 2: Train a Word2Vec word vector embedding model based on a large corpus, and then process each tweet into {Word1, Word2, …, Word n}, where: i = 1, 2, …, n, and Word i is the Word2Vec word vector corresponding to the i-th word of the tweet.
[0092] Step 3: Use the writing style representation S = {lenght, qu_sys, uh_sys, w1, w2, ..., w n , ot_sys}, where: length represents the length of the tweet; qu_sys represents the usage frequency of "?"; uh_sys represents the usage frequency of "!"; w1, w2, ..., w n represents the usage frequency of common parts of speech; ot_sys represents the usage frequency of other special symbols.
[0093] Then, obtain the style feature representation S i corresponding to each tweet.
[0094] Step 4: Establish a multi-task rumor detection model based on a parallel information sharing module and train the model.
[0095] In step 4, the model training includes the following steps:
[0096] Step 4.1: Input the vector matrix of text encoding, the user information description matrix, the user discrete information, and the corresponding writing style features into the model to obtain Y c , Y so , and Y st .
[0097] Step 4.2: Then calculate the corresponding loss functions Loss c , Loss so , and Loss st according to the corresponding true labels and the results obtained in the previous step, where the loss functions all use the cross-entropy loss function, and Loss c , Loss so , and Loss st are the loss functions corresponding to the tasks of detecting based on text content, detecting based on social background information, and detecting based on writing style, respectively.
[0098] Step 4.3: Calculate the overall loss Loss total of the model = α c *Loss c + α so *Loss so + α st *Lossst , where α c , α so and α st are the coefficients of the loss functions corresponding to the three tasks respectively.
[0099] Step 4.4: Calculate the gradient based on Loss total and update the parameters through backpropagation.
[0100] Step 4.5: If the number of iterations is less than Epoch, go back to Step 4.1. Otherwise, go to the next step.
[0101] Step 4.5: Validate the model parameters on the training set to obtain the corresponding validation results.
[0102] The multi-task rumor detection model and method based on a parallelized information sharing module according to the present invention have the following technical effects:
[0103] 1) The present invention organizes three basic rumor detection tasks based on text content, writing style, and social background in a multi-task manner centered on text content; secondly, on this basis, a multi-task rumor detection method based on a parallelized sharing module is formed through two information sharing modules. The advantages are: compared with the traditional multi-task learning method in which one sharing module connects multiple tasks, it can reduce the information interference caused when multiple types of information are projected into the same feature space.
[0104] 2) Since the present invention uses the same data set during the training process, it can more fully mine and utilize the information of the data itself for detection.
[0105] 3) The present invention uses a hierarchical attention mechanism to focus on the words and tweets that are more helpful for rumor detection, which can effectively accelerate the convergence of the model, reduce the interference of noise, and thus improve the detection effect of the model. Description of the Drawings
[0106] Figure 1 is the overall structure diagram of the model of the multi-task rumor detection method based on a parallelized information sharing module proposed by the present invention.
[0107] Figure 2 is the structure diagram of the rumor detection module based on text content and social background information.
[0108] Figure 3 is the structure diagram of the rumor detection module based on text content and writing style.
[0109] Figure 4 is the model training flow chart. Detailed Embodiments
[0110] 1) Define the social background information feature So, which consists of the user's self-description information D and some other discrete social background information I = {user_frc, user_folc, user_stac, user_ver, user_fac, t_attc, t_rec, time_gap}. The user's self-description information D is the feature vector obtained by encoding the word vector matrix obtained by segmenting and word embedding the user information description through the Sentence Encoder.
[0111] Among them: user_frc, user_folc, user_stac, user_ver, user_fac are the number of friends, the number of fans, the number of tweets sent, whether it is certified, and the number of follows of the user respectively; time_gap represents the time difference between the registration time of the user account and the publication time of this tweet, and t_attc, t_rec represent the number of likes and retweets of the tweet.
[0112] 2) Train a Word2Vec word vector embedding model based on a large corpus. Then, process each tweet into {Word1, Word2, …, Word n}, where: i = 1, 2, …, n, Word i is the Word2Vec word vector corresponding to the i-th word of the tweet.
[0113] The large corpus used is the Chinese and English corpus of Wikipedia.
[0114] The Word2Vec word vector embedding model is a language model proposed by Google in 2013 to convert text into vectors or matrices.
[0115] 3) Use the writing style representation S = {lenght, qu_sys, uh_sys, w1, w2,..., w n , ot_sys}, where: length represents the tweet length; qu_sys represents the usage frequency of "?"; uh_sys represents the usage frequency of "!"; w1, w2,..., w n represents the usage frequency of common part-of-speech; ot_sys represents the usage frequency of other special symbols. Then, obtain the style feature representation S i corresponding to each tweet.
[0116] 4) Design and define a multi-task rumor detection model based on a parallelized information sharing module and its calculation process. Its overall model structure is as Figure 1 shown. It mainly consists of the following parts:
[0117] (I): A multi-task module based on text content and social background information:
[0118] This part includes a rumor detection task based on text content and a rumor detection task based on social background information. These two parts achieve information sharing through an attention-based sharing method, and the detection effects of each can be effectively enhanced during the information sharing process. Because in the shared space, the abnormal social background information features during the rumor spreading process will be projected near the text content features of false information. After passing through the shared layer respectively, they return to the layers unique to each task with the information in the shared space, thereby enhancing the feature representations of each task. The overall structure of this part is as Figure 2 shown.
[0119] Among them: D0 is the feature vector obtained after encoding the user information description of the source tweet poster through the Sentence Encoder module; D1 - D n are the feature vectors of the user information of the related tweets; I0 is the feature vector of the combined discrete social background information of the source tweet; D i and I i After splicing, they form the social background feature vector representation of the i-th tweet of the detection object;
[0120] X0 - X n are the feature vectors obtained after encoding the text content of the original tweet and related tweets through the Sentence Encoder module.
[0121] Among them: The related tweets are the tweets that retweet or comment on the original tweet;
[0122] The Sentence Encoder module is an encoder that encodes the text features originally represented as a word vector matrix into a fixed-dimensional vector.
[0123] For the detection task based on text content, first input the content feature vectors X0 - X n of the source tweet and related tweets into the text content and social background information sharing module to obtain the shared information containing content and social background information; at the same time, input them into the text content and writing style sharing module to obtain the shared information containing writing style information and text content. Then, combine the two sets of obtained shared information and the original text content information and input them into the layer unique to the text content detection task. Finally, through Sentence Attention, the final feature vector representation H content is obtained; finally, through the fully connected layer, the classification result is obtained.
[0124] The original text is the feature vector obtained after encoding the original tweet through the Sentence Encoder.
[0125] The text content and social background information sharing module is used to extract information shared based on text content and social background information detection tasks;
[0126] The text content and writing style sharing module is used to extract information shared based on text content and writing style detection tasks;
[0127] The layer unique to the text content detection task is used to further extract the layer based on text content features.
[0128] The fully connected layer converts the obtained vector representation into a vector containing only one number for the Sigmoid function to calculate.
[0129] The specific calculation formula is as follows:
[0130] X t = SentenceEncoder(W1,..., W n )
[0131] h st1 = SharedModel1(X t )
[0132] h st2 = SharedModel2(X t )
[0133]
[0134] h ct = ContentBiLSTM(h t )
[0135] H content = SentenceAttention(h c0 ,..., h cn )
[0136] Y c = Sigmoid(H content W c + b c )
[0137] Where: W i is the Word2Vec word vector representation of the i-th word of the t-th tweet, W i is the Word2Vec word vector representation of the i-th word of the t-th tweet, i = 1, 2,..., n,; h st1 and h st2 are the intermediate states of the text content and writing style sharing module and the text content and social background information sharing module at the t-th moment respectively; ht is the tweet vector representation X at time t t which is concatenated with the result after being calculated by two shared modules to obtain h ci is the intermediate state of the layer unique to the rumor detection task based on text content at time i; i = 0, 1, 2, …, n; W c and b c are the trainable parameters of the fully connected layer.
[0138] For the detection task based on social background information, first, the feature vector D of the user description is obtained through the Sentence Encoder i ; then, it is concatenated with the corresponding feature I of other discrete social background information i to jointly form the social background information feature vector; the social background information feature vector is first input into the shared module of text content and social background information to obtain the feature containing both text content information and social background information; then, it is combined with the original social background information feature and input into the layer unique to this task, and the layer unique to this task is used to further extract the feature based on social background information. Finally, the social background feature vector representation H social is obtained through Sentence Attention, and finally, the final detection result Y so is obtained through a Sigmoid layer, and its formula is as follows:
[0139] D t = SentenceEncoder(W1,..., W n )
[0140]
[0141] h st2 = SharedModel2(S t )
[0142]
[0143] h sot = SocialContextBiLSTM(h t )
[0144] H social = SentenceAttention(h so0 ,..., h son )
[0145] Y so = Sigmoid(H social W c + b c )
[0146] Where: W i is the Word2Vec word vector of the i-th word in the user description of the t-th user; I t is the corresponding other discrete information, which together with D t forms the feature vector So of the social background information of the t-th tweet t ; h st1 is the intermediate state of the text content and social background information sharing module at the t-th moment; h t is the tweet vector So at the t-th moment t obtained by concatenating it with the result calculated by the text content and social background information sharing module; h soi is the intermediate state of the layer unique to the rumor detection task based on social background information at the i-th moment, i = 0, 1, 2,..., n; W c and b c are the trainable parameters of the fully connected layer.
[0147] (2): Multi-task module based on text content and writing style:
[0148] The information sharing method based on text content and writing style realizes information interaction and sharing in the form of sharing parameters through the rumor detection task based on text content and the rumor detection task based on writing style. At the same time, the information sharing method based on the attention mechanism is also used to reduce noise interference, so as to effectively improve the detection effect of each. Because in the shared feature space, the projection of the text content features of false information often tends to be very close to the corresponding writing style features. After their respective calculations through the shared layer, they will carry rich shared information and return to the unique layer of their own tasks, thereby improving the quality of their respective features. Its overall structure is as Figure 3 shown.
[0149] Where: S0 is the representation of the writing style feature vector of the source tweet, and S1 - S n are the writing style features of the relevant context; while X i is the text content feature vector of the corresponding tweet;
[0150] For the rumor detection task based on writing style features, first input S0 - S n into the alignment layer for preliminary feature extraction, then input it into the text content and writing style sharing module to obtain intermediate features containing both text content information and writing style information. Then, combine the intermediate features containing text content information and writing style information with the features passed through the alignment layer and input them into the layer unique to writing style; finally, take the output of the last Bi-LSTM unit as the writing style feature H style, and obtain the final detection result Y through a fully connected layer st ,
[0151] The alignment layer performs preliminary feature extraction on the writing style feature vector and makes the dimension of S i consistent with that of X i ;
[0152] The writing style unique layer is used to further extract writing style features.
[0153] Its calculation formula is as follows:
[0154] h at = AlignBiLSTM(S t )
[0155] h st = SharedBiLSTM(h at )
[0156]
[0157] H styl e = StyleBiLSTM(h n )
[0158] Y st = Sigmoid(H style W s + b s )
[0159] Among them, h at is the intermediate state obtained by the writing style feature S t of the t-th tweet passing through the alignment layer; h st is the corresponding intermediate state obtained by inputting h at into the text content and writing style sharing module; h t is the intermediate feature vector obtained by concatenating h at and h st , h n is the intermediate vector corresponding to the last tweet, W c and b c are the trainable parameters of the fully connected layer.
[0160] For the rumor detection task based on text content, first encode the word vector matrix of each tweet into a feature vector X through the Sentence Encoder i , and then X0 - X nThey are respectively input into the text content and writing style sharing module and the text content and social background information sharing module to obtain intermediate features containing both types of information; then the intermediate features containing both types of information are combined with the text content features of itself and input into the content layer; finally, a text content feature vector H containing three types of information is obtained through a SentenceAttention content ; The calculation formula of Sentence Attention is as follows. Finally, classification is performed through a fully connected layer
[0161] The text content of itself is the text feature vector X obtained after passing through the Sentence Encoder i 。
[0162] The Sentence Attention is used to generate the attention mechanism for the weight occupied by each X i 。
[0163] u ct =tanh(h ct W c +b c )
[0164]
[0165]
[0166] Among them: u ct is the parameter obtained by the attention mechanism according to h ct , α ct is the weight occupied by the t-th tweet, u c , W c , b c are parameters to be trained. The finally obtained H c is the weighted sum of the feature representations of each tweet in the detection object, and its weight represents the importance of each tweet for detection. By assigning different weights to different tweet features, tweets useful for rumor detection are focused on, while those noises are ignored, thereby achieving the effects of accelerating the convergence speed of the model and improving the accuracy of the model
[0167] 5) Model training: The training process of the model proposed by the present invention is as Figure 4 shown. Among them: Y tc is the true label of the content-based rumor detection task, Y ts is the true label of the writing style-based rumor detection task, Y tso is the true label of the rumor detection task based on social background information, Y c , Y s and Y soThey respectively correspond to their predicted labels, where epoch is the number of iterations and Adam is the parameter optimization method.
[0168] Examples:
[0169] To verify the effectiveness of the method of the present invention, experiments were conducted and verified on datasets collected from Twitter, Pheme, and Chinese Weibo. The present invention uses four evaluation metrics, namely accuracy, precision, recall, and F1-score, to reflect the effectiveness of the method of the present invention. Their corresponding calculation formulas are as follows:
[0170]
[0171]
[0172]
[0173]
[0174] Among them: TP represents the true positive rate, TN represents the true negative rate, FP represents the false positive rate, and FN represents the false negative rate. Accuracy can reflect the method's precision, while precision and recall reflect the detection effect of the method of the present invention on positive samples, and F1 reflects the comprehensive performance of the model. The experimental results of the present invention are shown in Table 1:
[0175] Table 1 Experimental Results on Pheme Dataset and Chinese Weibo Dataset
[0176]
[0177] As can be seen from Table 1 above, good detection results have been achieved on both the English Pheme dataset and the Chinese Weibo dataset.
Claims
1. A multi-task rumor detection model based on a parallel information sharing module, characterized in that The model includes: A multi-task module based on text content and social background information, which includes: Feature vector D0, which is obtained after encoding the user information of the source tweet poster by the Sentence Encoder module; Feature vectors D1 - D n , which are feature vectors of user information of related tweets; Feature vector I0, which is the feature vector of the discrete social background information combination of the source tweet; Feature vectors X0-X n , which are obtained after the text content of the source tweet and related tweets is encoded by the Sentence Encoder module; A multi-task module based on text content and writing style, which includes: Feature vector S0, which is the writing style feature vector of the source tweet; Feature vectors S1 - S n , which is the writing style feature vector of the context associated therewith; Feature vector X i , which is the feature vector of the text content of the corresponding tweet; When detecting rumors based on text content, S1: First, input the source tweet and the text content feature vectors X0-X of the related tweets n into the text content and social background information sharing module to obtain the shared information containing content and social background information; at the same time, input them into the text content and writing style sharing module to obtain the shared information containing writing style information and text content; S3: Combine the two sets of shared information obtained and the original in-text content information, and input them into the layer unique to the rumor detection task based on text content. Finally, obtain the final feature vector: H through Sentence Attention content ; S3: Obtain the classification result through the fully connected layer; The specific calculation formula is as follows: X t = SentenceEncoder(W1,…,W n ) where: X t represents the feature vector obtained after encoding the coding matrix of the t-th by the SentenceEncoder, and W i is the Word2Vec word vector representation of the i-th word of the t-th tweet, where i = 1, 2, …, n h st1 = SharedModel1(X t ) h st2 = SharedModel2(X t ) where: h st1 and h st2 are the intermediate states of the text content and writing style sharing module and the text content and social background information sharing module at the t-th moment, respectively; where: h t is the tweet vector representation X at time t t concatenated with the result after being calculated by two shared modules; h ct = ContentBiLSTM(h t ) Where: H content = SentenceAttention(h c0 , …, h cn ) Where: h ci is the intermediate state of the unique layer for the rumor detection task based on text content at time i; i = 0, 1, 2, …, n; Y c = Sigmoid(H content W c + b c ) Where: W c , b c are the trainable parameters of the fully connected layer; Y c represents the final output result of the rumor detection task based on the text content.
2. The multi-task rumor detection model based on the parallel information sharing module according to claim 1, characterized in that: When detecting based on social background information, Step1: Obtain the feature vector D of the user description through the Sentence Encoder i ; Step2: Concatenate the feature vector D described by the user i with the corresponding discrete social background information feature I i to jointly form a social background information feature vector; Step3: First, input the social background information feature vector into the text content and social background information sharing module to obtain features containing both text content information and social background information, and then combine them with the original social background information features and input them into the layer unique to this task; Step4: Obtain the social background feature vector H through Sentence Attention social ; Step5: Obtain the final detection result Y through a Sigmoid layer so 。 3. The multi-task rumor detection model based on the parallel information sharing module according to claim 2, characterized in that: The specific calculation formula is as follows: D t = SentenceEncoder(W1,…,W n ) Among them: D t represents the feature vector obtained by SentenceEncoder for the personal description text of the publisher corresponding to the t-th tweet; where: I t is the other discrete information corresponding to the t-th tweet, which together with D t forms the feature vector So of the social background information of the t-th tweet t ; h st2 = SharedModel2(S t ) h sot = SocialContextBiLSTM(h t ) H social = SentenceAttention(h so0 , …, h son ) Where: h soi is the intermediate state of the unique layer of the rumor detection task based on social background information at time i, where i = 0, 1, 2,..., n; Y so = Sigmoid(H social W c + b c )。 4. The multi-task rumor detection model based on the parallel information sharing module according to claim 1, characterized in that: When detecting rumors based on writing style features, A1: Input S0 - S n into the alignment layer Align Layer for preliminary feature extraction; A2: Input the features extracted in A1 into the text content and writing style sharing module to obtain intermediate features containing both text content information and writing style information; A3: Combine the intermediate features obtained in A1 with the features passed through the alignment layer and input them into the layer unique to the writing style, A4: Use the output of the last Bi-LSTM cell as the writing style feature H style and obtain the final detection result Y through a fully connected layer st ; The specific calculation formula is as follows: h at = AlignBiLSTM(S t ) where: h at is the writing style feature S of the t-th tweet t the intermediate state obtained through the alignment layer; h st = SharedBiLSTM(h at ) where: h st is the corresponding intermediate state obtained by inputting h at into the text content and writing style sharing module; where: h t is an intermediate feature vector obtained by concatenating h at and h st ; H style = StyleBiLSTM(h n ) Where: h n is the intermediate vector corresponding to the last tweet; Y st = Sigmoid(H style W s + b s ) Where: W s , b s represent the parameters to be trained of the Sigmoid layer function.
5. The multi-task rumor detection model based on the parallel information sharing module according to claim 1, characterized in that: When detecting rumors based on text content, B1: Encode the word vector matrix of each tweet into a feature vector X through the Sentence Encoder i ; B2: Input X0-X n into the text content and writing style sharing module and the text content and social background information sharing module respectively to obtain intermediate features containing both types of information; B3: Combine the obtained intermediate features with the text content features of itself and input them into the content layer; B4: Through a Sentence Attention, obtain the text content feature vector Hcontent containing all three types of information; B5: Classify through the fully connected layer; The specific calculation formula is as follows: u ct = tanh(h ct W c + b c ) Where: u ct is the parameter obtained by the attention mechanism based on h ct ; α ct is the weight of the t-th tweet; u c , W c , b c are the parameters to be trained; H c It is the weighted sum of the feature representations of each tweet in the detection object, and its weight represents the importance of each tweet for detection.
6. A multi-task rumor detection method based on a parallel information sharing module, characterized in that It includes the following steps: Step 1: Define the social background information feature representation So, including the user self-description information D and the discrete social background information I; Step 2: Train a Word2Vec word vector embedding model based on a corpus, and then process each tweet into {Word1, Word2, …, Word n}, where: i = 1, 2, …, n, Word i is the Word2Vec word vector corresponding to the i-th word of the tweet; Step 3: Represent S = {lenght, qu_sys, uh_sys, w1, w2,..., w n , ot_sys} using the writing style, where: length represents the tweet length; qu_sys represents the usage frequency of "?"; uh_sys represents the usage frequency of "!"; w1, w2,..., w n represents the usage frequency of common parts of speech; ot_sys represents the usage frequency of other special symbols. Then, obtain the style feature representation S corresponding to each tweet i ; Step 4: Establish a multi-task rumor detection model based on the parallel information sharing module according to any one of claims 1 to 5, and train the model.
7. The multi-task rumor detection method based on the parallel information sharing module according to claim 6, characterized in that: In the said step 1: The user self-description information D is the feature vector obtained after encoding the word vector matrix obtained by segmenting the user information description and performing word embedding through the Sentence Encoder; The discrete social background information I = {user_frc, user_folc, user_stac, user_ver, user_fac, t_attc, t_rec, time_gap}, where: user_frc, user_folc, user_stac, user_ver, user_fac respectively represent the number of friends, the number of fans, the number of tweets sent, whether it is authenticated, and the number of follows of the user; time_gap represents the time difference between the registration time of the user account and the publication time of the tweet; t_attc and t_rec respectively represent the number of likes and the number of retweets of the tweet.
8. The multi-task rumor detection method based on the parallelized information sharing module according to claim 6, characterized in that: In the step 4, the model training includes the following steps: Step 4.1: Input the vector matrix of text encoding, the user information description matrix, the user's discrete information, and the corresponding writing style features into the model to obtain Y c , Y so , and Y st ; Step 4.2: Then calculate the corresponding loss function Loss according to the corresponding true label and the result obtained in the previous step c , Loss so and Loss st , where the loss functions all use the cross-entropy loss function, Loss c , Loss so and Loss st are the loss functions corresponding to the text content detection, social background information detection, and writing style detection tasks respectively; Step 4.3: Calculate the overall loss Loss of the model total = α c * Loss c + α so * Loss so + α st * Loss st , where α c , α so , and α st are the coefficients of the loss functions corresponding to the three tasks respectively; Step 4.4: According to Loss total Calculate the gradient and backpropagate to update the parameters; Step 4.5: If the number of iterations is less than Epoch, go back to step 4.1; otherwise go to the next step; Step 4.5: Obtain the model parameters, verify them on the training set, and obtain the corresponding verification results.
Citation Information
Patent Citations
A multi-task rumor detection method based on bidirectional propagation graph
CN113094596B
Comment-fused multi-task joint rumor detection method
CN113158075A
Multitask rumor detection method based on bidirectional propagation graph
CN113094596A
Writing style-based multi-task rumor detection method, device and equipment
CN114491025A