Rumor detection method and device based on text and user features

By combining graph convolution network, gated loop unit and reinforcement learning algorithm, the problems of insufficient feature extraction and difficult timing judgment in rumor detection are solved, efficient early rumor detection is achieved, and detection accuracy and timeliness are improved.

CN120354829AInactive Publication Date: 2025-07-22ZHEJIANG HUAXUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310903076.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-07-21
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing rumor detection methods ignore the timing characteristics between comments when extracting rumor characteristics, and the optimization strategies of early rumor detection decisions are difficult to converge, making it difficult to accurately judge the timing of rumor detection.

Method used

Combining the graph convolution network and gated loop unit, the rumor propagation and timing characteristics are extracted, and the early rumor detection timing is optimized through reinforcement learning algorithms. Text and user features are extracted using the RoBERTa model, and rumor propagation graph is constructed, and feature fusion and detection are combined with attention mechanism and LSTM network.

Benefits of technology

It improves the accuracy of early rumor detection and can detect more than 80% of the incidents within 6 hours of rumors occurring, effectively preventing the spread of rumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354829A_ABST
    Figure CN120354829A_ABST
Patent Text Reader

Abstract

The invention discloses a rumor detection method and device based on text and user features, early rumor detection can be carried out on posts in social media, and the method comprises the steps that firstly, text features of rumors and user features are fused to form feature representations of rumor source posts and comment posts; secondly, deep features of rumor propagation are mined by using a graph convolutional network and a gating loop unit; and finally, through a reinforcement learning algorithm, guiding the model to carry out early rumor detection at a proper opportunity. According to the method, two neural network models are used for extracting multidimensional features of rumor propagation, and a reinforcement learning algorithm is used for enabling the models to automatically learn the detection opportunity of early rumor detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for rumor detection based on text and user characteristics, belonging to the technical fields of the Internet and artificial intelligence. Background Art

[0002] The rapid development of social media has facilitated the communication between people, but at the same time, it has also bred a lot of rumors. Due to the characteristics of fast information dissemination speed and wide influence range of social media, rumors have run wild on social media, seriously affecting people's access to normal information. And manual rumor refutation has high costs and low efficiency, and cannot judge rumors in time. Therefore, automatic rumor detection on social media has become a task worthy of exploration.

[0003] Early methods for rumor detection mainly relied on automatically detecting features extracted manually. This method first extracts features that are conducive to representing rumor information from rumor datasets, including post text features, post propagation features, and user features, etc. Then these manually extracted features are put into traditional machine learning models, such as support vector machines, decision trees, random forest models, etc. for training, and finally the results of rumor classification are output. Such methods rely heavily on manually extracted features, and the production of these features is a labor-intensive project, which requires a lot of time and manpower. At the same time, these manually made features are highly subjective and can only obtain shallow representations of rumors, and cannot learn deep features of rumors. Later, many researchers applied deep learning models to rumor detection. Based on LSTM, CNN, and RNN models, they proposed many methods for rumor detection, and input the input data stream into the deep learning model according to time slices, improving the accuracy of the model. However, these models ignore the structural relationship of comments and cannot effectively capture the characteristics of wide spread of rumors. In recent years, the emergence of graph neural networks has provided a new solution for rumor detection. Huang et al. proposed a rumor detection model based on graph convolutional neural networks. This model comprehensively considers the content, user, and propagation aspects of rumor detection, and consists of three modules, namely a user feature encoder, a propagation tree encoder, and a connector that integrates the outputs of the two modules. Tian et al. proposed a bidirectional graph convolutional network structure. In this model, the upward propagation and downward propagation modes of social media text are combined, effectively capturing the global features of the rumor structure.

[0004] However, although existing rumor detection methods have achieved some development, there are still deficiencies. The method based on graph convolutional network is insufficient in extracting the temporal features of rumors. When the graph convolutional network model extracts rumor features, it can only extract local features of rumors and ignores the temporal features between comments. The existing optimization strategies for early rumor detection decisions mainly use the Actor-Critic algorithm. This algorithm has no limiting conditions when updating the strategy, and the strategy update amplitude is large, which easily leads to the problem that it is difficult to converge in the training of the optimization strategy, resulting in difficulty in grasping the timing of early rumor detection. Summary of the Invention

[0005] Aiming at the problems and deficiencies of the existing technology, the present invention proposes a rumor detection method based on text and user features, which can perform early rumor detection on posts in social media. This method first fuses the text features and user features of rumors to form feature representations of rumor source posts and comment posts; secondly, the present invention uses a graph convolutional network and a gated recurrent unit to mine the deep features of rumor propagation; finally, through a reinforcement learning algorithm, the model is guided to perform early rumor detection at an appropriate time.

[0006] To achieve the above object, the technical solution of the present invention is as follows: A rumor detection method based on text and user features, which mainly includes processes such as extracting text features and user features, using a graph convolutional network to mine rumor propagation features, using a gated recurrent unit to mine rumor temporal features, and using a reinforcement learning algorithm to find an appropriate timing for early rumor detection and perform rumor detection, etc., which can mine the propagation features and temporal features of rumors and enable the early rumor detection model to find an appropriate detection timing. This method mainly includes four steps, specifically as follows:

[0007] Step 1, use the RoBERTa model to extract text features and user features and fuse the two;

[0008] Step 2, use a graph convolutional network to mine the propagation features of rumors. First, construct a rumor propagation graph according to the comment relationship, and then use the graph convolutional network to mine rumor propagation features;

[0009] Step 3, use a gated recurrent unit to mine the temporal features of rumors. First, divide the comment structure according to the time of comments, and then use the gated recurrent unit to extract the temporal features of rumors and weight the features at each moment in combination with the attention mechanism;

[0010] Step 4, first fuse the features extracted by the graph convolutional network and the features extracted by the gated recurrent unit to form global features, then use a reinforcement learning algorithm to find an appropriate time point for early rumor detection, and finally detect early rumors.

[0011] The present invention also provides a rumor detection device based on text and user characteristics, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the above-mentioned rumor detection method based on text and user characteristics is implemented.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] 1. This method combines a graph convolutional network and a gated recurrent unit, which can not only mine the temporal and semantic information between comments but also extract the deep features of rumor propagation, solving the problem of insufficient feature extraction of a single neural network and improving the accuracy of early rumor detection.

[0014] 2. This method divides rumors into sub-comment structures unfolded in time series according to the time points of comments, constructs a rumor propagation graph, and then uses the proximal policy optimization (PPO) algorithm of reinforcement learning as an optimization strategy for early rumor detection decisions. A KL divergence constraint is added to limit the update amplitude of the strategy in the optimization strategy, making the optimization strategy easy to converge and easier to find a suitable detection point for early rumor detection.

[0015] 3. The early rumor detection method provided by this method enables more than 80% of events to be detected within 6 hours after the rumor occurs, which is conducive to timely preventing the spread of rumors. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a general framework diagram of the method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In order to deepen the understanding and recognition of the present invention, the present invention will be further clarified below with reference to specific embodiments.

[0018] Embodiment 1: A rumor detection method based on text and user characteristics, the overall framework of which is as Figure 1 shown, and the method includes the following steps:

[0019] Step 1, use the RoBERTa model to extract text features and user features and fuse the two. Specifically, this implementation step is divided into 2 specific sub-steps:

[0020] Sub-step 1-1, for the source post, comment text, and corresponding user characteristics of the rumor dataset, use the RoBERTa model to extract features. RoBERTa will first add <s>Identifier, add < / s> identifiers at the beginning of the sentence. Then, the tokens in the sentence are added to RoBERTa one by one for training to obtain the encoded feature representation:

[0021] R = RoBERTa( <s>,Sentence1, <sep>at the end of the sentence ,Sentence2,…,< / sep> < / s>) Select " <s>"Vector R at the position" s as the extracted feature.

[0022] Sub-step 1-2, fusion of text features and user features. After the word embedding process, the source post text content and the early comment text are subjected to feature extraction to obtain the text feature vector R text , and the user features are subjected to feature extraction to obtain the user feature vector R user . Among them, the user feature vector is obtained by concatenating the user nickname feature, the user profile feature, and the Weibo authentication type feature, and is specifically represented as follows:

[0023] R user = R nickname ||R profile ||R certification

[0024] where R nickname represents the feature vector encoded by the RoBERTa model for the user nickname, R profile represents the feature vector encoded by the RoBERTa model for the user profile, R certification represents the feature vector encoded by the RoBERTa model for the Weibo authentication type, and || represents the concatenation operation. After concatenating the text features and the user features, the fused feature is obtained by inputting them into the fully connected layer:

[0025] h = W * (R text ||R user ) + b

[0026] where R text represents the text feature, R user represents the user feature, h represents the fused feature, and W and b are the weights and bias terms of the linear transformation respectively.

[0027] Step 2, use the graph convolutional network to mine the propagation features of rumors. First, construct a rumor propagation graph according to the comment relationship, and then use the graph convolutional network to mine the rumor propagation features. Specifically, the implementation process of this step is divided into 2 sub-steps:

[0028] Sub-step 2-1, construction of the dynamic propagation graph. In this paper, according to the time when the rumor comments are released, a dynamic propagation graph is constructed for the rumor dataset. Specifically, first, the structure of the rumor comments is divided according to the release time, starting from the rumor sub-structure G1 with only the source post content. Each time a comment information is added in chronological order, the new comment information is added to the original sub-comment to form the next rumor sub-structure until all the comment information is added to form the final rumor structure G n-1 , where n - 1 represents the number of comments in this event. Finally, the dynamic propagation structure of event C j can be represented as For each rumor sub-structure can be represented as where t - 1 is the number of comments in this rumor sub - structure.

[0029] According to the rumor sub - structure formed above, construct its propagation graph. The rumor sub - structure is constructed into a graph structure form, specifically represented as represents the set of nodes in the rumor sub - structure, represents the comment action, denoted as If node N u comments on node N v , then is 1, otherwise is 0. During the construction of the propagation graph, we construct the vertices and edges of the propagation graph according to the nodes and the comment relationships between nodes. The feature of each vertex is represented as a fused feature obtained by fusing the text feature and the user feature.

[0030] Sub - step 2 - 2: Use the graph convolutional network to extract node features. After obtaining the fused feature of each node and constructing the dynamic propagation graph, we choose to use the graph convolutional neural network to extract the features of the rumor sub - propagation structure to obtain the propagation features of the rumor sub - propagation structure. In order to be able to extract the features of the rumor sub - propagation structure presented in the graph structure, this chapter uses two - layer graph convolutional network layers to perform convolutional operations on the rumor nodes. The input of the graph convolutional network GCN allows for a graph structure of any shape, and generates a node - level embedding representation through convolutional operations. In this paper, the input of the GCN layer is the rumor sub - structure generated above which contains t nodes and the edges between nodes. The feature vector corresponding to each node N i is the fused feature after fusing the text feature and the user feature Graph The corresponding adjacency matrix is represented as Let D be a diagonal matrix and satisfy D ii = ∑ j A ij . The input of the GCN layer is the fused feature matrix H (0) of all nodes in the current rumor sub - propagation structure. Each layer obtains the hidden representation of the next layer through convolutional operations. The convolutional process is as follows:

[0031]

[0032] where H (l) represents the hidden representation matrix of the l - th layer; I is the identity matrix; is 's degree matrix, which is a diagonal matrix and satisfies W (l) is the trainable weight matrix of the l-th layer, and σ represents the activation function. In this paper, the sigmoid function is used.

[0033] After encoding through two graph convolutional network layers, average pooling is performed on the hidden features of all nodes in the last layer to obtain the features extracted by the graph convolutional network module. The calculation method is as follows:

[0034]

[0035] where represents the vector of the i-th row in the hidden feature matrix extracted by the second graph convolutional network layer, represents the rumor sub-propagation structure is the feature extracted by the graph convolutional network module.

[0036] Step 3: Use a gated recurrent unit to mine the temporal features of the rumor. First, divide the comment structure according to the time of the comments, and then use the gated recurrent unit to extract the temporal features of the rumor and combine the attention mechanism to weight the features at each moment. Specifically, the implementation process of this step is divided into 2 sub-steps:

[0037] Sub-step 3-1: Divide the comment structure of the rumor. First, arrange the source post and all comment contents in a time series. Event C j can be expressed as where represents the fused feature obtained by fusing the source post text feature and the user feature, represents the fused feature obtained by fusing the i-1-th comment text feature and the user feature.

[0038] Sub-step 3-2: Use a gated recurrent unit to extract the temporal features of rumor propagation. For a certain moment t, the input of the GRU layer is the fused feature and the hidden state h t-1 passed down from the previous node, and the output is the hidden state h t of the next node. In the calculation of the GRU layer, there are two important gates, one is the reset gate and the other is the update gate. The role of the reset gate is to determine how to combine the input at the current moment with the previous memory, and the role of the update gate is to determine how much of the previous memory plays a role. The specific calculation methods of the reset gate and the update gate are as follows:

[0039] z r = σ(W r * [h t-1 , x t + b r )

[0040] z u = σ(W u * [h t-1 , x t + b u )

[0041] where z r represents the reset gate, and z u represents the update gate. W r and W u are the trainable weight matrices of the reset gate and the update gate respectively, and b r and b u are the bias terms of the reset gate and the update gate respectively. σ is the activation function, which is the sigmoid function here and can convert the input into data in the range of [0, 1] to act as the gating signal. After obtaining the gating signal, first use the reset gate to combine the hidden state passed down from the previous node with the current input to obtain the candidate hidden state. The calculation formula is:

[0042] h ′ = tanh(W * [h t-1 ⊙ z r , x t + b)

[0043] where ⊙ represents the Hadamard product, that is, element-wise multiplication; [h t-1 ⊙ z r , x t represents calculating the Hadamard product of the hidden state at the previous moment and the reset gate and then concatenating it with the current input; tanh is the activation function that scales the result to the range of [-1, 1]. Then use the update gate to update the memory to obtain the hidden state of the next node. The calculation method is:

[0044] h t = (1 - z) ⊙ h t-1 + z ⊙ h ′

[0045] where the range of z is [0, 1]. When z is equal to 0, it means directly using the hidden state at the previous moment without updating the hidden state; when z is equal to 1, it means using the currently updated hidden state; the value of z represents the proportion of the currently updated hidden state in the final hidden state. The larger the value of z, the more the past hidden state is forgotten. h t is the output of the hidden state at the current moment and is also the input of the hidden state passed to the next moment.

[0046] For the hidden state extracted by the gated recurrent unit GRU, use the attention mechanism to calculate the probability weights for the fusion features of the inputs at different times, so that the fusion features of some comments can receive more attention, thereby improving the quality of feature extraction by the gated recurrent unit layer.

[0047] The input of this attention mechanism layer is the hidden state outputs {h 1 , h 2 , …, h t} at each moment of the feature extraction layer. For a certain moment t, the similarity between the input and the output is calculated, and the dot product is used to calculate the similarity. The calculation formula is:

[0048] e i = tanh(w i h i + b i )

[0049] e i represents the similarity between the hidden state h i at the i-th moment and the output vector. The similarities at all moments with the output vector are calculated, and then a normalization operation is performed on them. The calculation formula is:

[0050]

[0051] where α i represents the weight coefficient of the hidden state h i at the i-th moment. The softmax function is used for the normalization operation. After obtaining the weight coefficients of the hidden states at each moment, the weighted sum of the hidden state outputs at each moment is calculated according to the weight coefficients to obtain the features finally extracted by the gated recurrent unit module. The calculation formula is:

[0052]

[0053] Step 4: First, fuse the features extracted by the graph convolutional network and the features extracted by the gated recurrent unit to form global features, then use the reinforcement learning algorithm to find the appropriate early rumor detection time point, and finally detect the early rumor. Specifically, the implementation process of this step is divided into 4 sub-steps:

[0054] Sub-step 4-1: Fuse the features extracted by the GCN network and the GRU network. The calculation method is:

[0055]

[0056] For the rumor sub-structure at the t-th moment, the feature extracted by the graph convolutional network module is The feature extracted by the gated recurrent neural network module is W GCN and W GRU are the parameter matrices corresponding to the graph convolutional network module and the gated recurrent neural network module respectively, and b t is the bias term at the t-th moment. Finally, the global feature extracted by GCN-GRU at the t-th moment is combined.

[0057] Sub-step 4-2, perform LSTM network encoding on the global features. At a certain moment t, the inputs of the LSTM network are three, namely the memory cell c passed from the previous moment t-1 , the hidden state h passed from the previous moment t-1 , and the input H at the current moment t . The outputs are two, namely the cell state c at this moment t and the hidden state h at this moment t . There are three gates in the LSTM network, namely the forget gate, the input gate, and the output gate. The calculation methods of the three gates are as follows:

[0058] z i = σ(W xi H t + W hi h t-1 + b i )

[0059] z f = σ(W xf H t + W hf h t-1 + b f

[0060] z o = σ(W xo H t + W ho h t-1 + b o )

[0061] Among them, z i , z f , and z o represent the input gate, the forget gate, and the output gate respectively; W xi and W hi are the linear transformation matrices of the input at the current moment and the hidden state at the previous moment in the input gate, and b i is the bias term of the input gate; W xf and W hf are the linear transformation matrices of the input at the current moment and the hidden state at the previous moment in the forget gate, and b f is the bias term of the forget gate; W xo and W ho are the linear transformation matrices of the input at the current moment and the hidden state at the previous moment in the output gate, and b o is the bias term of the output gate, and σ is the activation function, where the sigmoid function is used. Here, the role of the input gate is to determine whether to ignore the input data, the role of the forget gate is to reduce the value towards 0, and the role of the output gate is to determine whether to use the hidden state. In addition to these three gated states, there is also a candidate memory cell in the LSTM network for processing the input data, and the calculation method is:

[0062]

[0063] where is the candidate memory cell, W xc and W hc are the linear transformation matrices of the input at the current time and the hidden state at the previous time in the candidate memory cell, b c is the bias term of the candidate memory cell, and tanh is the activation function that converts the result to the range [-1, 1]. After obtaining the three gated signals and the candidate memory cell, the memory cell can be calculated through the following calculation method:

[0064]

[0065] c t represents the memory cell. First, perform the Hadamard product operation on the forget gate signal and the memory cell at the previous time, then calculate the Hadamard product of the input gate signal and the candidate memory cell, and then combine the two to obtain the memory cell. Since both the forget gate signal and the input gate signal are in the range [0, 1] and are independent of each other, the memory cell at this time in the LSTM can take into account both the memory cell at the previous time and the input at the current time, or ignore both the memory cell at the previous time and the input at the current time, which is more flexible than the GRU. After obtaining the memory cell, the hidden state can be calculated through the memory cell and the output gate, and the calculation formula is:

[0066] h t = z o ⊙tanh(c t )

[0067] where h t represents the hidden state at time t. Since the memory cell at time t is obtained by adding two parts, after time accumulation, the memory cell at time t can be a relatively large number. Therefore, the tanh activation function is used to remap the value of the memory cell to the range [-1, 1], and then perform the Hadamard product with the output gate signal to obtain the hidden state at time t. The gated signal of the output gate is used to control whether to output.

[0068] Sub-step 4-3: Use the reinforcement learning algorithm to find the appropriate timing for early rumor detection. Set up the environment. In the early rumor detection task, the environment is the rumor detection classifier. Each action output by the agent will act on the rumor detection classifier, and the rumor detection classifier will return a reward value (or possibly a penalty).

[0069] Set the state. In the early rumor detection task, the state is the global feature extracted by GCN and GRU, denoted as:

[0070]

[0071] s t represents the state at time t. This state is obtained after the fusion of the rumor features extracted by the rumor detection module through the GCN neural network and the GRU neural network, and contains the deep features of rumor propagation.

[0072] Set the action. In the early rumor detection task, there are two actions. One is "continue to add subsequent comments", and the other is "stop adding comments and immediately conduct rumor detection". The action output of the agent at time t is:

[0073]

[0074] where a t represents the action output at time t, and the set of actions is {0, 1}. When a t = 0, the action selected is "continue to add subsequent comments"; when a t = 1, the action selected is "stop adding comments and immediately conduct rumor detection". When the agent believes that the current number of comments is insufficient for rumor detection based on the current state and policy, the agent will choose to output the action of "continue to add subsequent comments"; when the agent believes that the current number of comments is sufficient for rumor detection based on the current state and policy, the agent will choose to output the action of "stop adding comments and immediately conduct rumor detection".

[0075] Set the reward. In the early rumor detection task, there are three results brought about by the action output by the agent. First, the selected action is "continue to add subsequent comments". At this time, the rumor detection does not detect the rumor, and the next rumor sub - propagation structure will be passed to the reinforcement learning decision module. Second, the selected action is "stop adding comments and immediately conduct rumor detection", and the prediction result of the rumor detection classifier is correct. Third, the selected action is "stop adding comments and immediately conduct rumor detection", and the prediction result of the rumor detection classifier is incorrect. For the above three cases, this module sets corresponding rewards and punishments as the evaluation of the agent's output action. The reward function is:

[0076]

[0077]

[0078] Among them, r(s t , a t ) represents the reward function at time t when the state is s t and the action is a t ; n t is the number of comments currently in use, and n is the total number of comments; μ t is a coefficient used to balance the reward function and avoid the situation where the number of comments is too small to conduct rumor detection on it.

[0079] For the first case, in order to enable the reinforcement learning decision-making module to conduct rumor detection earlier, a slight penalty needs to be given. ε is a small value.

[0080] For the second case, the rumor detection classifier predicts correctly, indicating that the number of comments is sufficient to conduct rumor detection on it. A large reward needs to be given to encourage the reinforcement learning decision-making module to conduct rumor detection in a timely manner. M is a large value.

[0081] For the third case, the rumor detection classifier predicts incorrectly, indicating that the current number of comments is not sufficient to conduct rumor detection. A large penalty needs to be given to prevent the reinforcement learning decision-making module from drawing premature conclusions under insufficient information. P is a large value.

[0082] Set the optimization strategy and use the PPO algorithm. The overall objective function of the PPO algorithm, that is, the reward function is:

[0083]

[0084] Among them, θ is the policy parameter of the Actor network; p θ (a t |s t ) represents the probability of outputting action a t in state s t ; π θ represents the policy with parameter θ; A θ′ (s t , a t ) is the advantage function used to estimate the quality of taking action a t in state s t . When A θ′ (s t , a t ) is positive, the probability should be increased. When A θ′ (s t , a t ) is negative, the probability should be decreased; It represents the final expected reward value obtained by weighted summation of the probability of each action selection and the obtained reward. The overall objective function is to maximize the expected reward; βKL(θ,θ′) is the penalty term for the KL divergence constraint, and β is a coefficient that can be dynamically adjusted according to the policy.

[0085] Sub-step 4-4, rumor classification and loss function. The output of the final rumor detection classifier is first linearly transformed by the hidden state, and then the softmax function is used to normalize each rumor detection category to obtain the prediction probability of each category of the rumor. The calculation formula is:

[0086]

[0087] where h t represents the hidden state of the LSTM network at time t; b rumor is the bias term; W rumor is the linear transformation matrix used to transform the dimension of the hidden state into the number of categories of rumor classification.

[0088] The loss function uses the cross-entropy loss function to optimize the early rumor detection model. The calculation method of the loss function is:

[0089]

[0090] where Loss rumor represents the loss value calculated by the loss function, N rumor represents the number of samples of rumor detection, C rumor represents the number of categories of rumor detection, represents the true probability that the i-th event rumor detection category is the j-th category, represents the predicted probability that the i-th event rumor detection category is the j-th category.

[0091] In summary, the present invention first fuses the text features of the rumor and user features to form the feature representations of the rumor source post and comment posts; secondly, the present invention uses the graph convolutional network and the gated recurrent unit to mine the deep features of rumor propagation; finally, through the reinforcement learning algorithm, the model is guided to perform early rumor detection at the appropriate time.

[0092] Embodiment 2: A rumor detection device based on text and user features disclosed in an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-mentioned rumor detection method based on text and user features.

[0093] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.< / s>

Claims

1. A rumor detection method based on text and user characteristics, characterized in that, The method includes the following steps: Step 1, extract text features and user features and fuse the two; Step 2, use a graph convolutional network to mine the local features of the rumor propagation graph; Step 3, use a gated recurrent unit to mine the temporal features of rumor propagation; Step 4, fuse the features extracted by the graph convolutional network and the gated recurrent unit, and use a reinforcement learning algorithm to find a suitable early rumor detection time point.

2. The rumor detection method based on text and user characteristics according to claim 1, wherein Step 1 specifically includes the following sub-steps: Sub-step 1-1: For the source posts, comment texts, and corresponding user features in the rumor dataset, use the RoBERTa model to extract features. RoBERTa will first add <s>Identifier, add at the end position of the sentence< / s> identifiers at the beginning of the sentence, and then add the tokens in the sentence to RoBERTa one by one for training to obtain the encoded feature representation: R = RoBERTa( <s>,Sentence1, <sep>, Sentence2,…,< / sep> < / s> ) Select from the feature representation R encoded by RoBERTa " <s>"vector R in the position s as the extracted feature,< / s> <s> Sub-step 1-2, text feature and user feature fusion. After the word embedding process, text feature vectors R are obtained after feature extraction of the source post text content and early comment text text , and user feature vectors R are obtained after feature extraction of user features user , where the user feature vector is obtained by concatenating user nickname features, user profile features, and Weibo authentication type features, and is specifically expressed as follows: R user = R nickname ||R profile ||R certification where R nickname represents the feature vector encoded by the RoBERTa model for the user nickname, R profile represents the feature vector encoded by the RoBERTa model for the user profile, R certification represents the feature vector encoded by the RoBERTa model for the Weibo authentication type, || represents the concatenation operation, and the text features and user features are concatenated and then input into the fully connected layer to obtain the fused features: h = W*(R text ||R user ) + b where R text represents text features, R user represents user features, h represents fused features, and W and b are the weights and bias terms of the linear transformation respectively.

3. The rumor detection method based on text and user characteristics according to claim 1, wherein Step 2 specifically includes the following sub-steps: Sub-step 2-1, Dynamic propagation graph construction. In this paper, according to the time when rumor comments are released, a dynamic propagation graph is constructed for the rumor dataset, as follows. First, the structure of rumor comments is divided according to the release time. Starting from the rumor sub-structure G1 with only the source post content, each time a comment message is added in chronological order, and the new comment information is added to the original sub-comments to form the next rumor sub-structure until all comment information is added to form the final rumor structure G n-1 , where n - 1 represents the number of comments in this event. Finally, the dynamic propagation structure of event C j can be represented as For each rumor sub-structure is represented as where t - 1 is the number of comments in this rumor sub-structure, Based on the rumor sub-structures formed above, construct a propagation graph for them. The rumor sub-structures are constructed in the form of a graph structure, specifically represented as represents the set of nodes in the rumor sub-structure. represents the comment action, denoted as If node N u comments on node N v , then is 1, otherwise is 0. During the construction of the propagation graph, according to the nodes and the comment relationships between nodes, the vertices and edges of the propagation graph are constructed. The feature of each vertex is represented as a fused feature obtained by fusing text features and user features. Sub-step 2-2: Use a graph convolutional network to extract node features. After obtaining the fusion features of each node and constructing a dynamic propagation graph, select to use a graph convolutional neural network to extract features of the rumor sub-propagation structure to obtain the propagation features of the rumor sub-propagation structure. The input of the graph convolutional network GCN allows for a graph structure of any shape, and generates node-level embedding representations through convolutional operations. In this paper, the input of the GCN layer is the rumor sub-structure generated above. contains t nodes and the edges between the nodes, and each node N i The corresponding feature vector is the fusion feature after fusing the text feature and the user feature. Graph The corresponding adjacency matrix is represented as Let D be a diagonal matrix and satisfy D ii = ∑ j A ij The input of the GCN layer is the fusion feature matrix H of all nodes of the current rumor sub-propagation structure. (0) Each layer obtains the hidden representation of the next layer through convolutional operations. The convolutional process is as follows: Among them, H (l) represents the hidden representation matrix of the l-th layer; I is the identity matrix; is the degree matrix of, which is a diagonal matrix and satisfies W (l) is the trainable weight matrix of the l-th layer, σ represents the activation function, and the sigmoid function is used in this article. After being encoded by two layers of graph convolutional network layers, average pool the hidden features of all nodes in the last layer to obtain the features extracted by the graph convolutional network module. The calculation method is: Among them represents the vector of the i-th row in the hidden feature matrix extracted by the second-layer graph convolutional network layer, represents the rumor sub-spreading structure features extracted by the graph convolutional network module.

4. The rumor detection method based on text and user characteristics according to claim 1, wherein Step 3 specifically includes the following sub-steps: Sub-step 3-1: Divide the comment structure of the rumor. First, arrange the source post and all comment contents in chronological order. Event C j can be expressed as where represents the fusion feature obtained by fusing the source post text feature and the user feature, represents the fusion feature obtained by fusing the (i-1)-th comment text feature and the user feature; Sub-step 3-2: Use a gated recurrent unit to extract the temporal features of rumor propagation. For a certain moment t, the input of the GRU layer is the fused feature and the hidden state h passed down from the previous node t-1 , and the output is the hidden state h of the next node t . In the calculation of the GRU layer, there are two important gates. One is the reset gate, and the other is the update gate. The function of the reset gate is to determine how to combine the input at the current moment with the previous memory, and the function of the update gate is to determine how much of the previous memory plays a role. The specific calculation methods of the reset gate and the update gate are as follows: z r = σ(W r * [h t-1 , x t + b r ) z u = σ(W u * [h t-1 , x t [ + b u ) where z r represents the reset gate, and z u represents the update gate, W r and W u are the trainable weight matrices of the reset gate and the update gate respectively, b r and b u are the bias terms of the reset gate and the update gate respectively. σ is the activation function, which is the sigmoid function here, converting the input into data within the range of [0, 1] to act as the gating signal. After obtaining the gating signal, first use the reset gating to combine the hidden state passed down from the previous node with the current input to obtain the candidate hidden state. The calculation formula is: h ′ = tanh(W * [h t-1 ⊙ z r , x t + b) where ⊙ represents the Hadamard product, i.e., element-wise multiplication; [h t-1 ⊙ z r , x t represents concatenating the Hadamard product of the previous hidden state and the reset gate with the current input; tanh is the activation function, scale the result to the range of [-1, 1], and then use the update gate to update the memory to obtain the hidden state of the next node. The calculation method is: h t = (1 - z) ⊙ h t-1 + z ⊙ h' where the range of z is [0, 1]. When z equals 0, it means directly using the hidden state of the previous moment without updating the hidden state; when z equals 1, it means using the currently updated hidden state; the value of z represents the proportion of the currently updated hidden state in the final hidden state. The larger the value of z, the more the hidden state of the past is forgotten, and h t is the output of the hidden state at the current moment and also the input of the hidden state passed to the next moment; For the hidden state extracted by the gated recurrent unit GRU, use the attention mechanism to calculate the probability weights for the fused features input at different times, so that the fused features of some comments can receive more attention, thereby improving the quality of feature extraction by the gated recurrent unit layer; The input of this attention mechanism layer is the hidden state outputs {h 1 , h 2 , …, h t} at each moment of the feature extraction layer. For a certain moment t, the similarity between the input and the output is calculated, and the dot product method is used to calculate the similarity. The calculation formula is as follows: e i = tanh(w i h i + b i ) e i represents the hidden state h at time i i The similarity with the output vector, calculate the similarity with the output vector at all times, and then perform a normalization operation on it. The calculation formula is: where α i represents the weight coefficient of the hidden state h at time i i After obtaining the weight coefficients of the hidden states at each time by using the softmax function for the normalization operation, the weighted sum of the hidden state outputs at each time is calculated according to the weight coefficients to obtain the features extracted by the final gated recurrent unit module. The calculation formula is as follows:

5. The rumor detection method based on text and user characteristics according to claim 1, characterized in that Step 4 specifically includes the following sub-steps: Sub-step 4-1, fuse the features extracted by the GCN network and the GRU network. The calculation method is: For the rumor sub-structure at time t, the features extracted by the graph convolutional network module are The features extracted by the gated recurrent neural network module are W GCN and W GRU are the parameter matrices corresponding to the graph convolutional network module and the gated recurrent neural network module respectively. b t is the bias term at time t, and the final combined global features extracted by GCN-GRU at time t are Sub-step 4-2, perform LSTM network encoding on the global features. For a certain moment t, the inputs to the LSTM network are three, namely the memory cell c passed from the previous moment t-1 , the hidden state h passed from the previous moment t-1 , and the input H at the current moment t . There are two outputs, namely the cell state c at this moment t and the hidden state h at this moment t . There are three gates in the LSTM network, namely the forget gate, the input gate, and the output gate. The calculation methods of the three gates are as follows: z i = σ(W xi H t + W hi h t-1 + b i ) z f = σ(W xf H t + W hf h t-1 + b f ) z o = σ(W xo H t + W ho h t-1 + b o ) where z i , z f and z o represent the input gate, forget gate, and output gate respectively; W xi and W hi are the linear transformation matrices of the current input and the previous hidden state in the input gate, and b i is the bias term of the input gate; W xf and W hf is the linear transformation matrix of the input at the current time and the hidden state at the previous time in the forget gate, and b f is the bias term of the forget gate; W xo and W ho are the linear transformation matrices of the input at the current time and the hidden state at the previous time in the output gate, and b o is the bias term of the output gate, σ is the activation function, and the sigmoid function is used. Here, the role of the input gate is to determine whether to ignore the input data, the role of the forget gate is to reduce the value towards 0, and the role of the output gate is to determine whether to use the hidden state. In addition to these three gated states, there are also candidate memory units in the LSTM network for processing input data, and the calculation method is: Among them is a candidate memory cell, and W xc and W hc are linear transformation matrices of the input at the current moment and the hidden state at the previous moment in the candidate memory cell. b c is the bias term of the candidate memory cell. Tanh is an activation function that converts the result to the range of [-1, 1]. After obtaining the three gating signals and the candidate memory cell, the memory cell is calculated through the following calculation method: c t represents the memory cell. First, perform the Hadamard product operation on the forget gate signal and the memory cell at the previous moment, then calculate the Hadamard product of the input gate signal and the candidate memory cell, and then combine the two to obtain the memory cell. Since both the forget gate signal and the input gate signal are within the range of [0, 1] and are independent of each other, the memory cell at this moment in the LSTM can either take into account both the memory cell at the previous moment and the input at the current moment, or ignore both the memory cell at the previous moment and the input at the current moment, which is more flexible and free than the GRU. After obtaining the memory cell, the hidden state can be calculated through the memory cell and the output gate, and the calculation formula is: h t = z o ⊙tanh(c t ) where h t represents the hidden state at time t. Since the memory unit at time t is obtained by adding two parts, and after time accumulation, the memory unit at time t can be a relatively large number, the tanh activation function is used to remap the value of the memory unit back to the range of [-1, 1]. Then, it performs a Hadamard product with the output gate signal to obtain the hidden state at time t. The gating signal of the output gate is used to control whether to output or not. Sub-step 4-3, use a reinforcement learning algorithm to find a suitable early rumor detection opportunity, set the environment. In the early rumor detection task, the environment is the rumor detection classifier. Each action output by the agent will act on the rumor detection classifier, and the rumor detection classifier will return a reward value or a penalty value; Set the state. In the early rumor detection task, the state is the global features extracted by the GCN and GRU, expressed as: s t represents the state at time t, which is obtained after the rumor detection module fuses the rumor features extracted by the GCN neural network and the GRU neural network, and contains the deep features of rumor propagation. Set the action. In the early rumor detection task, there are two actions. One is "continue to add subsequent comments", and the other is "stop adding comments and immediately conduct rumor detection". The action output of the agent at time t is: where a t represents the action output at time t, and the set of actions is {0, 1}. When t a = 0, the action selection is "continue to add subsequent comments"; when t a = 1, the action selection is "stop adding comments and immediately conduct rumor detection". When the agent believes that the current number of comments is insufficient for rumor detection based on the current state and policy, the agent will choose to output the action of "continue to add subsequent comments"; when the agent believes that the current number of comments is sufficient for rumor detection based on the current state and policy, the agent will choose to output the action of "stop adding comments and immediately conduct rumor detection". Set the reward. In the early rumor detection task, there are three results brought by the action output by the agent. The first one is that the selected action is "continue to add subsequent comments". At this time, the rumor detection does not detect the rumor, and the next rumor sub-propagation structure will be passed to the reinforcement learning decision module; the second one is that the selected action is "stop adding comments and immediately conduct rumor detection", and the prediction result of the rumor detection classifier is correct; the third one is that the selected action is "stop adding comments and immediately conduct rumor detection", and the prediction result of the rumor detection classifier is wrong. For the above three situations, this module sets corresponding rewards and punishments as an evaluation of the action output by the agent. The reward function is: Among them, r(s t ,a t ) represents the reward function at time t when the state is s t and the action is a t ; n t is the number of comments currently used, and n is the total number of comments; μ t is a coefficient used to balance the reward function and avoid the situation where the number of comments is too small to conduct rumor detection on it For the first situation, in order to let the reinforcement learning decision module conduct rumor detection earlier, a slight penalty needs to be given. ε is a relatively small value, For the second situation, the rumor detection classifier predicts correctly, indicating that the number of comments is sufficient to conduct rumor detection on it. A relatively large reward needs to be given to encourage the reinforcement learning decision module to conduct rumor detection in time. M is a relatively large value, For the third case, the rumor detection classifier makes a wrong prediction, indicating that the current number of comments is not sufficient for rumor detection. A relatively large penalty needs to be given to prevent the reinforcement learning decision-making module from drawing premature conclusions under insufficient information. P is a relatively large value. Set the optimization strategy and use the PPO algorithm. The overall objective function of the PPO algorithm, that is, the reward function is: where θ is the policy parameter of the Actor network; p θ (a t |s t ) represents the probability of outputting action a t under state s t ; π θ represents the policy with parameter θ; A θ′ (s t ,a t ) is the advantage function, which is used to estimate the goodness of taking action a t under state s t . When A θ′ (s t ,a t ) is positive, the probability should be increased; when A θ′ (s t ,a t ) is negative, the probability should be decreased; represents the final expected reward value obtained by weighted summing the probability of each action selection and the obtained reward. The overall objective function is to maximize the expected reward; βKL(θ,θ′) is the penalty term for the KL divergence constraint, and β is a coefficient that is dynamically adjusted according to the policy. Sub-step 4-4, rumor classification and loss function. The output of the final rumor detection classifier is first linearly transformed by the hidden state, and then the softmax function is used to normalize each rumor detection category to obtain the predicted probability of each category of the rumor. The calculation formula is: where h t represents the hidden state of the LSTM network at time t; b rumor is the bias term; W rumor is the linear transformation matrix used to transform the dimension of the hidden state into the number of rumor classification categories, The loss function uses the cross-entropy loss function to optimize the early rumor detection model. The calculation method of the loss function is: Among them, Loss rumor represents the loss value calculated by the loss function, and N rumor represents the number of samples for rumor detection, and C rumor represents the number of categories for rumor detection. represents the true probability that the rumor detection category of the i-th event is the j-th category, represents the predicted probability that the rumor detection category of the i-th event is the j-th category.

6. A rumor detection device based on text and user characteristics, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the rumor detection method based on text and user characteristics described in any one of claims 1-5. < / s>