Social media robot detection method and system based on multi-view multi-modal network
By adopting a multi-view multi-modal network in social media robot detection and weighted integration of different model outputs combined with multi-head attention mechanism, the shortcomings of existing detection methods in dealing with new social robots are solved, and the accuracy and robustness of detection are improved.
Patent Information
- Application Number
- CN202510124175.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
Existing social media robot detection methods gradually lose their advantages in dealing with new social robots, and multimodal feature fusion strategies face problems such as data sparsity, noise influence, computing complexity and resource consumption.
The detection method based on multi-view multi-modal network is adopted, and feature extraction and preliminary classification prediction are performed through the Roberta model, BotRGCN model and MLP model, and the output of different models is weighted and integrated with the multi-head attention multi-view fusion model to improve the accuracy and robustness of the detection.
By better capturing higher-order correlations between different data types, the bias of classification results is reduced, and the accuracy of robot detection and model interpretability are improved.
Smart Images

Figure CN120067318A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social media robot detection, and particularly to a method and system for social media robot detection based on a multi-view multi-modal network. Background Art
[0002] Social media robot accounts are automated accounts that perform various functions. In recent years, significant progress has been made in the field of social media robot detection, and the research methods can be mainly classified into the following categories: feature-based methods, text sequence-based methods, user graph relationship-based methods, multi-modal feature fusion-based methods, and multi-modal classification result integration-based methods.
[0003] Feature-based methods extract features from user attribute features and tweets and use machine learning classifiers for robot detection. Text sequence-based methods mainly rely on analyzing the text content posted by users, using deep learning models to extract features and train, and finally obtaining a classifier. User graph relationship-based methods are capable of capturing complex interaction relationships and structural features in the social network and using these features for model training. Multi-modal feature fusion-based methods use different types of data, after feature encoding, and then through operations such as feature concatenation, feature weighting, or feature transformation, and input them into the model for training. Multi-modal classification result integration-based methods aim to improve the overall classification performance by combining the prediction results of multiple classifiers.
[0004] Such as Figure 1As shown, existing detection methods have their own advantages and disadvantages. In terms of detection effect, the detection method based on graph relationship is particularly prominent in terms of accuracy. However, with the continuous progress of artificial intelligence technology, social robot manipulators continuously optimize the attribute characteristics of robots, resulting in the detection methods relying on a single data type gradually losing their advantages when dealing with new types of social robots. The emergence of multi-modal technology enables researchers to integrate multiple information from different data sources. This can not only analyze various input types such as the text content, image information, audio data, and social relationship network of users simultaneously, but also, by fusing these different modal data, enable the detection system to more precisely capture the complex behavior patterns and hidden abnormal characteristics of different users, providing more comprehensive capabilities for social robot detection. For this reason, researchers in related fields have begun to explore the use of multi-modal feature fusion methods for social media robot detection. However, feature fusion strategies face many potential challenges and drawbacks, such as data sparsity, the impact of data noise, the complexity of data calculation, and large resource consumption. To address the above problems, existing methods based on the integration of multi-modal classification results (such as voting method, Stacking method, Boosting method, etc.) can, to a certain extent, alleviate the drawbacks brought by feature fusion. These methods are widely used in the field of machine learning to improve the performance of classifiers. However, the methods based on the integration of multi-modal classification results also have many limitations.
[0005] The high-order correlations between different data types are often ignored. Traditional methods based on the integration of classification results usually directly combine the classification results from different data types and fail to fully consider the high-order correlations and potential interactions between these data. This kind of neglect may lead to the model not fully utilizing the complementary information between multiple data types. The classification results may be biased. Simple methods for integrating classification results may be more inclined to the classification results of certain specific types of data, especially when these data perform well when used alone. This kind of bias may cause the overall model performance to be limited by the single data type with the best performance, and unable to fully explore and utilize the supplementary information of other data types. Summary of the Invention
[0006] To at least partially solve the problems that the high-order correlations between different data types in multi-modal data classification are often ignored and the classification results may be biased, the present invention provides a social media robot detection method and system based on a multi-view multi-modal network. The present invention uses three different models for feature extraction and preliminary classification prediction, namely the Roberta model, the BotRGCN model, and the MLP model. The Roberta model uses a variant of the pre-trained language model BERT for text feature extraction, the BotRGCN model captures heterogeneous relationship information of users in the social network through a graph neural network architecture, and the MLP model models the numerical and boolean attributes of users. Finally, the output probabilities of the Roberta model, the BotRGCN model, and the MLP model are integrated by using a multi-head attention multi-view fusion model (AttVCDN), and weighted according to the importance of different models for the final classification result during the integration process, further improving the accuracy and robustness of robot detection. The present invention better captures the high-order correlations between different data types through the multi-head attention multi-view fusion model, reducing the possible bias in the classification results.
[0007] To achieve the above object, the technical solution of the present invention is:
[0008] The first aspect of the present invention proposes a social media robot detection method based on a multi-view multi-modal network, including:
[0009] Step 1: Collect the text data of users, the user graph relationship type data, and the user attribute features for subsequent robot detection;
[0010] Step 2: Input the text data into the Roberta model to obtain the probability distribution of the text sequence data for enhancing the text representation ability;
[0011] Input the user graph relationship type data into the Roberta model to obtain the probability distribution of the user graph relationship type data for effectively capturing the information transmission in the multi-relationship graph structure;
[0012] Input the user attribute features into the MLP model to obtain the probability distribution of the user attribute features for effectively capturing the potential patterns of social media accounts;
[0013] Step 3: Input the probability distribution of the text sequence data, the probability distribution of the user graph relationship type data, and the probability distribution of the user attribute features into the multi-head attention multi-view fusion model to obtain the classification result and complete the social media robot detection.
[0014] Further, the text data includes user description features and tweet features;
[0015] The user graph relationship type data includes descriptive features, tweet features, numerical attributes, and categorical attributes;
[0016] The user attribute features include numerical features and boolean features.
[0017] Furthermore, the Roberta model is expressed by the following formula:
[0018] X combined = concat(W des X des + W tweet X tweet )
[0019] Z Roberta = Softmax(W output · Dropout(σ(W input · X combined )))
[0020] Where X combined is the concatenated feature, X des is the descriptive feature, X tweet is the tweet feature, concat is the concatenation function, Z Roberta is the probability distribution of the text sequence data, W des is the linear transformation matrix of the descriptive feature, W tweet is the linear transformation matrix of the tweet feature, Softmax is the normalization exponential function, W output is the linear transformation matrix, Dropout is random inactivation, σ is the activation function, W input is the linear transformation matrix of the concatenated feature.
[0021] Furthermore, the BotRGCN model is expressed by the following formula:
[0022] Z BoRGCN = Softmax(W out 2 · σ(W out 1 · H final ))
[0023] Where Z BotRGCN is the probability distribution of the user graph relationship type data, Softmax is the normalization exponential function, W out 2 is the output layer weight, σ is the activation function, W out 1 is the output layer weight, H final is the node representation after the relational graph convolutional network and residual connection.
[0024] Furthermore, the multi-head attention multi-view fusion model includes a multi-head attention mechanism and a VCDN model;
[0025] The multi-head attention mechanism is used to process the concatenated three-dimensional tensor after concatenating the probability distribution of text sequence data, the probability distribution of user graph relationship type data, and the probability distribution of user attribute features, to obtain the attention-weighted probability distribution of text sequence data, the attention-weighted probability distribution of user graph relationship type data, and the attention-weighted probability distribution of user attribute features;
[0026] The VCDN model is used for the attention-weighted probability distribution of text sequence data, the attention-weighted probability distribution of user graph relationship type data, and the attention-weighted probability distribution of user attribute features, to obtain a classification result and complete social media robot detection.
[0027] The second aspect of the present invention proposes a social media robot detection method based on a multi-view multi-modal network, including:
[0028] A collection module, configured to collect the text data, user graph relationship type data, and user attribute features of a user, facilitating subsequent robot detection;
[0029] A processing module, configured to input the text data into the Roberta model to obtain the probability distribution of text sequence data, for enhancing the text representation ability;
[0030] Input the user graph relationship type data into the Roberta model to obtain the probability distribution of user graph relationship type data, facilitating the effective capture of information transmission in the multi-relationship graph structure;
[0031] Input the user attribute features into the MLP model to obtain the probability distribution of user attribute features, facilitating the effective capture of potential patterns of social media accounts;
[0032] A classification module, configured to input the probability distribution of text sequence data, the probability distribution of user graph relationship type data, and the probability distribution of user attribute features into the multi-head attention multi-view fusion model to obtain a classification result and complete social media robot detection.
[0033] The third aspect of the present invention proposes an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the social media robot detection method based on a multi-view multi-modal network as described in the first aspect above.
[0034] The fourth aspect of the present invention proposes a computer-readable storage medium, where the storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the social media robot detection method based on a multi-view multi-modal network as described in the first aspect above.
[0035] Advantages of the present invention:
[0036] (1) The present invention extracts features and conducts preliminary classification prediction through three different models, and then detects and identifies social media robots according to the multi-head attention multi-view fusion model, improving the accuracy of detection. The present invention can better capture the high-order correlations between different data types and reduce the possible biases in the classification results.
[0037] (2) The present invention first introduces VCDN into the social media robot detection task. The results show that VCDN is effective in improving the performance of the model. To further improve the performance of the model, the present invention optimizes the VCDN module and adds a multi-head attention mechanism to learn the influence of different modality data on the detection results. The final experimental results show that after adding the attention mechanism, both the detection accuracy and F1-score are further improved, and the interpretability of the model is also significantly enhanced. By introducing the attention mechanism, AttVCDN can not only capture and utilize the relationships between different modality views, but also accurately identify the influence of different modality data on the classification results, making the decision-making process of the model more transparent and interpretable. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of the current social media robot detection related algorithm provided by an embodiment of the present invention.
[0039] Figure 2 It is one of the flowcharts of the social media robot detection method based on a multi-view multi-modal network provided by an embodiment of the present invention.
[0040] Figure 3 It is another flowchart of the social media robot detection method based on a multi-view multi-modal network provided by an embodiment of the present invention.
[0041] Figure 4 It is a schematic diagram of the Roberta model provided by an embodiment of the present invention.
[0042] Figure 5 It is a schematic diagram of the BotRGCN model provided by an embodiment of the present invention.
[0043] Figure 6 It is a schematic diagram of the multi-head attention multi-view fusion model provided by an embodiment of the present invention.
[0044] Figure 7 It is a schematic diagram of the comparison of different multi-modal data combinations provided by an embodiment of the present invention.
[0045] Figure 8 It is a schematic diagram of the comparison between the VCDN training process and the AttVCDN training process provided by an embodiment of the present invention.
[0046] Figure 9 This is the architecture diagram of the social media robot detection system based on the multi-view multi-modal network provided by the embodiments of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Embodiment 1
[0049] As Figure 2 and Figure 3 shown, the social media robot detection method (BotAttVCDN) based on the multi-view multi-modal network includes:
[0050] S101: Collect the text data, user graph relationship type data and user attribute features of users.
[0051] Specifically, the text data includes user description features and tweet features. The user graph relationship type data includes description features, tweet features, numerical attributes and categorical attributes. The user attribute features include numerical features and boolean features.
[0052] S102: Input the text data into the Roberta model to obtain the probability distribution of the text sequence data.
[0053] Input the user graph relationship type data into the Roberta model to obtain the probability distribution of the user graph relationship type data.
[0054] Input the user attribute features into the MLP model to obtain the probability distribution of the user attribute features.
[0055] S103: Input the probability distribution of the text sequence data, the probability distribution of the user graph relationship type data and the probability distribution of the user attribute features into the multi-head attention multi-view fusion model to obtain the classification result and complete the social media robot detection.
[0056] The present invention uses three different models for feature extraction and preliminary classification prediction, namely the Roberta model, the BotRGCN model, and the MLP model. The Roberta model uses a variant of the pre-trained language model BERT for text feature extraction, the BotRGCN model captures heterogeneous relationship information of users in the social network through a graph neural network architecture, and the MLP model models the numerical and boolean attributes of users. Finally, by using AttVCDN to integrate the output probabilities of the Roberta model, the BotRGCN model, and the MLP model, and weighting according to the importance of different models to the final classification result during the integration process, the accuracy and robustness of robot detection are further improved.
[0057] Example 2
[0058] Based on the above embodiments, the embodiment of the present invention provides the specific structure of the Roberta model, which specifically includes:
[0059] The Roberta model is a natural language processing model pre-trained and fine-tuned based on BERT (Bidirectional Encoder Representations from Transformers), aiming to enhance text representation capabilities. The model is mainly pre-trained on a large-scale corpus in an unsupervised manner, and then fine-tuned for specific tasks through supervised learning. In the implementation of the present invention, the Roberta model realizes the robot detection classification task through a multi-layer perceptron (MLP) structure and a specific loss function design.
[0060] In the implementation of the Roberta model, the key to data preprocessing is to convert text data into numerical features. The text data of tweet texts and user descriptions are used to extract features through a pre-trained language model. The pre-trained feature representation used in the present invention is 768-dimensional. The description feature and the tweet feature are respectively denoted as X des , X tweet , and the extracted tweet feature and description feature are converted through a linear layer, and then they are concatenated to obtain a concatenated feature representation X combined . The concatenated feature is represented by the following formula:
[0061] X combined = concat(W des X des + W tweet X tweet )
[0062] where X combined is the concatenated feature, X des is the description feature, and X tweetFor tweet features, concat is the concatenation function, and W des is the linear transformation matrix for describing features, and W tweet is the linear transformation matrix for tweet features.
[0063] The core of the Roberta model is a classifier composed of a multi-layer perceptron (MLP). The input of the MLP is the concatenated feature X combined , and through multiple non-linear transformations, the model gradually extracts and refines text features and finally maps them to the classification label space, which can be expressed by the following formula.
[0064] Z Roberta = Softmax(W output · Dropout(σ(W input · X combined )))
[0065] where Z Roberta is the probability distribution of the text sequence data, σ is the activation function, Softmax is the normalized exponential function, W output is the linear transformation matrix, Dropout is random inactivation, σ is the activation function, and W input is the linear transformation matrix for the concatenated features.
[0066] The specific implementation process of Roberta is as Figure 4 shown. First, the text data of the tweet is extracted. Here, the tweet text data and the description text data are selected in the present invention. Secondly, the obtained tweet text data and description text data are respectively subjected to feature extraction through a pre-trained language model to obtain their respective feature representations, and then the two parts of the features are concatenated (concat) to obtain the final text feature representation. The final text feature representation is subjected to non-linear transformation through an MLP (multi-layer perceptron), and the result is mapped to the classification label space to obtain the final probability distribution.
[0067] Roberta is pre-trained based on the BERT model and deeply trained through unsupervised learning on a large-scale corpus. By combining tweet text and user description features, the model can integrate multiple information sources, enabling it to capture complex language patterns and context information, improving the generalization ability and accuracy of the model, and making it perform well in the robot detection task.
[0068] Example 3
[0069] Based on the above embodiments, the present invention provides the specific structure of the BotRGCN model, which specifically includes:
[0070] The BotRGCN model is a deep learning architecture based on the Graph Neural Network (GNN), specifically designed to process social network data containing different types of features. Based on the traditional Graph Convolutional Network (GCN), it effectively captures information transmission in multi-relational graph structures by introducing the Relational Graph Convolutional Network (RGCN). This design enables BotRGCN to handle complex heterogeneous data, fuse various node features (such as text descriptions, tweet content, numerical attributes, and categorical attributes), and propagate node information through the graph structure, thereby enhancing the performance of node classification tasks.
[0071] RGCN extends the GCN paradigm by supporting multiple types of edges (relationships) to model complex node relationships. In traditional GCN, the convolution operation is limited to unweighted graphs or homogeneous graphs, while RGCN enables the model to distinguish and learn different types of relationships by separately learning convolutional kernel parameters for each relationship type. This idea of multi-relational graph convolution provides theoretical support for processing social network data containing multiple heterogeneous features and complex relationships.
[0072] The input features of the model mainly include four types: descriptive features (des), tweet features (tweet), and numerical attributes (num_prop). These features are embedded through a linear transformation layer and combined into a complete node representation after the activation function, which can be expressed by the following formula:
[0073] H des =σ(W des X des )H tweet =σ(W tweet X tweet )
[0074] H num =σ(W num X num )H cat =σ(W cat X cat )
[0075]
[0076] where X des , X tweet , X num and X cat are descriptive features, tweet features, numerical attributes, and categorical attributes respectively, and W des , W tweet , W num and Wcat They are linear transformation matrices for describing features, tweet features, numerical attributes, and categorical attributes respectively. σ is the representation obtained after activation function embedding, and H des , H tweet , H num and H cat are the feature embeddings of descriptive features, tweet features, numerical attributes, and categorical attributes respectively. H input is the overall feature embedding, is the concatenation operation of features.
[0077] After obtaining the feature representation of the nodes, the model uses a Relational Graph Convolutional Network (RGCN) to propagate node information in the graph structure. RGCN extends the traditional GCN by supporting multiple types of edges to model complex node relationships. For each node, the convolutional operation of RGCN is shown in the following formula:
[0078] H cat = σ(W cat X cat )
[0079]
[0080] where H v (l+1) is the feature representation of node v at the l-th layer, σ is the activation function, is the set of all relationship types, represents the set of neighbor nodes related to node v under the relationship type r, c v,r is the normalization coefficient, and W r (l) is the weight matrix corresponding to the relationship type r. W 0 (l) is the self-loop weight matrix. Through multiple layers of RGCN, the model can capture the features of nodes and their neighbor nodes, thereby enhancing the representation ability of nodes.
[0081] To prevent gradient vanishing and model overfitting, the model introduces residual connections and regularization operations in the RGCN layer. Residual connections allow the model to learn incremental features at each layer without relying entirely on feature transformation between layers. Finally, the model maps the node representation to the final class probability space through a fully connected layer and an output layer. The calculation process of the output layer is shown in formula (5):
[0082] Z BotRGCN = Softmax(W out 2 ·σ(W out 1 ·H final ))
[0083] Among them, Z BotRGCN is the probability distribution of user graph relationship type data, W out 2 is the output layer weight, σ is the activation function, W out 1 is the output layer weight, H final is the node representation after the graph convolutional network and residual connection. Softmax is the normalized exponential function. Through the Softmax operation, the model maps the nodes to the classification label space and predicts the probability that the nodes belong to each category.
[0084] The specific process of the BotRGCN model is as Figure 5 shown. The architecture of BotRGCN has high scalability. It can embed and fuse various features by adjusting the types of input features and the depth of convolutional layers, and use RGCN for information propagation and representation learning between nodes. Techniques such as residual connection and regularization are used in the model to ensure the stability and generalization ability of the model. These operations enable the BotRGCN model to capture rich node relationships and context information when dealing with complex graph structures and heterogeneous data, thereby improving the performance of social robot detection and classification tasks.
[0085] Example 4
[0086] Based on the above embodiments, the embodiments of the present invention provide an introduction to the BotRGCN model, specifically including:
[0087] In social network analysis, the behavior and attribute characteristics of users can usually reveal the authenticity and activity of accounts. The present invention constructs a multi-layer perceptron (MLP) model and uses various features related to user accounts to perform classification tasks to accurately distinguish real users and potential fake accounts.
[0088] These features include numerical features and boolean features. The numerical features cover the number of followers, the number of following, the number of tweets, the number of times listed, etc., while the boolean features reflect whether there is a personal profile, whether to use the default avatar, whether to provide location information, etc. These features are combined to form an input feature vector X∈R^(N×D), where N is the number of samples and D is the feature dimension.
[0089] The MLP model effectively captures the potential patterns of social media accounts by modeling various features of user behavior and attributes. Although the structure is relatively simple, the MLP model has strong generalization ability and is especially suitable for processing high-dimensional and heterogeneous data. In classification tasks, combined with appropriate feature engineering, regularization methods and optimization strategies, the MLP can still maintain excellent performance in the case of class imbalance, making it outstanding in social robot detection tasks.
[0090] Example 5
[0091] Based on the above embodiments, an embodiment of the present invention provides a structure of a multi-head attention multi-view fusion model (AttVCDN), which specifically includes:
[0092] In multi-modal data, the influence degrees of different types of data views on the classification results may vary significantly. The traditional VCDN model fails to learn this difference, resulting in limited model performance in the case of large differences in classification accuracy. To solve this problem and to be able to learn the influence degree of multi-modal data on the results in the social media robot detection task, the present invention introduces a multi-head attention mechanism into the VCDN module. Through this mechanism, the present invention hopes to capture the interaction relationships between different data modalities and learn the importance of each modality in the classification task. For the social media robot detection task, different types of data have different influences on the final results, and there are also significant differences in the quantity and quality of multi-modal data. Therefore, learning the importance factors of different modalities not only helps to improve the performance of the model, but also can further enhance the generalization ability and robustness of the model.
[0093] The specific implementation of the AttVCDN module is as Figure 6 shown. First, assume that the prediction probabilities in three models are respectively Then, the output probabilities of these three models are concatenated into a three-dimensional tensor P J , and this tensor represents the prediction probabilities from different models. As shown in the following formula:
[0094]
[0095] where P J is a three-dimensional tensor with dimensions (3, C), 3 represents the number of models, and C is the number of categories.
[0096] Next, the present invention inputs the concatenated tensor into the multi-head attention mechanism. First, queries, keys, and values are generated through linear transformation. Each attention head calculates these three vectors through its own weight matrix, and the specific process is as shown in the following formula:
[0097] Q = P J W Q
[0098] K = P J W K
[0099] V = P J W V
[0100] where W Q , W K and WV They are the weight matrices of the linear transformation respectively, used to map the concatenated input P J to query, key, and value. K is the key, Q is the query, and V is the value.
[0101] Next, the dot - product attention mechanism is used to calculate the interaction relationship between different model outputs. The calculation of dot - product attention is shown in the following formula:
[0102]
[0103] Among them, Attention(Q, K, V) is the value of dot - product attention, d k is the dimension of the key matrix, used to scale the dot - product result to prevent the numerical value from being too large or too small. softmax is the normalized exponential function, which is used to normalize the attention weights so that the sum of the attention weights is 1.
[0104] In the multi - head attention mechanism, multiple attention heads calculate the attention scores in parallel. Each head independently performs the above - mentioned attention operation. For the i - th head, the calculation is shown in the following formula:
[0105] head i = Attention(Q i , K i , V i )
[0106] Among them, head i is the i - th head.
[0107] The outputs of all heads will be concatenated to obtain the final attention result, as shown in the following formula:
[0108] MultiHead(P)=Concat(head 1 , …, head h )W O
[0109] Among them, h is the number of attention heads, P is the probability distribution, Concat means concatenating the outputs of all attention heads, W O is a linear transformation matrix, used to project the concatenated result to the original dimension to obtain the output of multi - head attention. After that, the present invention uses residual connection to add it to the original input and processes it through layer normalization to obtain the attention - weighted probability distribution of text sequence data, the attention - weighted probability distribution of user graph relationship type data, and the attention - weighted probability distribution of user attribute features, as shown in the following formula:
[0110] P out = ReLU(LayerNorm(P + MultiHead(P))Wf )
[0111] Among them, P out is the attention-weighted probability distribution of text sequence data, user graph relationship type data, or user attribute features, ReLU is a function, LayerNorm is layer normalization, and W f is a weight matrix used to perform a linear transformation on the processed data (after steps such as multi-head self-attention, residual connection, layer normalization, ReLU activation, etc.).
[0112] Input the three probability distributions after attention-weighted processing into the VCDN (View Correlation Discovery Network) model, fuse them through a fully connected network, and output the final classification result to obtain the final classification probability.
[0113] Preferably, in order to further utilize the VCDN to obtain the correlation between different types of data, the core module in the VCDN (directly fusing probability distributions of different modalities in an element-wise multiplication manner) is changed to use a more advanced self-supervised feature contrastive learning method to obtain the correlation between different types of data. In self-supervised feature contrastive learning, by constructing positive and negative sample pairs, the model can learn the similarity and difference between different modalities in the feature space. This method enhances the cohesion of modal features of the same category in the feature space through a contrastive loss function, while improving the separability of modal features of different categories, thereby effectively capturing the interaction information between modalities.
[0114] Specifically, in the training implementation of the VCDN, for the attention-weighted probability distribution of text sequence data, the attention-weighted probability distribution of user graph relationship type data, and the attention-weighted probability distribution of user attribute features and First, calculate the similarity between them, usually represented by cosine similarity:
[0115]
[0116] Among them, and respectively represent different attention-weighted probability distributions, and sim is the cosine similarity.
[0117] Construct a similarity matrix S, and each element S ij represents the similarity between modalities and :
[0118]
[0119] Among them, τ is the temperature parameter, which controls the smoothness of the similarity.
[0120] For each mode The contrast loss can be defined as:
[0121]
[0122] Among them, S ij and S ik Representing modality Similarity with positive samples (same category) and negative samples (different category), For each mode Contrastive loss increases the similarity between positive samples and decreases the similarity between negative samples.
[0123] For the contrast loss of the three modalities, the final total contrast loss is:
[0124]
[0125] Among them, L contrastive is the total contrast loss.
[0126] By optimizing the total contrastive loss, the model is able to better bring similar modalities closer together and push different modalities further apart in the feature space.
[0127] The multi-head attention mechanism provides a weight adjustment mechanism for the model by learning the interaction between each modality, enabling it to focus on more important modal information. The VCDN model uses these attention-weighted outputs to further enhance the model's performance and generalization capabilities for multimodal data, improving the classification accuracy of social media robot detection tasks.
[0128] Example 6
[0129] Based on the above embodiments, the embodiments of the present invention provide a verification process of a social media robot detection method based on a multi-view multimodal network, which specifically includes:
[0130] In order to evaluate and analyze the performance of the AttVCDN method, the present invention uses three open source datasets for experiments, namely the Cresci-2015 dataset as shown in Table 1, the TwiBot-20 dataset as shown in Table 2, and the TwiBot-22 dataset as shown in Table 3, and compares them with 11 social media robot detection methods.
[0131] Table 1. Cresci-2015 dataset feature distribution
[0132] Table 2 Feature distribution of the Cresci-2015 dataset
[0133]
[0134] The statistical data of the Cresci-2015 dataset is shown in Table 1: HUM (Human Dataset) represents data generated by real human users. FAK (Fake Dataset) represents data generated by fake or robot accounts. Accounts is the number of accounts, tweet is the total number of tweets, followers is the number of followers, friends is the number of friends, and total is the sum of the number of followers and friends.
[0135] Table 2 Feature distribution of the TwiBot-20 dataset
[0136]
[0137] The statistical data of the TwiBot-20 dataset is shown in Table 2, where User is the number of users, Property is the number of user property items, Tweet is the total number of tweets, and Follow is the number of follow relationships.
[0138] Table 3 Feature distribution of the TwiBot-22 dataset
[0139]
[0140] The statistical data of the TwiBot-22 dataset is shown in Table 3, where User is the number of users, Tweet is the number of tweets, Follow is the number of follow relationships, Follower is the number of followers, Bot is the number of robot accounts, and Human is the number of human accounts.
[0141] The comparative analysis of human and robot accounts in these three datasets helps researchers better understand the behavioral characteristics of human users, thereby designing more effective robot detection algorithms.
[0142] The present invention compares BotAttVCDN with Twitter robot detection models based on user attribute features, text, and graphs. The specific methods and datasets adopted by each model are shown in Table 4.
[0143] Table 4 Summary of Twitter robot detection models based on user attribute features, text, and graphs
[0144]
[0145]
[0146] The experimental results on the three datasets of Cresci-2015, TwiBot-20 and TwiBot-22 are shown in Table 5. By analyzing the experimental results shown in Table 5, it can be found that the AttVCDN model has demonstrated its significant advantages in Twitter bot detection. First of all, as the first model to apply VCDN to the social media bot detection task, AttVCDN achieved the highest accuracy and F1-score on all datasets, significantly outperforming other advanced models such as RGT and BotMOE. Secondly, AttVCDN is superior to the original VCDN model on each dataset, which shows the effectiveness of the attention mechanism in enhancing feature learning and importance modeling. Compared with other models, AttVCDN not only performs well in terms of accuracy, but also reaches the optimal level in the balanced recognition ability (F1-score) of positive and negative samples, proving its excellent performance in dealing with diverse and complex features. Therefore, AttVCDN has strong competitiveness and generalization ability in the social media bot detection task. Table 5 is as follows.
[0147] Table 5 Performance of AttVCDN relative to the baseline on three Twitter bot detection datasets
[0148]
[0149] By analyzing Figure 7 the changes in the bar chart, it can be found that under the AttVCDN model, the advantages of the three-model combination (MLP + BERT + GCN) are becoming more and more obvious. In the Cresci-2015, TwiBot-20 and TwiBot-22 datasets, as the task complexity and data feature diversity increase continuously, the ability of single models and two-model combinations to capture complex data relationships is gradually limited, showing a bottleneck in accuracy. The three-model combination can effectively integrate the characteristics of each model: MLP is good at dealing with linear features, BERT can capture context information, and GCN is good at mining graph structure features. This diversity plays an important role when facing complex and diverse datasets, making the three-model combination far ahead in terms of performance and generalization ability. Therefore, as the data types and features increase, the three-model combination is more adaptable and can achieve better performance in various tasks, which is the best strategy for dealing with complex data.
[0150] From Figure 8From the movement trend of the line chart, it can be seen that the performance trends of AttVCDN and VCDN are basically the same. Both show a rapid upward trend in the initial stage and then tend to be stable, indicating that both methods can improve performance relatively quickly in the initial stage. However, as time goes by, the performance improvement gradually slows down. In all the charts, AttVCDN (red line) is always slightly better than VCDN (black line) in the later stage, which shows that as the number of training times increases, AttVCDN has more advantages in performance and performs better in the optimization of specific scenarios, parameters or data. The curve of AttVCDN is smoother, meaning that the process of its performance improvement is more stable, while there are some fluctuations in VCDN, indicating that AttVCDN is more robust to changing conditions or inputs. As time or the number of samples increases, the performance growth finally tends to saturation, indicating that both reach the optimal performance state under specific conditions, and AttVCDN still maintains a relatively high performance level when reaching the saturation state. Generally speaking, AttVCDN is superior to VCDN in overall performance, showing higher efficiency and stability, and performing more excellently throughout the process.
[0151] To evaluate the impact of each module in the AttVCDN model on the detection results, ablation experiments were conducted by removing or modifying different modules in the model and analyzing their impact on the overall performance. The specific implementation methods are as follows:
[0152] (1) BotMOE: Remove the VCDN module and adopt the method of multi-modal feature fusion.
[0153] (2) KNN, XGBoost, Randomforest: Remove the VCND module from the three models and adopt the method of integrating multi-modal classification results to detect social media robots. Specifically, first use multi-modal classifiers, such as MLP for user attribute features, Roberta for text sequence data, and BotRGCN for graph relationship data, to perform binary classification predictions on data of different modalities respectively. Then, combine the classification results of these different modalities and comprehensively process these multi-modal prediction results through traditional machine learning methods such as KNN, RandomForest, or an integrated algorithm such as XGBoost to integrate the classification results of each modality and obtain the final detection output.
[0154] (3) VCDN: Remove the attention mechanism module and only use the classification probability distribution of the VCDN part to integrate multi-modal data to achieve the detection of social media robots.
[0155] The ablation experiment results of the AttVCDN model on three Twitter datasets are shown in Table 6.
[0156] Table 6 Ablation Experiment Results of the AttVCDN Model on Three Twitter Datasets
[0157]
[0158]
[0159] Table 6 shows the ablation experiment results of AttVCDN on three Twitter datasets. By comparing the data in the table, it can be found that although methods such as BotMOE, KNN, Randomforest, and XGBoost can effectively detect social media robots to a certain extent, their accuracy and F1-score are lower than the detection method using VCDN. This further proves the effectiveness and superiority of the VCDN model in multi-modal data classification problems. Secondly, by comparing the accuracy and F1-score of AttVCDN shown in the table with the results of the VCDN model, it can be found that the performance of AttVCND has also been improved to a certain extent after adding the attention mechanism. Therefore, after introducing the attention mechanism, the AttVCDN model can better capture the importance differences between different modal data and further optimize the fusion process of multi-modal data. The experimental results show that the attention mechanism can not only improve the model's perception ability of different modal information but also help the model focus more effectively on key information, thus improving the overall detection performance.
[0160] Currently, although both feature fusion-based methods and classification result integration methods have achieved good results in social media robot detection, how to effectively combine multi-modal features, high-order correlations, and the importance of different modal information to classification results is still an unsolved problem. The present invention proposes the BotAttVCDN framework, the core of which is to enhance VCDN using the attention mechanism to better capture the importance differences and high-order correlations between different modalities. Experiments on three Twitter datasets prove that BotAttVCDN has a significant performance improvement compared to 11 other benchmark models, and its accuracy and F1-score are better than those of other benchmark models.
[0161] Example 7
[0162] Based on the above embodiments, as Figure 9 shown, the embodiment of the present invention provides a social media robot detection system based on a multi-view multi-modal network, including:
[0163] A collection module for collecting the text data of users, the user graph relationship type data, and the user attribute features.
[0164] A processing module for inputting text data into a Roberta model to obtain the probability distribution of text sequence data.
[0165] Input the user graph relationship type data into the Roberta model to obtain the probability distribution of the user graph relationship type data.
[0166] Input the user attribute features into an MLP model to obtain the probability distribution of the user attribute features.
[0167] A classification module for inputting the probability distribution of text sequence data, the probability distribution of user graph relationship type data, and the probability distribution of user attribute features into a multi-head attention multi-view fusion model to obtain a classification result and complete social media robot detection.
[0168] It should be noted that the social media robot detection system based on a multi-view multi-modal network provided in the embodiments of the present invention is to implement the above-mentioned social media robot detection method based on a multi-view multi-modal network. Its functions can be specifically referred to the above-mentioned method embodiments and will not be elaborated here.
[0169] Embodiment 8
[0170] Based on the above embodiments, the embodiments of the present invention further provide a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the social media robot detection method based on a multi-view multi-modal network in the above embodiments.
[0171] The present invention also provides a computer-readable storage medium, where the storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the social media robot detection method based on a multi-view multi-modal network in the above embodiments.
[0172] In summary, the present invention extracts features and conducts preliminary classification prediction through three different models, and then detects and identifies social media robots according to the multi-head attention multi-view fusion model, improving the detection accuracy. The present invention can better capture the high-order correlations between different data types and reduce the possible biases in the classification results. The present invention first introduces VCDN into the social media robot detection task, and the results show that VCDN is effective in improving the model performance. To further improve the model performance, the present invention optimizes the VCDN module and adds a multi-head attention mechanism to learn the influence of different modality data on the detection results. The final experimental results show that after adding the attention mechanism, both the detection accuracy and F1-score are further improved, and the interpretability of the model is also significantly enhanced. By introducing the attention mechanism, AttVCDN can not only capture and utilize the relationships between different modality views, but also accurately identify the influence of different modality data on the classification results, making the decision-making process of the model more transparent and interpretable.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A social media robot detection method based on multi-view multimodal network, characterized in that: include: Step 1: Collect user text data, user graph relationship type data and user attribute characteristics; Step 2: Input the text data into the Roberta model to obtain the probability distribution of the text sequence data; Input the user graph relationship type data into the Roberta model to obtain the probability distribution of the user graph relationship type data; Input the user attribute features into the MLP model to obtain the probability distribution of the user attribute features; Step 3: Input the probability distribution of text sequence data, the probability distribution of user graph relationship type data, and the probability distribution of user attribute features into the multi-head attention multi-view fusion model to obtain the classification results and complete social media robot detection.
2. The social media robot detection method based on multi-view multi-modal network according to claim 1 is characterized in that: The text data includes user description features and tweet features; The user graph relationship type data includes description features, tweet features, numerical attributes and category attributes; The user attribute features include numerical features and Boolean features.
3. The social media robot detection method based on multi-view multi-modal network according to claim 2 is characterized in that: The Roberta model is expressed by the following formula: X combined =concat(W des X des +W tweet X tweet ) Z Roberta =Softmax(W output ·Dropout(σ(W input ·X combined ))) Among them, X combined is the splicing feature, X des To describe the characteristics, X tweet is the tweet feature, concat is the connection function, Z Roberta is the probability distribution of text sequence data, W des is the linear transformation matrix describing the features, W tweet is the linear transformation matrix of tweet features, Softmax is the normalized exponential function, W output is the linear transformation matrix, Dropout is the random inactivation, σ is the activation function, W input is the linear transformation matrix of the splicing features.
4. The social media robot detection method based on multi-view multi-modal network according to claim 1 is characterized in that: The BotRGCN model is expressed as follows: WITH BotRGCN =Softmax(W out 2 ·σ(W out 1 ·H final )) Among them, Z BotRGCN is the probability distribution of user graph relationship type data, Softmax is the normalized exponential function, W out 2 is the output layer weight, σ is the activation function, W out 1 is the output layer weight, H final It is the node representation after the relational graph convolution network and residual connection.
5. The social media robot detection method based on multi-view multi-modal network according to claim 1 is characterized in that: The multi-head attention multi-view fusion model includes a multi-head attention mechanism and a VCDN model; The multi-head attention mechanism is used to process the spliced three-dimensional tensor after the probability distribution of the text sequence data, the probability distribution of the user graph relationship type data and the probability distribution of the user attribute features are spliced to obtain the attention-weighted probability distribution of the text sequence data, the attention-weighted probability distribution of the user graph relationship type data and the attention-weighted probability distribution of the user attribute features; The VCDN model is used for the attention-weighted probability distribution of text sequence data, the attention-weighted probability distribution of user graph relationship type data, and the attention-weighted probability distribution of user attribute features to obtain classification results and complete social media robot detection.
6. A social media robot detection system based on multi-view multimodal networks, characterized by: include: A collection module, used to collect user text data, user graph relationship type data and user attribute characteristics; A processing module is used to input text data into the Roberta model to obtain the probability distribution of text sequence data; Input the user graph relationship type data into the Roberta model to obtain the probability distribution of the user graph relationship type data; Input the user attribute features into the MLP model to obtain the probability distribution of the user attribute features; The classification module is used to input the probability distribution of text sequence data, the probability distribution of user graph relationship type data and the probability distribution of user attribute features into the multi-head attention multi-view fusion model to obtain the classification results and complete social media robot detection.
7. An electronic device, characterized in that: The invention comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the social media robot detection method based on a multi-view multimodal network as claimed in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the social media robot detection method based on a multi-view multi-modal network as described in any one of claims 1 to 5.