Multi-scale feature fusion robot water army detection method based on comparative learning

Through the large language model extracting multi-scale features and combining comparative learning methods, the problem of insufficient accuracy and robustness of robot naval army detection in the existing technology is solved, and more efficient feature fusion and model robustness are achieved.

CN120046113APending Publication Date: 2025-05-27HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209384.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the process of identifying and detecting robot naval forces, it is difficult for the prior art to effectively capture potential semantic patterns and differences in user behavior, and the model is not robust enough, resulting in low detection accuracy.

Method used

A large language model is used to help extract multi-scale features, combine contrast learning methods to generate augmented samples, fuse multi-scale features through channel attention mechanism, and introduce contrast learning in heterogeneous graph neural network model to improve the robustness of the model.

Benefits of technology

Through multi-scale feature fusion and comparative learning, the accuracy and robustness of robot water army detection have been significantly improved, and achieved good performance on widely used data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046113A_ABST
    Figure CN120046113A_ABST
Patent Text Reader

Abstract

Under limited hardware resources, the detection accuracy can be effectively improved, the robustness of the model is improved, and a detection result indicating whether the current user is a water army user or a real user is generated. The invention relates to a multi-scale feature fusion robot water army detection method based on comparative learning. The method comprises the steps of user feature extraction, multi-scale feature fusion, comparative learning optimization and heterogeneous graph neural network classification decision. Firstly, difference data of real users and water army users on a social media platform are collected, and preliminary user features are obtained; thirdly, weight distribution is carried out on the features of different scales through a channel attention mechanism; then, designing positive and negative sample pairs, and enhancing the capturing capability of the model on fine feature differences by using comparative learning; and finally, training a heterogeneous graph neural network model, inputting a heterogeneous graph structure, and outputting a detection result that the user is a water army user or a real-time user, thereby effectively improving the detection accuracy and the model robustness, and accurately judging the identity of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-scale feature fusion robot water army detection method based on contrast learning, belonging to the field of natural language processing. Background Art

[0002] The birth of large language models (such as GPT, ChatGLM, etc.) has shown excellent performance in various NLP tasks (such as intelligent question answering, machine translation, etc.), and can help extract deep semantic features of text and conduct comprehensive analysis in combination with context information. In the field of water army detection, large language models can help identify users with abnormal features by capturing the potential semantic patterns, sentiment tendencies, and behavioral characteristics of user-posted content. In addition, large language models can also provide effective feature assistance for water army detection by analyzing the language similarity, content generation patterns among water army groups, and the differences from normal users.

[0003] Contrast learning is a self-supervised learning method, and its core idea is to learn the feature representation of data by constructing positive and negative sample pairs, enabling the model to better capture the similarities and differences between data. Different from traditional supervised learning that requires manually labeled data, contrast learning does not require explicit labels and automatically generates learning signals by designing contrastive tasks.

[0004] Graph Neural Network (GNN) is a neural network model that can process graph-structured data, and its core design lies in using the topological structure contained in the nodes and edges of the graph to learn effective representations. The core idea of GNN is message passing. In each layer of the model, a node collects information from its neighboring nodes and aggregates this information. The aggregated information is combined with its own features, and the representation of the node is updated through the non-linear transformation of the neural network. This process can be recursively performed multiple times, enabling each node to finally capture the global structure information within a certain range.

[0005] Therefore, in the present invention, we focus on using large language models to help extract effective representations, contrast learning for sample augmentation to help learn the differences between positive and negative samples, and training a heterogeneous graph neural network model to learn effective representations. We propose a multi-scale feature fusion robot water army detection method based on contrast learning. This detection method not only cleverly fuses multiple effective features, but also enhances the robustness of the model, improves the accuracy of detection, and achieves good performance on widely used datasets. Summary of the Invention

[0006] The present invention designs a multi-scale feature fusion robot water army detection method based on contrastive learning. The present invention can, under limited hardware resources, use multi-scale fusion features to help the heterogeneous graph neural network model better learn the differences between real users and water army users, generate augmented samples through contrastive learning, and enhance the robustness of the model. Compared with traditional detection methods, a large language model is used to assist in feature extraction, and the feature hierarchy is richer. At the same time, using the channel attention mechanism to fuse multi-scale features is superior to the method of directly splicing multiple features, improving the detection performance. Introducing the contrastive learning method at the feature level enhances the robustness of the heterogeneous graph neural network detection model.

[0007] A multi-scale feature fusion robot water army detection method based on contrastive learning includes the following steps:

[0008] Step S101: User feature extraction, specifically including the extraction steps of the following multiple different features;

[0009] Step S1011: Extraction of descriptive features and tweet features. The descriptive information and tweet information of a user on the corresponding social media platform usually include content such as personal interests, professional backgrounds, and opinion expressions, which can reflect the user's personality characteristics and behavioral tendencies. Extract the user's descriptive information and tweet information, and use RoBERTa as a language model to encode the text content to generate high-quality semantic feature representations.

[0010] Step S1012: Category feature extraction. Convert information such as whether the user's account is protected, whether it is authenticated, and whether the background image is changed into boolean values of true or false, use one-hot encoding, cascade and transform them with a fully connected layer and leaky-relu to obtain the user category features;

[0011] Step S1013: Numerical feature extraction. Extract numerical information such as the number of followers and the number of being followed of the user account, perform z-score normalization, and use a fully connected layer to obtain the user numerical features;

[0012] Step S1014: Sentiment feature extraction. The tweet content of a user usually shows different sentiment polarities. Select a training model for sentiment analysis to obtain the user's sentiment features;

[0013] Step S1015: Domain distribution feature extraction. Use a large language model to classify the user's tweets into five domains, count the tweet domain distribution, and use the KL divergence to obtain the domain distribution features of each user.

[0014] Step S1016: Graph feature extraction. Regard the follow and be followed relationships between users as edge relationships, establish an index of the edges between users, and define the types of edges.

[0015] Step S102: Use the channel attention mechanism to weight the user's features of multiple different scales to obtain the fused features of the user;

[0016] Step S103: Incorporate contrastive learning at the feature level, design positive and negative samples, and construct a loss function;

[0017] Step S104: Train the heterogeneous graph neural network classifier to complete the detection of the user.

[0018] For further improvement, the specific steps of step 1011 are as follows:

[0019] The user's description information usually has only one piece, which may include useful information such as personal interests, professional backgrounds, and viewpoints. There are differences between spammer users and real users. Let represent the user's description information consisting of L words. Use the pre-trained RoBERTa to encode the user's description. First, use RoBERTa to transform the words in the user's description:

[0020]

[0021] In the formula: represents the representation of the user's description, and D s is the embedding dimension of RoBERTa. Derive the representation vector of the user's description:

[0022]

[0023] where W D and b D are learnable parameters, is the activation function leaky–relu, and D is the embedding dimension of the user's features.

[0024] A user usually contains multiple tweets. Let represent the user's M tweets, where each tweet contains Q words. Use RoBERTa to encode the user's tweets in a similar way to the description information, and average the representations of all the user's tweets to obtain the final representation of the user's tweets.

[0025] For further improvement, the specific steps of step 1012 are as follows:

[0026] The user's account information contains the following categorical attributes, and the data type represents the current storage form of the data. If privacy settings are enabled, the boolean value is true, otherwise it is false, and the same applies to the representation of other information. Water army users are not as vigilant as real users in setting up account information, and they are not as distinct as real users in personalizing their accounts. Therefore, it can be inferred that there are differences between water army users and real users in categorical attributes.

[0027] The categorical attributes are composed of the user's account information, denoted by P cat Represent the categorical attributes, use one-hot encoding, cascade and transform them with the fully connected layer and leaky-relu to obtain the user categorical features

[0028] For further improvement, the specific steps of step 1013 are as follows:

[0029] In the user's account information, use P num To represent numerical attributes, extract numerical information such as the number of followers and the number of follows in the user's account, perform z-score standardization, and use the fully connected layer to obtain the user numerical features

[0030] For further improvement, the specific steps of step 1014 are as follows:

[0031] When users post tweets, water army users tend to post a large number of purposeful tweets in a short period of time, and the tweet sentiment will be more distinct and variable. Real users are more life-like, and the change in tweet sentiment will be smoother. Therefore, the sentiment polarity of tweets is also a feature worthy of study. Before extracting the user sentiment features, it is first necessary to clean and process the user's tweet data. Specifically, the behavior of tagging other users in the user's tweets (such as "@username") is uniformly converted to @USER to avoid individual information interfering with sentiment analysis. The hashtag in the tweet (such as "#topic") should be converted to #HASHTAG to ensure that the sentiment analysis focuses on the sentiment content rather than the specific topic. All links (such as "http: / / ..." or "www...") should also be uniformly identified as HTTPURL to avoid their impact on sentiment analysis. In addition, delete the emojis in the tweets to ensure that the sentiment polarity analysis is based on pure text data and avoid the interference of emojis on the results.

[0032] Then, select the model training dataset, and select the Sentimen140 dataset. 10,000 pieces of data were randomly selected from it, and the data was also divided using the stratified sampling method. Finally, a reasonable neural network structure was designed as the training model, one-hot encoding was used to record the sentiment polarity results of user tweets, and the sentiment features of users were constructed.

[0033] For further improvement, the specific steps of step 1015 are as follows:

[0034] The fields involved in the tweets posted by users are diverse. Sock puppet users are more inclined to post a large number of tweets with similar content in a short period of time to achieve the purpose of guiding public opinion, while real users are more inclined to occasionally express their views or share their daily lives with friends. Therefore, there are also differences between real users and sock puppet users in terms of field distribution. Users' tweets can be divided into the following five categories according to their content: "Politics", "Business", "Entertainment", "Sports", "Others". First, consider using a large language model to classify tweet content, design a Prompt, and let the large language model generate a response based on the given prompt. Then, count the fields based on the response results of the large language model, and finally use the KL divergence to evaluate the similarity between the reference distribution and the current distribution, and add the similarity as a new feature to the numerical features.

[0035] For further improvement, the specific steps of step 1016 are as follows:

[0036] Based on the data of the user's follow and being followed lists, construct the user index id = [id 1 , id 2 , id 3 …id n , where id n represents the number of the nth user, n represents the total number of user numbers, and save the edge index between users as edge index = [[id 1 , id 2 ,[id 2 , id 3 …[id m , id n . If there is a follow relationship between users, then edge type = 0, and if there is a being followed relationship, then edge type = 1.

[0037] For further improvement, the specific steps of step 102 are as follows:

[0038] Design a class for weight allocation to complete the fusion of user multi-scale features. Adopt the channel attention mechanism, regard the user multi-scale features obtained in step S101 as different channels, and allocate weights to them according to the importance of each channel.

[0039] First, the user's multi-scale features contain information of multiple scales. First, initialize the user features as

[0040]

[0041] where W I and b I are learnable parameters, is the activation function leaky–relu. Each scale can be regarded as an independent channel, and the feature dimension of the user can be expressed as [c, h, w]. Then, by introducing the channel attention mechanism, first perform pooling on the user's multi-scale features x ′ = maxpooling(x), calculate an attention weight for each channel. At this time, the feature dimension of x ′ is [c, 1, 1]. The globally pooled vector passes through the MLP network to obtain the weight of each channel weight = MLP(x ′ ). The weight represents the influence of each channel on feature extraction and reflects the importance of this channel in subsequent detection tasks. Finally, multiply the obtained weight by the user's initial feature vector to obtain the weighted feature vector x weighted = x · weight. By weighting the features of each channel through the channel attention mechanism, the subsequent neural network classification model can focus on the features that make greater contributions to the detection task.

[0042] For further improvement, the specific steps of step 103 are as follows:

[0043] First, design a class that implements the self-supervised contrast learning method. For each sample in the subset of the training data in each iteration, generate a corrupted version using the marginal distribution of the user features as the positive sample. To this end, we extract some random features of each sample in the training data and replace each feature with a feature randomly selected from another sample in the dataset.

[0044] In a mini-batch of size N, let i ∈ I = {1, 2, 3.., N} be the index of the sample. Assume that each sample x i has a corresponding sample generated by augmentation Define the loss function as follows:

[0045]

[0046] where z iand z j is composed of x i and obtained through Encorder. τ is the temperature parameter that controls the scaling of similarity. The positive samples are user features generated through the margin loss, and the negative samples are user features without the margin loss. The contrastive loss prompts the model to pull the positive samples closer and push the negative samples farther apart in the embedding space, maximizing the similarity between similar features and minimizing the similarity between different category features, improving the model's ability to capture subtle feature differences and enhancing the robustness of detection.

[0047] For further improvement, the specific steps of step 104 are as follows:

[0048] In step S1014, a reasonable network structure is designed and the graph neural network model is trained, expecting that the graph neural network model can learn the differences between real users and spammer users from publicly available datasets widely used in spammer detection, and obtain experimental results by dividing the data on the widely used datasets.

[0049] In the training task, we regard it as a binary classification problem. The so-called binary classification task is to input the feature vectors of multiple nodes obtained in step S102 and the edge relationships obtained in S1016 into the neural network model to aggregate the features of neighbor nodes, and obtain the probabilities of classifying as spammer users and real users.

[0050] First, define the graph attention network SimpleHGN and initialize it. Create a module list containing SimpleHGNConv layers, and the number of network layers is num_layers. In the model training stage, initialize the model and the optimizer. Define the optimizer optimizer as Adam with a learning rate of learning_rate; define the loss function loss_fn as the cross-entropy loss CrossEntropyLoss. Then, define the number of iterations as epochs and the size of the input batch of training data as batch_size. In the first layer of the network, map the input training set sample features to the hidden layer, the dimension of the input features is in_channels, and the dimension of the hidden layer features is hidden_channels; the middle layer maps the hidden layer features to the next hidden layer; the last layer maps the hidden layer features to the output features, the dimension of the output features is out_channels, and the Dropout rate is dropout. In the forward propagation function of the model, its input is the node feature x weighted and the edge index edge index , traverse all convolutional layers, and apply the softmax layer in the last layer for probability prediction.

[0051] After the heterogeneous graph neural network model is trained, by inputting the node features and edge relationships of the user, the detection result of this user can be obtained to detect whether this user is a water army user or a real user.

[0052] Compared with other methods, the present invention has the following remarkable advantages:

[0053] The present invention uses a large language model to help extract effective features, and a channel attention mechanism to help perform multi-scale feature fusion. By introducing a contrast learning method into the heterogeneous graph neural network model, it has the following advantages: 1. Accuracy. By extracting features that are different between real users and water army users from the user's account information and using the channel attention mechanism for weighted processing, it helps the heterogeneous graph neural network detection model better distinguish real users and water army users, effectively improving the detection accuracy; 2. Robustness. By performing contrast learning at the feature level and using the marginal distribution of user features to generate a corrupted version as a positive sample, the contrast loss can maximize the similarity between similar features, effectively improving the robustness of the detection model. Description of the Drawings

[0054] Figure 1 Shows the overall flowchart of the "Robot Water Army Detection Method Based on Contrast Learning and Multi-scale Feature Fusion" of the present invention.

[0055] Figure 2 Shows the comparison chart of the robustness experiment of the "Robot Water Army Detection Method Based on Contrast Learning and Multi-scale Feature Fusion" of the present invention.

[0056] Figure 3 Shows the training loss curve of the heterogeneous graph neural network model of the "Robot Water Army Detection Method Based on Contrast Learning and Multi-scale Feature Fusion" of the present invention.

[0057] Figure 4 Shows the validation loss curve of the heterogeneous graph neural network model of the "Robot Water Army Detection Method Based on Contrast Learning and Multi-scale Feature Fusion" of the present invention. Detailed Embodiment

[0058] To more clearly show the features and advantages of this patent, the following provides a detailed description of the embodiments. It should be clear that the following detailed description is only for example and is intended to provide further explanation for this application. Unless otherwise stated, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. Figure 1 Illustrates a robot water army detection method based on contrast learning and multi-scale feature fusion provided by an embodiment of this application. The method includes the following steps:

[0059] Step 1: User Feature Extraction Step

[0060] In the user feature extraction step, we extract data from three datasets widely used in robot water army detection, namely: Twibot20, Cresci-15, and Twibot22. Since the Twibot22 dataset is too large, we adopt a random sampling method to construct a sample set for subsequent experiments under limited computing power.

[0061] Regarding the description information of users, it may include useful information such as personal interests, professional backgrounds, and opinion expressions. There are differences between water army users and real users. Let \(u\) represent the user description information consisting of \(L\) words. The pre-trained RoBERTa is used to encode the user description. First, RoBERTa is used to transform the words in the user description:

[0062]

[0063] In the formula: \(h_{u}\) represents the representation of the user description, and \(D\) s is the embedding dimension of RoBERTa, which is 768. Therefore, when the number of users in the dataset is USER_NUM, the dimension of the description feature is \([USER_NUM, 768]\). The representation vector of the user description is derived:

[0064]

[0065] where \(W\) D and \(b\) D are learnable parameters, \(\sigma\) is the activation function leaky–relu, and \(D\) is the embedding dimension of the user feature, which is 128.

[0066] Regarding the tweet information of users, a user usually contains multiple tweets. Let \(t_{i}\) represent the \(M\) tweets of the user, where each tweet contains \(Q\) words. RoBERTa is used to encode the user's tweets in a similar way to the description information, and the representations of all users' tweets are averaged to obtain the final representation of the user's tweets. Similar to the user's description information, when the number of users in the dataset is USER_NUM, the dimension of the tweet feature is \([USER_NUM, 768]\).

[0067] Regarding the user's category attributes, the user's account information includes privacy settings, etc. The data type represents the storage format of the current data. If privacy settings are made, the Boolean value is true, otherwise it is false, and the same is true for other information. Water army users are not as vigilant as real users in setting account information, and their personalized account appearance is not as distinct as real users. Therefore, it can be inferred that there are differences in category attributes between water army users and real users. cat Represents category attributes, uses one-hot encoding, cascades and transforms them with the fully connected layer and leaky-relu to obtain user category features When the number of users in the data set is USER_NUM, the dimension of the category attribute feature is [USER_NUM,5].

[0068] For the user's numerical attributes, the user's account information is represented by P num Represents numerical attributes, extracts numerical information such as the number of followers and the number of followers of the user account, performs z-score standardization, and uses a fully connected layer to obtain user numerical features When the number of users in the dataset is USER_NUM, the dimension of the category attribute feature is [USER_NUM,7]. Regarding the user's sentiment polarity, when users post tweets, water army users tend to post a large number of purposeful tweets in a short period of time, and the sentiment of the tweets will be more distinct and changeable, while real users are more lifelike and the change of sentiment in tweets will be more gradual. Therefore, the sentiment polarity of tweets is also a feature worth studying. Based on the number of user tweets in different datasets, we select the number of tweets with the largest number of tweets among users as the number of tweets for sentiment polarity analysis.

[0069] Before extracting user sentiment features, the user's tweet data must first be cleaned and processed. Specifically, the behavior of marking other users in the user's tweet (such as "@username") is uniformly converted to @USER to avoid individual information interfering with sentiment analysis. Topic tags in tweets (such as "#topic") should be converted to #HASHTAG to ensure that sentiment analysis focuses on sentiment content rather than specific topics. All links (such as "http: / / ..." or "www...") should also be uniformly marked as HTTPURL to avoid their impact on sentiment analysis. In addition, emoticons in tweets are deleted to ensure that sentiment polarity analysis is based on plain text data to avoid interference from emoticons in the results.

[0070] Then, the Sentimen140 dataset was used. The Sentiment140 dataset contains 1,600,000 tweets scraped from Twitter, and each tweet has its corresponding sentiment polarity (0 = negative, 4 = positive). 10,000 pieces of data were randomly selected from it, and the data was also divided using the stratified sampling method, with 80% of the data selected as the training set. Finally, RoBERTa + BiLSTM + ATTENTION was selected as the training model for sentiment classification, one-hot encoding was used, and the sentiment polarity results of the user's tweets were recorded to form the user's sentiment features When the number of tweets is N and the number of users is USER_NUM, the dimension of the sentiment polarity feature is [USER_NUM, N].

[0071] The fields involved in the tweets posted by users are diverse. Sockpuppet users are more inclined to post a large number of tweets with similar content in a short period of time to achieve the purpose of guiding public opinion, while real users are more inclined to occasionally express their views or share their daily lives with friends. Therefore, there are also differences between real users and sockpuppet users in terms of field distribution. Users' tweets can be classified into the following five categories according to their content: "Politics", "Business", "Entertainment", "Sports", "Others". First, consider using a large language model to classify the tweet content, design a Prompt, and let the large language model generate a response based on the given prompt. Then, count the fields based on the response results of the large language model, and finally use the KL divergence to evaluate the similarity between the reference distribution and the current distribution, and add the similarity as a new feature to the numerical features When the number of users in the dataset is USER_NUM, the dimension of the new numerical feature is [USER_NUM, 8].

[0072] According to the data in the user's follow and follower lists, that is, extract the user's id from the sequences after "following" and "follower", and construct the user index id = [id 1 , id 2 , id 3 … id n , where id n represents the number of the nth user, n represents the total number of user numbers, and save the edge index between users as edge index = [[id 1 , id 2 , [id 2 , id 3 … [id m , id n , if there is a follow relationship between users, then edgetype = 0. If there is a relationship of being concerned about, then edge type = 1. Complete the multi-scale feature extraction of the user;

[0073] Step 2: Multi-scale feature fusion step

[0074] We design a class for weight allocation to complete the fusion of the user's multi-scale features. Adopting the channel attention mechanism, regard the user's multi-scale features obtained in step S101 as different channels, and assign weights to them according to the importance of each channel.

[0075] First of all, the user's multi-scale features contain information of multiple scales. First, initialize the user features as

[0076]

[0077] where W I and b I are learnable parameters, is the activation function leaky–relu. Each scale can be regarded as an independent channel. Then, by introducing the channel attention mechanism and setting the initial parameters according to the user's feature dimension, the user's feature dimension can be expressed as [c, h, w]. First, perform pooling on the user's multi-scale features x ′ = maxpooling(x) to calculate an attention weight for each channel. At this time, the feature dimension of x ′ is [c, 1, 1]. The globally pooled vector passes through the MLP network to obtain the weight of each channel weight = MLP(x ′ ). The weight represents the influence of each channel on feature extraction and reflects the importance of this channel in subsequent detection tasks. Finally, multiply the obtained weight by the user's initial feature vector to obtain the weighted feature vector x weighted = x · weight. By weighting the features of each channel through the channel attention mechanism, the subsequent neural network classification model can focus on the features that make greater contributions to the detection task.

[0078] Step 3: Self-supervised contrastive learning step

[0079] First of all, design a class that implements the self-supervised contrastive learning method. For each sample in the subset of the training data in each iteration, generate a corrupted version as the positive sample using the marginal distribution of the user features weighted by the channel attention mechanism. The selected encoder is SimpleHGN.

[0080] To this end, we extract some random features of each sample in the training data and replace each feature with the feature of another sample randomly selected from the dataset. In a mini-batch of size N, let i ∈ I = {1, 2, 3.., N} be the index of the sample, assuming each sample x i has a corresponding sample generated through augmentation Define the loss function as follows:

[0081]

[0082] where z i and z j are obtained from x i and through the Encorder. τ is the temperature parameter that controls the scaling of similarity. The positive samples are the user features generated through the margin loss, and the negative samples are the user features without the margin loss. The contrastive loss prompts the model to pull the positive samples closer in the embedding space while pulling the negative samples apart, maximizing the similarity between similar features and minimizing the similarity between different category features, improving the model's ability to capture subtle feature differences and achieving better robustness. Select a method with better performance on the Twibot20 dataset as a comparison, mask the features by 10%-90%, and observe the results of the model in terms of accuracy and F1 score. It can be seen from the experimental results that our method still shows good performance when the masking rate is 10%-90% compared with the method with better performance on the Twibot20 dataset, and is superior to the comparison method in both the accuracy and F1 score metrics, proving that this method has good robustness. Moreover, it can be seen that the lower the masking rate, the more complete the features available for detection, and accordingly, the accuracy and F1 score of the detection will also show an upward trend, and the experimental results of the robustness meet the expectations.

[0083] Step 4: Training steps of the heterogeneous graph neural network model

[0084] In step S1014, design a reasonable network structure and train the graph neural network model, expecting that the graph neural network model can learn the differences between real users and spammer users from publicly available datasets widely used in spammer detection. Perform data splitting on three widely used datasets, and divide them into a training set, a validation set, and a test set according to 7:2:1.

[0085] In the training task, we treat it as a binary classification problem. A binary classification task means that based on the feature vectors of multiple nodes obtained in step S102 and the edge relationships obtained in S1016, they are input into a neural network model to aggregate the features of neighboring nodes, and the probabilities of classifying as spammer users and real users are obtained. First, define the Graph Attention Network SimpleHGN and initialize it. Create a module list containing SimpleHGNConv layers, and the number of network layers is num_layers. In the model training stage, initialize the model and the optimizer. Define the optimizer optimizer as Adam with a learning rate of learning_rate; define the loss function loss_fn as CrossEntropyLoss. Then, define the number of iterations as epochs and the size of the input batch training data as batch_size. In the first layer of the network, map the input training set sample features to the hidden layer. The dimension of the input features is in_channels, and the dimension of the hidden layer features is hidden_channels; the middle layer maps the hidden layer features to the next hidden layer; the last layer maps the hidden layer features to the output features. The dimension of the output features is out_channels, and the Dropout rate is dropout. In the forward propagation function of the model, its input is the node features x weighted and the edge index edge index , traverse all convolutional layers, and apply a softmax layer in the last layer for probability prediction.

[0086] During the training process, for different datasets, we control the training parameters of the heterogeneous graph neural network model as shown in Table 1:

[0087] Table 1 Training parameter table of the heterogeneous graph neural network model

[0088]

[0089] After the heterogeneous graph neural network model is trained, by inputting the node features and edge relationships of the user, the detection result of this user can be obtained, and it can be detected whether this user is a spammer user or a real user. Taking the Twibot20 dataset as an example, the training loss and validation loss of this heterogeneous graph neural network model during the training process are given. In fact, the loss continuously iteratively optimized by the model is the sum of the loss of contrastive learning and the loss between the true label and the predicted label, Figure 3 and Figure 4 shows that the model converges on both the training set and the validation set.

[0090] In the specific experimental verification process, we verified the performance on the publicly available dataset widely used for water army detection, selected relatively advanced existing methods for comparison, and in terms of evaluation metrics, we used three metrics, namely accuracy, F1-score, and Mcc, to evaluate the method proposed in the present invention. The experimental results are shown in Table 2:

[0091] Table 2 Experimental results of different comparison methods on the dataset

[0092]

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-scale feature fusion robot water army detection method based on contrastive learning, characterized in that: The following steps are involved: Step S101: User feature extraction, specifically including the following steps of extracting multiple different features; Step S1011: Extracting description features and tweet features. The description information and tweet information on the user's corresponding social media platform usually include personal interests, professional background, and opinions, which can reflect the user's personality characteristics and behavioral tendencies. The user's description information and tweet information are extracted, and RoBERTa is used as a language model to encode the text content to generate high-quality semantic feature representation; Step S1012: Category feature extraction, converting information such as whether the user's account is protected, authenticated, and whether the background image is changed into true or false Boolean values, using one-hot encoding, cascading and transforming them with the fully connected layer and leaky-relu to obtain user category features; Step S1013: extracting numerical features, extracting numerical information such as the number of followers and the number of followers of the user account, performing z-score standardization, and obtaining user numerical features using a fully connected layer; Step S1014: Extracting sentiment features. The tweets of users usually show different sentiment polarities. A training model for sentiment analysis is selected to obtain the sentiment features of the users. Step S1015: extracting domain distribution features, using a large language model to classify user tweets into five domains, counting the domain distribution of tweets, and using KL divergence to obtain the domain distribution features of each user; Step S1016: extract graph features, treat the following and followed relationships between users as edge relationships, establish edge indexes between users, and define edge types; Step S102: using the channel attention mechanism, weighting the features of multiple user scales to obtain the user's fusion features; Step S103: adding contrastive learning at the feature level, designing positive and negative samples, and constructing a loss function; Step S104: Train the heterogeneous graph neural network classifier to complete the detection of users.

2. According to claim 1, a multi-scale feature fusion robot navy detection method based on contrastive learning is characterized in that: The specific steps of step S1011 are as follows: There is usually only one description of a user, which may include useful information such as personal interests, professional background, and opinions. There are differences between water army users and real users. Represents user description information consisting of L words. The pre-trained RoBERTa is used to encode the user description. First, RoBERTa is used to convert the words in the user description: In the formula: Represents the representation of the user description, D s is the embedding dimension of RoBERTa, and the representation vector of the user description is derived: Where W D and b D is a learnable parameter, is the activation function leaky-relu, D is the embedding dimension of user features, and users usually contain multiple tweets. represents a user's M tweets, each of which contains Q words, Use RoBERTa to encode the user's tweets similar to the description information, average the representations of all users' tweets, and get the representation of the final user's tweets 3. According to the multi-scale feature fusion robot navy detection method based on contrastive learning in claim 1, it is characterized in that: The specific steps of step S1012 are as follows: The user's account information contains the following category attributes. The data type represents the storage format of the current data. If privacy is set, the Boolean value is true, otherwise it is false. The same is true for other information. Water army users do not have a strong sense of prevention in setting account information, and their personalized account dressing is not as distinct as that of real users. Therefore, it can be inferred that there are differences in category attributes between water army users and real users. The category attributes are composed of account information and P is used. cat Represents category attributes, uses one-hot encoding, concatenates and transforms them with the fully connected layer and leaky-relu to obtain user category features 4. According to the multi-scale feature fusion robot navy detection method based on contrastive learning in claim 1, it is characterized in that: The specific steps of step S1013 are as follows: The user's account information is stored in P num Represents a numerical attribute. The numerical attributes selected in the design are as follows: Extract the numerical information in the table above, such as the number of fans and the number of followers of the user account, perform z-score standardization, and use the fully connected layer to obtain the user's numerical features 5. The method for detecting robot navy using multi-scale feature fusion based on contrastive learning according to claim 1, characterized in that: The specific steps of step S1014 are as follows: When users post tweets, water army users tend to post a large number of purposeful tweets in a short period of time, and the emotions of tweets will be more distinct and changeable. Real users are more lifelike, and the changes in the emotions of tweets will be more gradual. Therefore, the emotional polarity of tweets is also a feature worthy of study. Before extracting user emotional features, the user's tweet data needs to be cleaned and processed first. Specifically, the behavior of marking other users in the user's tweets (such as "@username") is uniformly converted to @USER to avoid individual information interfering with sentiment analysis. The topic tags in the tweets (such as "#topic") should be converted to #HASHTAG to ensure that the sentiment analysis focuses on the emotional content rather than the specific topic. All links (such as "http: / / ..." or "www...") should also be uniformly marked as HTTPURL to avoid their influence on sentiment analysis. In addition, the emoticons in the tweets are deleted to ensure that the sentiment polarity analysis is based on pure text data to avoid the interference of emoticons on the results. Then, we select the model training data set, choose the Sentimen140 data set, randomly select 10,000 data from it, and use the stratified sampling method to divide the data; finally, we design a reasonable neural network structure as the training model, use one-hot encoding, record the emotional polarity results of user tweets, and constitute the user's emotional features.

6. The multi-scale feature fusion robot navy detection method based on contrastive learning according to claim 1 is characterized in that: The specific steps of step S1015 are as follows: The fields covered by the tweets posted by users are diverse. Water army users tend to post a large number of tweets with similar content in a short period of time to guide public opinion, while real users tend to occasionally express their opinions or share their daily lives with friends. Therefore, there are differences in field distribution between real users and water army users. Users can be divided into the following five categories according to the content of their tweets: "Politics", "Business", "Entertainment", "Sports", and "Others". First, consider using a large language model to divide the content of tweets and design Prompt to let the large language model generate replies based on the given prompts. Then, count the fields based on the reply results of the large language model. Finally, use KL divergence to evaluate the similarity between the reference distribution and the current distribution, and add the similarity as a new feature to the numerical feature.

7. The method for detecting robot navy using multi-scale feature fusion based on contrastive learning according to claim 1, characterized in that: The specific steps of step S1016 are as follows: According to the data of the user's follow and followed lists, construct the user index id = [id1, id2, id3...id n ], id n Indicates the number of the nth user, n indicates the total number of user numbers, and the edge index between users is stored as edge index =[[id1,id2],[id2,id3]…[id m ,id n ]], if there is a following relationship between users, edge type =0, if there is a concerned relationship, then edge type =1.

8. The method for detecting robot navy using multi-scale feature fusion based on contrastive learning according to claim 1, characterized in that: The specific steps of step S102 are as follows: Design a class for weight assignment to complete the fusion of user multi-scale features. Use the channel attention mechanism to treat the user multi-scale features obtained in step S101 as different channels and assign weights to each channel according to its importance. First, the user multi-scale features contain information at multiple scales. Initialize the user features as Where W I and b I is a learnable parameter, is the activation function leaky–relu, each scale can be regarded as an independent channel, and the user's feature dimension can be expressed as [c, h, w]. By introducing the channel attention mechanism, the user's multi-scale features are first pooled x′=maxpooling(x), and an attention weight is calculated for each channel. At this time, the feature dimension of x′ is [c, 1, 1]. The vector after global pooling passes through the MLP network to obtain the weight of each channel weight=MLP(x′). The weight represents the influence of each channel on feature extraction and reflects the importance of the channel in subsequent detection tasks. The obtained weight is multiplied by the user's initial feature vector to obtain the weighted feature vector x weighted =x·weight, the features of each channel are weighted through the channel attention mechanism, so that the subsequent neural network classification model can focus on the features that contribute more to the detection task.

9. The method for detecting robot navy using multi-scale feature fusion based on contrastive learning according to claim 1, characterized in that: The specific steps of step S103 are as follows: First, we design a class that implements the self-supervised contrastive learning method. For each sample in the subset of the training data in each iteration, we use the marginal distribution of user features to generate a corrupted version as a positive sample. To this end, we extract some random features of each sample in the training data and replace each feature with the feature of another sample randomly selected from the dataset. In a small batch of size N, let i∈I={1,2,3..,N} be the index of the sample, then we can assume that: for each sample x i There is a corresponding sample generated through enhancement The loss function is defined as follows: where z i and z j By x i and Obtained by Encorder, τ is the temperature parameter that controls the scaling of similarity. The positive sample is the user feature generated by edge loss, and the negative sample is the user feature that has not undergone edge loss. The contrast loss prompts the model to bring the positive samples closer in the embedding space and pull the negative samples apart, thereby maximizing the similarity between similar features and minimizing the similarity between features of different categories, improving the model's ability to capture subtle feature differences, and improving the robustness of detection.

10. The method for detecting robot navy using multi-scale feature fusion based on contrastive learning according to claim 1, characterized in that: The specific steps of step S104 are as follows: In step S1014, a reasonable network structure is designed and a graph neural network model is trained. It is expected that the graph neural network model can learn the difference between real users and water army users from a public dataset that is widely used for water army detection. Experimental results are obtained by data partitioning on a widely used dataset. In the training task, we treat it as a binary classification problem. The so-called binary classification task is to aggregate the neighbor node features according to the feature vectors of multiple nodes obtained in step S102 and the edge relationships obtained in S1016, and obtain the probability of classification as water army users and real users. First, define the graph attention network SimpleHGN and initialize it, create a module list containing the SimpleHGNConv layer, and the number of network layers is num_layers. In the model training stage, initialize the model and Optimizer, define optimizer optimizer as Adam, learning rate as learning_rate, and loss function loss_fn as CrossEntropyLoss; then, define the specified number of iterations as epochs and the input batch training data size as batch_size. In the first layer of the network, the input training set sample features are mapped to the hidden layer, the dimension of the input features is in_channels, and the dimension of the hidden layer features is hidden_channels. The middle layer maps the hidden layer features to the next hidden layer; finally, the hidden layer features are mapped to the output features, the dimension of the output features is out_channels, and the Dropout rate is dropout. In the forward propagation function of the model, its input is a network containing node features x. weighted and edge index edge index , traverse all convolutional layers and apply a softmax layer to the last layer for probability prediction; After the heterogeneous graph neural network model training is completed, the user's node features and edge relationships are input to obtain the user's detection results, and the user can be detected as a water army user or a real user.