Social robot detection method based on multi-view evidence fusion

By mining implicit connections and using D-S evidence theory to fuse multi-view evidence, the problem of relying on explicit relationships and ignoring implicit connections in the prior art is solved, and higher social robot detection accuracy and reliability are achieved.

CN119989277APending Publication Date: 2025-05-13ZHENGZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203823.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing graph-based social robot detection methods rely on explicit relationships and ignore implicit connections. The camouflage behavior of social robots violates the assumption of homogeneity, resulting in limited detection accuracy.

Method used

By mining and analyzing implicit connections between users, establish multiple views, and use D-S evidence theory to fuse evidence from multiple views, thereby generating a more comprehensive and representative graph structure, improving the accuracy and reliability of social robot detection.

Benefits of technology

It effectively alleviates the problem of homogeneity assumption, improves the accuracy and reliability of detection, reduces the rate of misjudgment, and proves in the experiment that it is better than other baseline models in terms of indicators such as accuracy, accuracy and F1 score.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989277A_ABST
    Figure CN119989277A_ABST
Patent Text Reader

Abstract

The invention discloses a social robot detection method based on multi-view evidence fusion, and belongs to the technical field of social robot detection.The method comprises the steps that firstly, user features are preprocessed to obtain a user representation vector, and then an encoder combines a variational auto-encoder and a supervised contrast learning strategy to obtain a user representation vector; distinguishing features between the social robot and real users are learned, an implicit connection relation graph between the users is obtained through a decoder, a reconstructed graph containing three different types of edges is established based on the implicit connection relation graph, then multi-view evidence extraction is conducted on the reconstructed graph through a graph attention network and evidence deep learning, and the reconstructed graph is obtained. And finally, fusing evidences of a plurality of views by using a D-S evidence theory to obtain final joint belief quality and uncertainty, and completing social robot detection. The effectiveness and superiority of the method are verified through experiments, and the detection performance of the social robot is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of social robot detection, and specifically relates to a social robot detection method based on multi-view evidence fusion. Background Art

[0002] With the popularity of social networks and the rapid development of artificial intelligence technology, the participation of automated programs (i.e. social robots) in online social activities has increased significantly. These robots sneak into social media platforms through complex imitation strategies and engage in various harmful behaviors such as spreading false information, posing a serious threat to the security of cyberspace. Since ordinary users have difficulties and biases in identifying these malicious accounts, it is particularly important to develop and implement efficient social robot detection tools to expose and curb the improper behavior of social robots in a scientific way to maintain the authenticity of information and the security of cyberspace.

[0003] In recent years, methods based on graph neural networks (GNNs) have shown significant advantages in the field of social robot detection. These methods can simultaneously utilize user characteristics and network structure information and effectively integrate them into user representation. Usually, users are regarded as nodes, and the social relationships between users constitute edges, forming a complex social network graph. Alhosseini et al. were the first to introduce GNNs into this field and verified its advantages over traditional methods. With the deepening of research and the evolution of social robots, traditional GNNs can no longer meet the requirements. In order to meet the challenges, researchers have proposed a variety of improved methods, such as introducing relational graph convolutional neural networks, combining Markov random fields, introducing attention mechanisms, combining contrastive learning, and utilizing the community relationships of users.

[0004] However, existing graph-based detection methods still face several challenges. First, existing methods mainly rely on explicit relationships such as following or forwarding to construct graph structures. While previous studies have shown that combining multiple types of relationships can improve robot detection, implicit connections, such as potential coordinated activities or common interest communities, have been largely overlooked. These implicit connections contain valuable signals that help distinguish social robots from real users. However, the lack of these connections in current models limits the accuracy of detection. Second, the disguised nature of social robots violates the homogeneity assumption. The homogeneity assumption holds that nodes with edges in the graph are more likely to belong to the same category or have similar characteristics. However, for social robots, especially advanced social robots, they usually disguise themselves by stealing the attributes of real users and establishing connections with real users, resulting in limited model performance. Summary of the invention

[0005] The purpose of the present invention is to provide a social robot detection method based on multi-view evidence fusion, which establishes multiple different views by mining and analyzing implicit connections between users, uses DS evidence theory to fuse evidence from multiple views, comprehensively analyzes user relationships and behavior patterns from multiple angles, and improves the accuracy and reliability of social robot detection.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a social robot detection method based on multi-view evidence fusion, comprising the following steps:

[0007] Step 1: preprocessing user features to obtain user representation vectors, wherein the user features include semantic features, classification features, and digital features;

[0008] Step 2: Input the user representation vector into an encoder, and the encoder combines a variational autoencoder and a supervised contrastive learning strategy to learn the distinguishing features between the social robot and the real user;

[0009] Step 3: input the output result of the encoder to the decoder, and the decoder calculates the probability that there is an implicit connection between users according to the distinctive features, and obtains an implicit connection relationship graph containing only implicit connections;

[0010] Step 4: by analyzing the explicit connection edges in the social network and the implicit connection edges of the implicit connection relationship graph, a reconstructed graph containing three different types of edges is generated, wherein the three types of edges include enhanced edges, strong relationship edges, and weak relationship edges. The enhanced edges refer to edges included in the implicit connections but not in the explicit connections, the strong relationship edges refer to edges included in both the implicit connections and the explicit connections, and the weak relationship edges refer to edges included in the explicit connections but not in the implicit connections;

[0011] Step 5: Use graph attention network and evidence deep learning to extract evidence from multiple views of the reconstructed graph, and analyze the belief quality and overall uncertainty of each category based on the extracted evidence vector, so as to obtain evidence from multiple views;

[0012] Step 6: Use the evidence fusion rules of DS evidence theory to fuse the evidence from multiple views to obtain the final joint belief quality and uncertainty to complete social robot detection.

[0013] Furthermore, the implementation process of step 2 includes: firstly, mapping the user representation vector to a latent space, converting the high-dimensional input data into a low-dimensional representation, respectively calculating the mean and variance of the Gaussian distribution of the low-dimensional representation through two multi-layer perceptrons MLP, using the re-parameterization technique to perform random sampling from the standard normal distribution, and calculating the latent variable according to the following formula:

[0014] In the formula, z i represents the potential representation after random sampling of user i, μ i represents the mean, σ i represents the variance, ε represents the value randomly sampled from the standard normal distribution;

[0015] Then, samples with the same label and corresponding VAE sampling results are marked as positive samples, and samples with different labels are marked as negative samples. Combined with supervised contrastive learning, the following contrastive loss function is designed to train the model:

[0016] In the formula, i∈I≡{1...2N} represents the index of the sample after data enhancement, P(i) represents the index set of positive samples, A(i) represents the index set of negative samples of different categories from sample i, τ is the temperature parameter, z i is the feature representation of the i-th sample, z p is with z i The feature representation of samples belonging to the same category, z a is with z i Feature representation of samples belonging to different categories.

[0017] Furthermore, in step 3, the calculation formula for the probability of implicit connection between users is as follows:

[0018] In the formula, p i,j represents the probability of implicit connection between users, sigmoid(·) represents the sigmoid function, z i 、z j They represent the feature representation of the i-th sample and the j-th sample respectively, and T represents the transpose of the matrix;

[0019] If p i,j is greater than the set threshold, then node z i 、z j There is an implicit connection between them.

[0020] Furthermore, in step 4, the calculation methods of the three different types of edges are as follows: E enh =E imp ∪E exp -E exp E str =E imp ∩E exp E weak =E imp ∪E exp -Eimp

[0021] In the formula, E enh represents the set of enhanced edges, E str represents the set of strong relationship edges, E weak represents the set of weak relationship edges, E exp represents the set of explicit connection edges, E imp Represents the set of implicitly connected edges.

[0022] Furthermore, in the step five, a graph attention network and evidence deep learning are used to perform evidence extraction on the reconstructed graph in three views, namely, an enhanced relationship view, a strong relationship view and a weak relationship view. The enhanced relationship view only contains enhanced edges, the strong relationship view only contains strong relationship edges, and the weak relationship view only contains weak relationship edges.

[0023] Furthermore, in step 5, the process of evidence extraction is as follows:

[0024] (1) The importance weight between two nodes under relationship r is calculated through the graph attention network. The calculation process is as follows:

[0025] In the formula, {l} represents the lth layer of the graph neural network, represents the query vector of node i under relation r, represents the key vector of node j under relation r, represents the value vector of node j under relation r, W r,{l} , b r,{l} They represent the learnable parameter matrices of the lth layer under the relation r, Respectively represent the node z of the lth layer under the relationship r i , node z j , N r (i) represents node z under relation r i Neighborhood, d represents the vector dimension, Represents the l-th layer node z under relation r j For node z i The importance of , T represents the transpose of the matrix;

[0026] (2) A multi-head attention mechanism is used to learn the neighborhood information of nodes. There are multiple attention heads under each relationship. The importance of neighboring nodes in each attention head is calculated, and then the neighborhood information is summarized to represent the node:

[0027] In the formula, The representation vector of node i under the l-th layer of C attention heads under the relationship r, represents the attention weight calculated by the cth attention head, Represents the corresponding value vector, N r (i) represents the neighborhood of node i under relation r, C represents the number of attention heads, and σ represents the activation function;

[0028] (3) Use the RELU activation function to extract evidence from the enhanced relationship view, strong relationship view, and weak relationship view, respectively, as follows:

[0029] Where j∈{1,2,3}, represents the evidence vector of node i under relation rj, represents the strength of evidence that node i is a social robot under relation rj, represents the strength of evidence proving that node i is a real user under relationship rj, fc represents the fully connected layer, The final representation vector learned by node i under the representation relationship rj.

[0030] Furthermore, in the step 5, the belief quality is the proportion of the evidence value in the total evidence, and the overall uncertainty measures the confidence of the model in the current node classification;

[0031] Quality of Belief The calculation formula is as follows:

[0032] In the formula, represents the belief quality that node i is category k under relation rj, represents the strength of evidence proving that node i is of category k under relation rj, represents the global sum of evidence, K represents the number of classification categories, and +1 is the normalized bias term;

[0033] Overall uncertainty The calculation formula is as follows:

[0034] In the formula, represents the overall uncertainty of node i in relation rj, It is a positive value, and K represents the number of classification categories.

[0035] Furthermore, the implementation process of step 6 includes: firstly, performing belief distribution fusion of two perspectives, and then combining all views using the Dempster rule to obtain the final joint belief quality. The belief distribution fusion calculation formula of the two perspectives is as follows:

[0036] In the formula, Respectively represent the joint belief quality and uncertainty after the fusion of two views, They represent the belief quality of node i being category k under relations r1 and r2 respectively. represents the overall uncertainty of node i in relations r1 and r2, Represents the conflict coefficient, which is used for normalization processing. m and n represent different classifications.

[0037] The prediction with uncertainty is then handled by the cross entropy loss function, and the difference between the true label distribution and the model prediction distribution is measured using the KL divergence. The cross entropy loss function is as follows:

[0038] In the formula, yij is the true category label of the ith sample, pij is the probability that the ith sample belongs to category j, αij is the Dirichlet distribution parameter, which is used to indicate the strength of evidence that the ith sample belongs to category j, B(αi) is the normalization constant of the Dirichlet distribution, K represents the number of classification categories, and b i,k 、u i They represent the joint belief quality and uncertainty after fusion of multiple views, respectively.

[0039] The present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned social robot detection method based on multi-view evidence fusion when executing the computer program.

[0040] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned social robot detection method based on multi-view evidence fusion.

[0041] The beneficial effects of the above scheme are:

[0042] (1) The present invention introduces implicit connections. On the one hand, it combines implicit connections with explicit connections to construct a more comprehensive and representative graph structure. On the other hand, the implicit connections constructed are mostly homogeneous connections, which effectively alleviates the homogeneity assumption problem. This method uses implicit connections for analysis, which can reveal the deeper relationship structure in social networks, break through the limitations of traditional methods, and effectively identify social robots with clever disguises.

[0043] (2) Through the multi-view fusion method based on evidence theory, uncertainty quantification and evidence fusion of multiple views are performed, which makes more effective use of a more comprehensive graph structure, further solves the homogeneity assumption problem, and reduces the misjudgment rate.

[0044] (3) Through a large number of experiments on two widely used social robot detection datasets, it is demonstrated that the present invention outperforms other baseline models in key indicators such as accuracy, precision and F1 score, which verifies the effectiveness and superiority of the present method and highlights its potential in promoting the development of the field of social robot detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 The present invention is based on the implicit connection mining flow chart of the variational autoencoder;

[0046] Figure 2 Flow chart of the multi-view fusion method based on evidence theory of the present invention;

[0047] Figure 3 Experimental diagram of the sensitivity of implicit connection threshold parameters of the present invention. DETAILED DESCRIPTION

[0048] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0049] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0050] Embodiment 1:

[0051] A social robot detection method based on multi-view evidence fusion, comprising the following steps:

[0052] Step 1: preprocessing user features to obtain user representation vectors, wherein the user features include semantic features, classification features, and digital features;

[0053] Step 2: Input the user representation vector into an encoder, and the encoder combines a variational autoencoder and a supervised contrastive learning strategy to learn the distinguishing features between the social robot and the real user;

[0054] Step 3: input the output result of the encoder to the decoder, and the decoder calculates the probability that there is an implicit connection between users according to the distinctive features, and obtains an implicit connection relationship graph containing only implicit connections;

[0055] Step 4: by analyzing the explicit connection edges in the social network and the implicit connection edges of the implicit connection relationship graph, a reconstructed graph containing three different types of edges is generated, wherein the three types of edges include enhanced edges, strong relationship edges, and weak relationship edges. The enhanced edges refer to edges included in the implicit connections but not in the explicit connections, the strong relationship edges refer to edges included in both the implicit connections and the explicit connections, and the weak relationship edges refer to edges included in the explicit connections but not in the implicit connections;

[0056] Step 5: Use graph attention network and evidence deep learning to extract evidence from multiple views of the reconstructed graph, and analyze the belief quality and overall uncertainty of each category based on the extracted evidence vector, so as to obtain evidence from multiple views;

[0057] Step 6: Use the evidence fusion rules of DS evidence theory to fuse the evidence from multiple views to obtain the final joint belief quality and uncertainty to complete social robot detection.

[0058] Each step of the present invention is described in detail below:

[0059] Step 1: Preprocess user features to obtain user representation vectors. User features consist of three parts: semantic features, classification features, and digital features.

[0060] First, Z-score normalization is performed on each digital feature of the user. Taking the i-th digital feature of user u as an example, the processing method is as follows:

[0061] In the formula, Represents the i-th digital feature of user u The standardized processing results, represents the average value of all user digital features i in the dataset, It represents the standard deviation of the i-th numerical feature of all users in the dataset. n is the abbreviation of numerical, which means numerical feature.

[0062] The final representation vector u of the user u's digital features n The result of concatenating its various digital features is transformed through MLP. The specific method is as follows:

[0063] In the formula, w n and b n is a learnable parameter, r is the total number of digital features, σ is the activation function, and cat represents the concatenation operation.

[0064] The categorical features are first one-hot encoded:

[0065] In the formula, Represents the jth classification feature of user u The result after one-hot encoding processing, c is the abbreviation of classification, which means classification feature.

[0066] For each user, if a categorical feature is None, the feature is assigned a default value of 0. The final representation vector u of the numeric feature of user u c The result of concatenating its various digital features is transformed through MLP. The specific method is as follows:

[0067] In the formula, w c and b c is a learnable parameter, t is the total number of classification features, and σ is the Leaky-ReLU activation function.

[0068] Semantic features include user tweets and user descriptions. For user u, the semantic features of his tweets are u t , the user description feature is u d First, use the pre-trained RoBERTa model to encode each word in the user description and obtain the preliminary vector representation s of the user description features of user u:

[0069] In the formula, Represents the user description information of user u, including L words, s i represents the i-th word in its user description, D r is the embedding dimension of the pre-trained model RoBERTa.

[0070] Then, user description feature u of user u d Calculated by the following method: u d =σ(w d ·s+b d )

[0071] Where w d and b d is a learnable parameter and σ is the Leaky-ReLU activation function.

[0072] Similar to processing user description features, we first use the pre-trained RoBERTa model to encode each word in the user's tweet to obtain a preliminary vector representation of the user's tweet:

[0073] In the formula, Represents the user tweet information of user u, containing H words, t i represents the i-th word in the user’s tweet, D r is the embedding dimension of the pre-trained model RoBERTa.

[0074] Then, user tweet feature u of user u t Calculated by the following method: u t =σ(w t ·t+b t )

[0075] In the formula, w t and b t is a learnable parameter and σ is the Leaky-ReLU activation function.

[0076] By combining the above-mentioned user's digital features, classification features, tweet features and description features, we can get the overall representation vector x of user i. i : x i =cat(u n ,u c ,u t ,u d )

[0077] Step 2: Input the user representation vector obtained in step 1 into the encoder, which combines the node potential representation learning of the variational autoencoder and the strategy of supervised contrastive learning to learn the distinctive features between the social robot and the real user.

[0078] Implicit connection mining methods based on variational autoencoders such as Figure 1 shown.

[0079] First, the user's representation is mapped to a latent space, that is, the high-dimensional input is converted into a low-dimensional representation: h i =RELU(wx i +b)

[0080] In the formula, h i represents the low-dimensional representation of user i, w and b are learnable parameters, and RELU is a nonlinear activation function.

[0081] Then, two multi-layer perceptrons (MLPs) are used to calculate the mean μ and variance σ of the Gaussian distribution of the low-dimensional representation:

[0082] Where θ1 and θ2 represent the parameters of the two MLP models respectively.

[0083] Next, we use the reparameterization trick to randomly sample ε from the standard normal distribution, making the implicit connection calculation trainable. Also, because the representation of the node is no longer a deterministic vector, but a random sample in the latent space, the model is more robust to noisy data. The latent variable is calculated according to the following formula:

[0084] In the formula, z i represents the potential representation after random sampling of user i, μ i represents the mean, σ i represents the variance, and ε represents the value randomly sampled from a standard normal distribution.

[0085] In order to prevent model overfitting and improve model performance, on the one hand, the distribution is restricted to the standard normal distribution through KL divergence:

[0086] Where n represents the dimension of vector z.

[0087] At the same time, the idea of ​​supervised contrastive learning (SCL) is combined. Although our dataset does not show obvious data sparsity, the disguised behavior of social robots introduces a large number of false edges in the network topology, which undermines the traditional homogeneity assumption. This method can train the model to focus on learning implicit connections that are more likely to be homogeneous. The contrast loss function of SCL is defined as follows:

[0088] In the formula, i∈I≡{1...2N} represents the index of the sample after data enhancement. For example, the original data set has N samples, and each sample is sampled twice by VAE to obtain 2N enhanced samples. P(i) represents the index set of positive samples, and A(i) represents the index set of negative samples that are different from sample i. τ is the temperature parameter, z i is the feature representation of the i-th sample, z p is with z i Samples belonging to the same category are the feature representation of positive samples, z a is with z i Feature representation of samples belonging to different categories.

[0089] Combining the node potential representation learning of variational autoencoders with supervised contrastive learning can effectively improve the model's ability to capture the distinctive features of social robots and real users. VAE latent space modeling: By learning the potential distribution of social network nodes (users) through the encoder, it can capture implicit behavioral patterns and structural features (such as posting topic rules, social interaction patterns, etc.). Supervised contrastive learning uses label information (social robots / real users) to optimize contrast loss, forcing similar samples to aggregate in the latent space and heterogeneous samples to separate, which can enhance the model's sensitivity to local structural features and improve the clarity of classification boundaries.

[0090] Step 3: Input the output result of the encoder into the decoder, calculate the probability that there is an implicit connection between users, and if the probability exceeds a specified threshold, infer that there is an implicit connection between the two points, thereby obtaining an implicit connection relationship graph containing only implicit connections.

[0091] The probability that there is an implicit connection between users is calculated as follows:

[0092] In the formula, p i,j represents the probability of implicit connection between users, sigmoid(·) represents the sigmoid function, z i 、z j They represent the feature representation of the i-th sample and the j-th sample respectively, and T represents the transpose of the matrix.

[0093] If p i,j If it is greater than the threshold, then the node z is considered i 、z j There is an implicit connection between them.

[0094] Step 4: By analyzing the implicit connection edge E of the implicit connection relationship graph imp and the explicit connection edge E in the data set exp , generating a reconstructed graph containing three different types of edges.

[0095] These three edges include the enhancement edge E enh , strong relationship edge E str and weak tie edge E weak . Enhanced edge E enh , that is, the edge included in the implicit connection but not in the explicit connection, which supplements the potential relationship that cannot be captured in the explicit connection; the strong relationship edge E str , that is, the edge included in both implicit and explicit connections; weak relationship edge E weak , that is, the edges included in the explicit connections but not in the implicit connections, such as the follow relationship with the official account initially established by the social platform, and the follow relationship or like relationship established out of courtesy with unfamiliar classmates.

[0096] These three types of edges provide another perspective on user interaction. Enhanced edges and strong relationship edges contain more heterogeneous edges compared to weak relationship edges. Dealing with these three types of edges and possible conflicts separately helps us to more comprehensively handle the structural information of the graph and the problem of the failure of the homogeneity assumption.

[0097] The calculation methods for the three different types of edges are as follows: E enh =E imp ∪E exp -E exp E str =E imp ∩E exp E weak =E imp ∪E exp -E imp

[0098] In the formula, E enh represents the set of enhanced edges, E str represents the set of strong relationship edges, E weak represents the set of weak relationship edges, E exp represents the set of explicit connection edges, E imp Represents the set of implicitly connected edges.

[0099] Step 5: Use graph attention network and evidence deep learning to extract multi-view evidence of the reconstructed graph. In this embodiment, evidence is extracted from three reconstructed graphs, i.e., three views, and the belief quality and overall uncertainty of each category are analyzed based on the extracted evidence vector to obtain evidence from multiple views.

[0100] The three views are enhanced relationship view, strong relationship view and weak relationship view. The enhanced relationship view only contains enhanced edges, the strong relationship view only contains strong relationship edges, and the weak relationship view only contains weak relationship edges.

[0101] Specifically, for a node pair (z i ,z j ), the calculation process of importance weight is as follows:

[0102] In the formula, {l} represents the lth layer of the graph neural network, represents the query vector of node i under relation r, represents the key vector of node j under relation r, represents the value vector of node j under relation r, W r,{l} , b r,{l}They represent the learnable parameter matrices of the lth layer under the relation r, Respectively represent the node z of the lth layer under the relationship r i , node z j , N r (i) represents node z under relation r i Neighborhood, d represents the vector dimension, Represents the l-th layer node z under relation r j For node z i The importance of , T represents the transpose of the matrix.

[0103] Multi-head attention is used to learn neighborhood information. There are C attention heads under each relationship. The importance of neighbor nodes in each attention head is calculated in the above way, and then the neighborhood information is summarized to learn node representation.

[0104] In the formula, The representation vector of node i under the l-th layer of C attention heads under the relationship r, represents the attention weight calculated by the cth attention head, Represents the corresponding value vector, N r (i) represents the neighborhood of node i under relation r, C represents the number of attention heads, and σ represents the activation function.

[0105] Use h r1 、h r2 、h r3 denote the evidence extracted from the enhanced relation view, strong relation view, and weak relation view, respectively, with h r1 Take as an example, use the RELU activation function to extract evidence to ensure that the evidence value is non-negative. The expression of evidence extraction is as follows:

[0106] Where j∈{1,2,3}, represents the evidence vector of node i under relation rj, represents the strength of evidence that node i is a social robot under relation rj, represents the strength of evidence proving that node i is a real user under relationship rj, fc represents the fully connected layer, The final representation vector learned by node i under the representation relationship rj.

[0107] After evidence is extracted, further analysis is required based on the extracted evidence vector To calculate the belief quality and overall uncertainty of each category. These quantitative indicators can be used to measure the model's support for a category and the confidence level of the classification.

[0108] Each element in the evidence vector represents the degree of support for a certain category, but these evidence values ​​themselves are not normalized, so a global evidence sum S needs to be calculated. rj , as the basis for normalization. The formula is as follows:

[0109] In the formula, represents the global sum of evidence, It represents the strength of evidence that proves that node i is of category k under relationship rj, K represents the number of classification categories, and +1 is the normalized bias term, which ensures that the belief quality and uncertainty can always form a complete probability distribution numerically.

[0110] Quality of Belief It is the normalized support of a node belonging to a certain category, which is defined as the proportion of the evidence value in the total evidence. The formula is as follows:

[0111] In the formula, represents the belief quality that node i is category k under relation rj, represents the strength of evidence proving that node i is of category k under relation rj, Represents the global sum of evidence.

[0112] Higher The value indicates that the model has strong support for category k, reflecting the model's classification confidence in this category.

[0113] The overall uncertainty u measures the model's confidence in the classification of the current node. If the total evidence S is large, the u value is small, indicating that the model's classification confidence is high; conversely, if S is small, the u value is large, indicating that the model's classification confidence is insufficient. Belief Quality The calculation formula is as follows:

[0114] In the formula, represents the overall uncertainty of node i in relation rj, K represents the number of classification categories, and uncertainty u is a positive value that satisfies

[0115] Step 6: After the belief quality and uncertainty of different views are calculated, the evidence of multiple views is fused using the DS (Dempster-Shafer, DS) evidence theory fusion rule to generate the final joint belief quality and uncertainty.

[0116] Multi-view fusion methods based on evidence theory are Figure 2Specifically, after uncertainty quantization, we first obtain the views of the three views. For the node i in the enhanced relationship view, it is represented as Node i in the strong relationship view is represented as Node i in the weak relationship view is represented as

[0117] Then use the Dempster-Shafer rule to take the belief distribution of two perspectives for fusion as an example:

[0118] In the formula, Respectively represent the joint belief quality and uncertainty after the fusion of two views, They represent the belief quality of node i being category k under relations r1 and r2 respectively. represents the overall uncertainty of node i in relations r1 and r2, represents the conflict coefficient, which is used for normalization processing, and m and n represent different categories. In this embodiment, m and n are equal to 1 or 2.

[0119] When multiple views are involved, combine them sequentially using Dempster's rule as follows:

[0120] After integrating all views, the final joint belief representation M is obtained i = {b i,1 ,b i,2 ,u i}. Generate the Dirichlet distribution parameter α according to the following equation:

[0121] Where b i,k 、u i They represent the joint belief quality and uncertainty after fusion of multiple views, respectively, and K represents the number of classification categories.

[0122] In order to adapt to the Dirichlet distribution and uncertainty quantification, the cross entropy loss is modified on this basis so that it can handle predictions with uncertainty.

[0123] Modified cross entropy loss: In the context of uncertainty quantification, the class probability output by the model is not a standard probability distribution, but is based on the Dirichlet distribution. Therefore, the loss function needs to be calculated under the framework of the Dirichlet distribution. Specifically, the cross entropy loss function is defined by integrating the Dirichlet distribution:

[0124] In the formula, y ij is the true category label of the i-th sample, p ij is the probability that the i-th sample belongs to the j-th category, α ij is the Dirichlet distribution parameter, which is used to represent the strength of evidence for the i-th sample, category j, B(α i ) is the normalization constant of the Dirichlet distribution, and K represents the number of classification categories.

[0125] The evidence adjustment loss is used to adjust and optimize the evidence output of the Dirichlet distribution to reduce the noise introduced by model errors or uncertainties. By maximizing the consistency of evidence between data and predictions, the evidence adjustment loss helps the model gradually improve the prediction confidence of each category during the training process. This embodiment uses KL divergence (Kullback-Leibler Divergence) as the loss function for evidence adjustment to measure the difference between the true label distribution and the model prediction distribution. Specifically, the expression of the KL divergence loss function is:

[0126] In the formula, is the Dirichlet parameter after removing non-misleading evidence in the predicted parameters, α i refers to the original Dirichlet parameter, y i Refers to the true classification label of the sample, D(p i |1) is a uniform Dirichlet distribution.

[0127] By mining implicit connections and constructing different views, and then using DS evidence theory to fuse evidence from multiple views to detect social robots, the detection performance can be effectively improved, the misjudgment rate can be reduced, and the detection accuracy and precision can be improved.

[0128] We use two widely used social robot detection datasets on Twitter to evaluate the effectiveness of the proposed method: TwiBot-20 and MGTAB. These datasets provide a wide range of entities and relations. The TwiBot-20 dataset was made public in 2020, which includes 229,573 users, 8,723,736 user attributes, 33,488,192 tweets, and 455,958 follow-up relations. The MGTAB dataset was made public in 2023 and contains 10,199 expert-annotated users and 7 relationship types. Since MGTAB only provides network structure and processed user features, we rely on the account and tweet features provided by the benchmark instead of deploying our feature engineering-based processing strategy. At the same time, some methods that rely on raw data cannot be applied to this benchmark. Therefore, some methods cannot be reproduced in the MGTAB dataset.

[0129] The accuracy, F1 score, precision and recall of the proposed method on the TwiBot-20 and MGTAB datasets are better than other baseline models, including the most advanced graph-based baseline model. Compared with previous graph-based methods, the method of the present invention not only mines the potential relationships between users but also enhances the use of more comprehensive data by using DS evidence theory. The specific experimental results are shown in Table 1, where Ours represents the social robot detection method based on multi-view evidence fusion described in this embodiment. Table 1 Experimental results of different social robot detection methods on TwiBot-20 and MGTAB datasets

[0130] Experimental results show that methods that utilize the structural information of graphs generally have better performance, and compared with only utilizing attention relationships, diversified interactive relationships can further improve model performance. The accuracy, F1 score, precision, and recall of the method of the present invention on the TwiBot-20 and MGTAB datasets are better than those of the compared baseline models. Compared with previous graph-based methods, the method of the present invention not only utilizes implicit connections between users and mines potential connections between users, but also enhances the use of more comprehensive data through DS evidence theory, thereby improving the accuracy and reliability of detection.

[0131] In order to study the impact of implicit connections between users on social robot detection, we compared the experimental results of the method of the present invention without using implicit connections, and enhanced several graph-based baseline methods using implicit connection calculation modules, and compared the performance before and after enhancement. The experimental results listed in Table 2 show that the use of implicit connections between users has a positive effect on social robot detection. Enhanced in Table 2 refers to adding an implicit connection calculation module to the original graph structure for social robot detection. Table 2 Ablation experiment of implicit connection calculation module

[0132] In order to study the effectiveness of the multi-view fusion method using evidence theory proposed in the present invention, we input the three relationship edges extracted by the present invention into the social robot detection model of HGT, BotRGCN, and RGT for processing heterogeneous graphs. The experimental results are shown in Table 3, where Ours represents the social robot detection method based on multi-view evidence fusion described in this embodiment. Table 3 Multi-view fusion ablation experiment based on evidence theory

[0133] Different implicit connection thresholds will directly affect the recognition effect of implicit connections. If the threshold is too low, it may cause noise interference, causing many irrelevant node pairs to be mistakenly marked as having implicit connections, thereby reducing the detection accuracy; on the contrary, setting the threshold too high may cause the omission of truly meaningful implicit connections and also reduce the performance of the model. Therefore, we evaluated the impact of different thresholds on the detection effect on TwiBot-20. The results are shown in Figure 2. Figure 3 The experimental results show that when the threshold is set to about 0.6, the model performance reaches the best.

[0134] In summary, the present invention can effectively improve the social robot detection performance and meet the actual needs of social platforms and regulatory agencies for ensuring the security of the network environment and maintaining the authenticity of information. This embodiment establishes three different views and performs evidence fusion to comprehensively analyze user relationships and behavior patterns from multiple perspectives, thereby capturing richer and more complex information, which helps to reveal the deeper relationship structure in social networks and more effectively identify those social robots that are cleverly disguised. Experiments have shown that it has achieved significant results in improving the social robot detection performance.

[0135] Embodiment 2:

[0136] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the social robot detection method based on multi-view evidence fusion described in Example 1 is implemented.

[0137] Embodiment 3:

[0138] This embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the social robot detection method based on multi-view evidence fusion described in Embodiment 1.

[0139] Finally, it should be noted that the parts of the present invention that are not described in detail are all prior art. Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not intended to limit the invention. Although the invention is described in detail with reference to the aforementioned examples, those of ordinary skill in the art can still modify the technical solutions recorded in the aforementioned examples, or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, etc. made within the spirit and principles of the invention should be included in the scope of protection of the invention.

Claims

1. A social robot detection method based on multi-view evidence fusion, characterized in that: The following steps are involved: Step 1: preprocessing user features to obtain user representation vectors, wherein the user features include semantic features, classification features, and digital features; Step 2: Input the user representation vector into an encoder, and the encoder combines a variational autoencoder and a supervised contrastive learning strategy to learn the distinguishing features between the social robot and the real user; Step 3: input the output result of the encoder to the decoder, and the decoder calculates the probability that there is an implicit connection between users according to the distinctive features, and obtains an implicit connection relationship graph containing only implicit connections; Step 4: by analyzing the explicit connection edges in the social network and the implicit connection edges of the implicit connection relationship graph, a reconstructed graph containing three different types of edges is generated, wherein the three types of edges include enhanced edges, strong relationship edges, and weak relationship edges. The enhanced edges refer to edges included in the implicit connections but not in the explicit connections, the strong relationship edges refer to edges included in both the implicit connections and the explicit connections, and the weak relationship edges refer to edges included in the explicit connections but not in the implicit connections; Step 5: Use graph attention network and evidence deep learning to extract evidence from multiple views of the reconstructed graph, and analyze the belief quality and overall uncertainty of each category based on the extracted evidence vector, so as to obtain evidence from multiple views; Step 6: Use the evidence fusion rules of DS evidence theory to fuse the evidence from multiple views to obtain the final joint belief quality and uncertainty to complete social robot detection.

2. The social robot detection method based on multi-view evidence fusion according to claim 1 is characterized in that: The implementation process of step 2 includes: firstly, mapping the user representation vector to a latent space, converting the high-dimensional input data into a low-dimensional representation, respectively calculating the mean and variance of the Gaussian distribution of the low-dimensional representation through two multi-layer perceptrons MLP, using the re-parameterization technique to perform random sampling from the standard normal distribution, and calculating the latent variables according to the following formula: z i =μ i +εσ i ,ε~N(0,1) In the formula, z i represents the potential representation after random sampling of user i, μ i represents the mean, σ i represents the variance, ε represents the value randomly sampled from the standard normal distribution; Then, samples with the same label and corresponding VAE sampling results are marked as positive samples, and samples with different labels are marked as negative samples. Combined with supervised contrastive learning, the following contrastive loss function is designed to train the model: In the formula, i∈I≡{1…2N} represents the index of the sample after data enhancement, P(i) represents the index set of positive samples, A(i) represents the index set of negative samples of different categories from i, τ is the temperature parameter, z i is the feature representation of the i-th sample, z p is with z i The feature representation of samples belonging to the same category, z a is with z i Feature representation of samples belonging to different categories.

3. The social robot detection method based on multi-view evidence fusion according to claim 2 is characterized in that: In step 3, the calculation formula for the probability of implicit connection between users is as follows: In the formula, p i,j represents the probability of implicit connection between users, sigmoid(·) represents the sigmoid function, z i 、z j They represent the feature representation of the i-th sample and the j-th sample respectively, and T represents the transpose of the matrix; If p i,j is greater than the set threshold, then node z i 、z j There is an implicit connection between them.

4. The social robot detection method based on multi-view evidence fusion according to claim 1 is characterized in that: In step 4, the calculation methods of the three different types of edges are as follows: AND enh =And imp ∪E exp -AND exp AND str =And imp ∩E exp AND weak =And imp ∪E exp -AND imp In the formula, E enh represents the set of enhanced edges, E str represents the set of strong relationship edges, E weak represents the set of weak relationship edges, E exp represents the set of explicit connection edges, E imp Represents the set of implicitly connected edges.

5. The social robot detection method based on multi-view evidence fusion according to claim 1 or 4, characterized in that: In the step five, a graph attention network and evidence deep learning are used to extract evidence of three views of the reconstructed graph, where the three views are an enhanced relationship view, a strong relationship view, and a weak relationship view. The enhanced relationship view only contains enhanced edges, the strong relationship view only contains strong relationship edges, and the weak relationship view only contains weak relationship edges.

6. The social robot detection method based on multi-view evidence fusion according to claim 5 is characterized in that: In step 5, the process of evidence extraction is as follows: (1) The importance weight between two nodes under relationship r is calculated through the graph attention network. The calculation process is as follows: In the formula, {l} represents the lth layer of the graph neural network, represents the query vector of node i under relation r, represents the key vector of node j under relation r, represents the value vector of node j under relation r, W r,{l} 、b r,{l} They represent the learnable parameter matrices of the lth layer under the relation r, Respectively represent the node z of the lth layer under the relationship r i , node z j , N r (i) represents node z under relation r i Neighborhood, d represents the vector dimension, Represents the l-th layer node z under relation r j For node z i The importance of , T represents the transpose of the matrix; (2) A multi-head attention mechanism is used to learn the neighborhood information of nodes. There are multiple attention heads under each relationship. The importance of neighboring nodes in each attention head is calculated, and then the neighborhood information is summarized to represent the node: In the formula, The representation vector of node i under the l-th layer of C attention heads under the relationship r, represents the attention weight calculated by the cth attention head, Represents the corresponding value vector, N r (i) represents the neighborhood of node i under relation r, C represents the number of attention heads, and σ represents the activation function; (3) Use the RELU activation function to extract evidence from the enhanced relationship view, strong relationship view, and weak relationship view, respectively, as follows: Where j∈{1,2,3}, represents the evidence vector of node i under relation rj, represents the strength of evidence that node i is a social robot under relationship rj, represents the strength of evidence proving that node i is a real user under relationship rj, fc represents the fully connected layer, The final representation vector learned by node i under the representation relationship rj.

7. The social robot detection method based on multi-view evidence fusion according to claim 1 is characterized in that: In step 5, the belief quality is the proportion of the evidence value in the total evidence, and the overall uncertainty measures the confidence of the model in the current node classification; Quality of Belief The calculation formula is as follows: In the formula, represents the belief quality that node i is category k under relation rj, represents the strength of evidence proving that node i is of category k under relation rj, represents the global sum of evidence, K represents the number of classification categories, and +1 is the normalized bias term; Overall uncertainty The calculation formula is as follows: In the formula, represents the overall uncertainty of node i in relation rj, It is a positive value, and K represents the number of classification categories.

8. The social robot detection method based on multi-view evidence fusion according to claim 7 is characterized in that: The implementation process of step six includes: First, the belief distribution fusion of the two perspectives is performed, and then all views are combined using the Dempster rule to obtain the final joint belief quality. The belief distribution fusion calculation formula of the two perspectives is as follows: In the formula, Respectively represent the joint belief quality and uncertainty after the fusion of two views, They represent the belief quality of node i being category k under relations r1 and r2 respectively. represents the overall uncertainty of node i in relations r1 and r2, Represents the conflict coefficient, which is used for normalization processing. m and n represent different classifications. The prediction with uncertainty is then handled by the cross entropy loss function, and the difference between the true label distribution and the model prediction distribution is measured using the KL divergence. The cross entropy loss function is as follows: In the formula, y ij is the true category label of the i-th sample, p ij is the probability that the i-th sample belongs to category j, α ij is the Dirichlet distribution parameter, which is used to indicate the strength of evidence that the i-th sample belongs to category j. B(α i ) is the normalization constant of the Dirichlet distribution, K represents the number of classification categories, and b i,k 、u i They represent the joint belief quality and uncertainty after fusion of multiple views, respectively.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the social robot detection method based on multi-view evidence fusion described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the social robot detection method based on multi-view evidence fusion as described in any one of claims 1-8.