A method and device for detecting malicious users based on graph neural network
By creating potential graphs in malicious user detection and adopting biased sampling strategies, the problem of inability to effectively capture long-distance node features and aggregating high-similar node features in the prior art is solved, which significantly improves the performance of malicious user detection.
Patent Information
- Application Number
- CN202110266580.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-03-11
AI Technical Summary
The prior art has three main challenges in malicious user detection: it is impossible to effectively capture the feature information of long-distance nodes, cannot gather node features that are highly similar but not within the neighborhood, and the current graph neural network method's impact on different neighbors is improperly sampled.
A malicious user detection method based on graph neural network is proposed. By creating a potential graph to capture the characteristics of long-distance nodes, and using partial sampling strategies to distinguish the influence of different neighbors, gathering social and potential neighbor information to compute node embedding.
The characteristics of long-distance nodes are effectively captured, the detection capabilities of high-similar nodes are improved, and the quality of node embedding is improved through biased sampling strategies, significantly improving the performance of malicious user detection.
Smart Images

Figure CN115081581B_ABST
Abstract
Description
Technical Field
[0001] The present invention designs a malicious user detection method and device based on a graph neural network, which belongs to the intersection field of network information security and artificial intelligence. Technical Background
[0002] The rapid development of social media has enabled people to express their opinions freely online. However, it has also become a hotbed for hate speech. The dramatic increase in hateful content on the Internet has led to the emergence of conflicts. Therefore, the classification of hate speech has become a topic of increasing interest in industry and academia.
[0003] Despite the efforts in this area, hate speech detection remains challenging. First, some models oversimplify the problem, such as only considering tweets with hate-related words. These methods rely entirely on textual (e.g., lexical and semantic) features, ignoring user and community information. Second, some hate speech classifiers are very vulnerable to very simple and model-independent attacks. In some cases, these attacks reduce the detection recall by nearly 50%. Therefore, hate speech classifiers that rely only on text detection are not robust to attacks that intentionally mislead the classifier.
[0004] Studying hate speech can not only start from the text content in social networks, but also from the information related to the individual or organization itself. Studying such information provides a lot of opportunities to explore a richer feature space, which helps to identify online hate speech. In addition, when detecting hate speech, it is more common to consider the user's profile rather than just isolated tweets. In addition, directly identifying hate users who intentionally spread hate speech is an important and effective measure to combat online hate speech. Therefore, the present invention shifts the focus from hate speech to hate users.
[0005] Malicious users can be detected through users' tweets, profiles, and social relationships, so malicious users can be detected through graph neural networks (GNNs). Specifically, users in a social network are nodes in the graph, and social relationships can be regarded as edges. Each node contains rich attribute information, such as the user's profile, the user's tweets, etc. By using deep neural networks to aggregate feature information of neighboring nodes, neural networks can be applied to such node classification tasks. However, there are three main challenges in detecting malicious users based on graph neural networks (GNNs).
[0006] Aggregating neighbor information may be ineffective for nodes that have no or too few relationships with other nodes in the graph.
[0007] The neighbors of a node are defined as the set of first-order or multi-order neighbors. Conventional methods aggregate only the features of the neighborhood. However, there may be nodes that are very similar to the target node, but are not in the neighborhood of that node. Existing methods cannot aggregate such high-similarity (non-neighborhood) nodes.
[0008] Current graph neural networks (GNNs), especially spatial-based methods such as GraphSAGE, sample all neighbors equally when aggregating their information. It does not take into account that different neighbors may have different influences on a node. Summary of the invention
[0009] In order to solve the above limitations, the present invention proposes a malicious user detection method and device based on graph neural network, which uses a malicious user detection framework - HateGNN, to detect malicious users in social networks. HateGNN has the following significant features:
[0010] To address the first and second challenges mentioned above, in addition to the explicit social graph, the present invention also creates a latent graph based on node attribute information. This will be of great help to nodes that do not have any social neighbors or have too few neighbors. In addition, the latent graph has the ability to capture important features from distant but information-rich nodes.
[0011] To address the third challenge, the present invention adopts a biased sampling strategy for sample neighbors (not only immediate neighbors but also potential neighbors) to distinguish the influence of neighbors. The present invention proposes a sampling strategy to help select the most informative features.
[0012] Once the neighbors are selected, the present invention aggregates social and potential neighbors to compute the final node embedding.
[0013] In order to achieve the above object, the present invention provides the following technical solutions:
[0014] A malicious user detection method based on graph neural network, the steps of which include:
[0015] 1) Collect user tweets in social networks and build a social graph G o , where users are social graphs G o If there are social relationships between users, then a social graph G is constructed. o The edges of the corresponding nodes in ;
[0016] 2) According to the attributes of each node, calculate the initial features of each node, and according to the social graph G o The similarity of the initial features of any two nodes in the latent graph G is obtained l The edge of
[0017] 3) Through the potential graph Gl The edges and social graph G o Nodes on it, construct the potential graph G l ;
[0018] 4) Based on the initial features of each node, o The social neighborhood in the latent graph G l The latent neighborhood in generates the node embedding representation of each node, where the social neighborhood consists of the neighboring nodes of the target node, and the latent neighborhood consists of the nodes whose similarity with the target node is higher than the set threshold;
[0019] 5) Based on the node embedding representation of each node, the malicious user detection result is obtained.
[0020] Furthermore, the attributes include: content-related attributes, activity-related attributes, emotion-related attributes and / or structure-related attributes.
[0021] Furthermore, the content-related attributes include: tweets of the user.
[0022] Furthermore, the activity-related attributes include: the number of tweets, the number of retweets of tweets, the number of fans, the number of followings, the number of favorites, the number of tags, the number of citations, the number of URLs, the average number of mentions per tweet, the average time interval between each tweet, and the median time interval between each tweet.
[0023] Further, sentiment-related attributes include: sentiment of the tweet and use of profanity.
[0024] Furthermore, the structure-related attributes include: centrality, eigenvector, in-degree, out-degree and content-related attributes of the user's first-order neighbors, activity-related attributes, sentiment-related attributes and structure-related attributes.
[0025] Furthermore, methods for calculating initial features include: one-hot encoding, label encoding, or word embedding using GloVe.
[0026] Furthermore, the edges of the potential graph are obtained by the following steps:
[0027] 1) Based on the initial feature Z of node v v and the initial feature Z of node u u , calculate the similarity between node v and node u S(v,u) = PearsonSimilarity(Z v , Z u ), where v,u∈V, V is the node set of the social graph;
[0028] 2) When the similarity S(v,u) is greater than a set threshold When , the edge between node v and node u is taken as the edge between node v and node u in the potential graph.
[0029] Furthermore, the node embedding representation of each node is generated through the following steps:
[0030] 1) Based on the social graph G o With the latent graph G l , get a graph neural network;
[0031] 2) Obtaining neighbor nodes u of node v in the graph neural network, wherein the neighbor nodes include: nodes in the social neighborhood and nodes in the potential neighborhood;
[0032] 3) Determine whether the total number of neighbor nodes u is less than the set value β; if true, go to 4); if false, go to 5);
[0033] 4) Determine whether the similarity between node v and neighbor node u is greater than the threshold η; if true, go to 5); if false, go to 2);
[0034] 5) Take neighbor node u as sampled neighbor node u′ and add it to the sampled neighbor node set of node v;
[0035] 6) Mean aggregation samples the embedding representation of the k-1 layers of neighbor nodes u′ to generate the k-th layer neighbor node feature aggregation representation of node v And the k-th layer neighbor node feature aggregation representation of node v and the k-1th layer embedding representation of node v After concatenation and linear transformation, the k-th layer embedding representation of node v is obtained Where 1≤k≤K, K is the number of layers of aggregated neighbor nodes, is the set of sampled neighbor nodes of node v, is the initial feature;
[0036] 7) Embedding representation for the kth layer After normalization;
[0037] 8) Loop through steps 2)-7) to get the Kth layer embedding representation as the node embedding representation of node v.
[0038] A storage medium stores a computer program, wherein the computer program is configured to execute the method described above when running.
[0039] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer to execute the method described above.
[0040] Compared with the prior art, the present invention has the following advantages:
[0041] 1) The existing technology cannot capture the feature information of distant nodes well. The present invention creates a latent graph based on node attribute information, which can help capture rich feature information of distant nodes;
[0042] 2) The prior art ignores the differences between nodes when sampling nodes. The present invention proposes a biased sampling strategy that considers the similarities between nodes when sampling, which can help select nodes with rich feature information. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Schematic diagram of the HateGNN framework of the present invention.
[0044] Figure 2 The present invention is a flow chart of the sampling method. DETAILED DESCRIPTION
[0045] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention and to make the objects, features and advantages of the present invention more obvious and easy to understand, the technical core of the present invention is further described in detail below with reference to the accompanying drawings and examples.
[0046] The HateGNN method proposed in this invention is described below. The overall framework is as follows: Figure 1 First, obtain the social graph The nodes are users. If there is a social relationship between nodes, there is an edge between them. Second, the nodes are characterized. Each node has its attributes (such as the attributes listed in Table 1). Third, when the similarity of the feature vectors between two nodes exceeds a threshold, When , they are connected by a potential edge. These potential edges and corresponding nodes form a potential graph Fourth, a biased neighborhood sampling strategy is implemented. Neighbors that are more similar to the target node (including social neighbors and potential neighbors) have higher sampling priority until a fixed-size neighbor set is obtained. Fifth, choose your neighbors Afterwards, the present invention aggregates their attribute information and obtains the node embedding X by multiplying the weight matrix trained with the neural network v Finally, embed X through the node v Determine whether the user node is a malicious user.
[0047] Next, the creation of the potential graph in the present invention is described. For malicious user detection, the attributes of each node v∈V include three categories, namely content-related, activity-related, and sentiment-related. In addition, the topological structure of each node can also be regarded as an attribute. The attributes used in the present invention are detailed in Table 1. The present invention obtains a vector representation of each attribute (for example, one-hot encoding, label encoding, word embedding using GloVe), and obtains the initial features of the node v∈V, represented as z v For u, v∈V, their similarity is defined as follows:
[0048] S(v,u)=PearsonSimilarity(z v , z u )
[0049] When the similarity between two nodes exceeds a threshold , they are connected by latent edges, which ultimately creates the latent graph When a node has no potential edge with any other node, the node is considered an isolated node and remains in the potential graph.
[0050] Table 1: Various node attributes used in HateGNN
[0051]
[0052]
[0053] When a graph has high-degree nodes (i.e., nodes have a large number of neighbors), it is usually inefficient and unnecessary to consider all neighbors for aggregation. The GraphSAGE method assumes that the neighbors of a node in the graph do not have an order like sentences or images, so a set of neighbors of a fixed size is randomly sampled. The HateGNN framework proposed in this paper improves GraphSAGE by selecting a set of sampled neighbors based on similarity. The reason is that similar neighbors can consolidate and enhance the node embedding results.
[0054] Figure 2 is a flow chart of the sampling method of the present invention, wherein the initial features z of all nodes are v , the weight matrix W of the neural network k , σ is a nonlinear activation function, K is the order of neighbors to be aggregated, V is all nodes, k is the order, v, u are nodes, N (v) is the neighbor node of node v, is the sampling neighbor node, is the number of sampled neighbor nodes, β is the specified number of sampled nodes, S(v,u) is the PearsonSimilarity(z v , z u ), η is the set similarity threshold, is the k-th layer neighbor node feature aggregation representation of node v, Embedding the k-th layer of node v, the steps include:
[0055] 1) Input the initial features of all nodes
[0056] 2) Loop through step 3)-10) K times, and then end the loop and execute step 1);
[0057] 3) Traverse all nodes V, and for each node v, perform operations 4)-9);
[0058] 4) Traverse the neighbors of node v and perform operations 5)-7) for each neighbor node u;
[0059] 5) Determine whether the total number of sampled neighbor nodes of node v is less than the set value β. If true, go to 6); if false, go to 7);
[0060] 6) Determine whether S(v, u) between nodes v and u is greater than the threshold η. If true, go to 7); if false, go to 4);
[0061] 7) Add node u to the neighbor nodes of v;
[0062] 8) Mean aggregate the embedding representations of the k-1 layers of neighboring nodes to generate the k-th layer of neighboring node feature aggregation representation of node v The formula is
[0063] 9) It is concatenated with the k-1th layer embedding representation of node v and, after nonlinear transformation, the kth layer embedding representation is obtained The formula is Go to 3);
[0064] 10) Yes Perform normalization and go to 2);
[0065] 11) Output the final node embedding.
[0066] The neighbors of node v N(v) = {N o (v),N l (v)} includes its neighbors in the social graph and the latent graph. Social Neighborhood Social Graph The set of adjacent nodes of v in the potential neighborhood are those nodes whose similarity to node v is higher than parameter During the aggregation process, the present invention combines social neighborhoods and latent neighborhoods to generate node embeddings. The original intention is that different types of neighbors will make different contributions to the final node representation. For social neighbors, it represents the influence of the user's social nature. Compared with this explicit relationship, latent neighbors represent long-term dependencies with nodes, which are invisible and cannot be captured directly. Therefore, step 8) in HateGNN can be updated with the following formula, where and are the sampled neighbors from the social graph and the latent graph respectively, and mean is the mean aggregation. That is:
[0067]
[0068] Positive Effects
[0069] In view of the increasing harmful impact of online hate speech, this paper addresses this problem by developing a framework. The framework is used to detect hate users in social networks. It not only utilizes the explicit social graph but also constructs a latent graph. In addition, an effective biased neighbor sampling technique is adopted. Experiments show that on two standard hate speech detection datasets, the proposed model has better performance in malicious user detection than existing graph neural network methods such as GraphSAGE, as shown in Table 2.
[0070] Table 2: Comparison of HateGNN with baseline models and state-of-the-art graph models
[0071]
[0072]
[0073] The above embodiments are provided only for the purpose of describing the present invention, and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present invention should all be included within the scope of the present invention.
Claims
1. A malicious user detection method based on graph neural network, the steps of which include: 1) Collect user tweets in social networks and build a social graph G o , where users are social graphs G o If there are social relationships between users, then a social graph G is constructed. o The edges of the corresponding nodes in ; 2) According to the attributes of each node, calculate the initial features of each node, and according to the social graph G o The similarity of the initial features of any two nodes in the latent graph G is obtained l The edge of 3) Through the potential graph G l The edges and social graph G o Nodes on it, construct the potential graph G l ; 4) Based on the initial features of each node, o The social neighborhood in the latent graph G l The node embedding representation of each node is generated by the potential neighborhood in the target node, where the social neighborhood consists of the neighboring nodes of the target node, and the potential neighborhood consists of the nodes whose similarity with the target node is higher than the set threshold; wherein the node embedding representation of each node is generated by the following steps: 4.1) Based on the social graph G o With the latent graph G l , get a graph neural network; 4.2) Obtaining neighbor nodes u of node v in the graph neural network, wherein the neighbor nodes include: nodes in the social neighborhood and nodes in the potential neighborhood; 4.3) Determine whether the total number of neighbor nodes u is less than the set value β; if true, go to 4.4); if false, go to 4.5); 4.4) Determine whether the similarity between node v and its neighbor node u is greater than the threshold η; if true, go to 4.5); if false, go to 4.2); 4.5) Take neighbor node u as the sampled neighbor node u ′ , and added to the sampled neighbor node set of node v; 4.6) Mean aggregation sampling of neighbor nodes u ′ The k-1 layer embedding representation of node v generates the k-th layer neighbor node feature aggregation representation And the k-th layer neighbor node feature aggregation representation of node v and the k-1th layer embedding representation of node v After concatenation and linear transformation, the k-th layer embedding representation of node v is obtained Where 1≤k≤K, K is the number of layers of aggregated neighbor nodes, is the set of sampled neighbor nodes of node v, is the initial feature; 4.7) Embedding representation for the kth layer After normalization; 4.8) Loop through steps 4.2)-4.7) to embed the obtained Kth layer into as the node embedding representation of node v; 5) Based on the node embedding representation of each node, the malicious user detection result is obtained.
2. The method according to claim 1, characterized in that The attributes include: content-related attributes, activity-related attributes, emotion-related attributes and / or structure-related attributes.
3. The method according to claim 2, characterized in that Content-related attributes include: User's tweets.
4. The method according to claim 2, characterized in that Activity-related attributes include: number of tweets, number of retweets, number of fans, number of followings, number of favorites, number of tags, number of citations, number of URLs, average number of mentions per tweet, average time interval between tweets, and median time interval between tweets.
5. The method according to claim 2, characterized in that Sentiment-related attributes include: sentiment of the tweet and use of profanity.
6. The method according to claim 2, characterized in that The structure-related attributes include: centrality, eigenvector, in-degree, out-degree and content-related attributes of the user's first-order neighbors, activity-related attributes, sentiment-related attributes and structure-related attributes.
7. The method according to claim 1, characterized in that Methods for calculating initial features include: one-hot encoding, label encoding, or word embedding using GloVe.
8. The method according to claim 1, characterized in that Get the edges of the potential graph by following these steps: 1) Based on the initial feature Z of node v v and the initial feature Z of node u u , calculate the similarity between node v and node u S(v,u) = PearsonSimilarity(Z v ,Z u ), where v,u∈V, V is the node set of the social graph; 2) When the similarity S(v,u) is greater than a set threshold When , the edge between node v and node u is taken as the edge between node v and node u in the potential graph.
9. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Graph embedding method and device, and storage medium
CN109614975A
Work skill prediction method based on multi-graph neural network joint learning
CN111667158A