Inductive attribute graph anomaly detection method based on node-local subgraph mode and generative adversarial network
By employing a generator and discriminator to create and differentiate node-local subgraph pairs, the method addresses the limitations of existing attribute network anomaly detection, enhancing generalization and performance in both inductive and transductive scenarios by focusing on node-local subgraph relationships.
Patent Information
- Application Number
- CN202410052125.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
In the detection of attribute network anomaly, it is difficult to effectively identify abnormal nodes under inductive settings. Especially when new nodes are added, the model needs to be retrained. The traditional method fails to fully utilize the relationship between nodes and local subgraphs, resulting in insufficient generalization capabilities.
Generative adversarial network is adopted to generate exception node-local subgraph instance pairs through the generator. The discriminator distinguishes between the real and generated instance pairs. The training target is directly related to exception detection. The decoupled representation generates exception instances, and the generalization ability of the discriminator is improved through multiple rounds of sampled subgraphs.
Under the inductive setting, the abnormal nodes can be identified without retraining the model, and the relationship between nodes and local subgraphs is explicitly mined, which improves the generalization ability of unseen abnormal patterns. The generator and discriminator promote each other to form a node representation of distinction.
Smart Images

Figure CN120316655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph neural networks, and specifically to an inductive attribute graph anomaly detection method based on node-local subgraph patterns and generative adversarial networks. Background Art
[0002] The task of attribute network anomaly detection plays an important role in many applications such as financial fraud detection, network intrusion detection, and social spam detection. Many modern information systems can be abstracted as attribute networks, where entities in the information system are represented as nodes in the attribute network, and the relationships between entities are represented as edges in the attribute network. For example, a social network can be abstracted as an attribute network, where users represent nodes in the network, the posts or personal information of users are node attributes, and the follow or like relationships between users can be abstracted as edges between nodes. Attribute network anomaly detection aims to detect nodes that deviate from the normal pattern. Generally speaking, abnormal nodes can be divided into two types according to their abnormal patterns: context-abnormal nodes and structure-abnormal nodes. Context-abnormal nodes have attributes that are either noisy or inconsistent with their adjacent nodes, while structure-abnormal nodes often have abnormal connections with other nodes. Attribute network anomaly detection is expected to discover these potential abnormal patterns and thus identify abnormal nodes.
[0003] Due to the difficulty of obtaining true labels, most of the anomaly detection methods for attribute networks are designed and trained in an unsupervised manner, which can be mainly divided into matrix-based methods and autoencoder-based methods. Matrix-based methods use shallow mechanisms such as residual analysis or ego-network to detect anomalies. However, these methods are limited by their shallow representation learning mechanisms and cannot handle high-dimensional data. With the development of deep learning technology, researchers have been committed to developing deep learning-based algorithms, such as using autoencoders to reconstruct the attributes and structures of attribute networks and using reconstruction errors, residual analysis, or density estimation to evaluate anomalies. Although these reconstruction-based methods have achieved considerable success in anomaly detection of attribute graphs, they still have some inevitable drawbacks.
[0004] First, reconstruction-based methods have difficulty dealing with newly observed nodes because most of these methods use graph neural networks as the main component, and the spectral calculations in graph neural networks require prior access to the global network structure. When new data arrives, these models need to be retrained, which incurs high computational costs. Second, the training objective of autoencoder-based methods is to reconstruct data, which has no direct relation to anomaly detection. Therefore, the learned latent representation is a general summary of the latent patterns of the data, making it impossible to distinguish between normal and abnormal nodes and resulting in poor performance. Finally, the abnormality of a node is usually reflected in its relationship with its neighbors, rather than just the node attributes themselves. For ordinary nodes, there is a potential consistent relationship between the node and its neighbors. For abnormal nodes, there are various inconsistencies between the node attributes and its local substructure, and traditional autoencoder-based methods do not explicitly utilize the relationship between the node and its adjacent nodes. Based on the above three limitations, autoencoder-based methods are difficult to work effectively in the inductive setting. Anomaly detection in the inductive setting must meet the following two points. First, when new nodes are added, there is no need to retrain the entire model to calculate the anomaly values of the new nodes. Second, the model has good generalization performance for the possible anomaly patterns in unknown data. Although the patterns of normal nodes tend to be stable, the anomaly patterns may be very different from the training data. Therefore, it is necessary to improve the generalization ability of inductive anomaly detection methods for unseen anomaly patterns. Summary of the Invention
[0005] The object of the present invention is to make up for the deficiencies of the prior art and provide a method for detecting abnormal nodes in an attributed graph in the inductive setting. When new nodes are added, this method does not require repeated training. In addition, this method fully explores the relationship between a node and its local subgraph and has good generalization performance for unseen anomaly patterns.
[0006] To discover the abnormal patterns between a point and its local subgraph, a single training sample of the method of the present invention is a node-local subgraph instance pair. When a new node is added, only its local subgraph needs to be sampled from the original graph and input into the method of the present invention to obtain the anomaly score of the node. Overall, the architecture of this method is a generative adversarial network, where the generator is used to generate abnormal node-local subgraph instance pairs, and the discriminator is used to distinguish whether the input instance pair is an abnormal instance pair generated by the generator or a normal instance pair sampled from the real graph. During the training process, the generator generates a large number of potential abnormal instance pairs, which greatly improves the discriminator's ability to distinguish unseen anomalies; and the improvement of the discriminator's recognition ability can in turn promote the generator to generate more abundant and more conforming abnormal instance pairs to the real data manifold. In summary, a unified whole that promotes each other is formed between the generator and the discriminator. In addition, the output of the discriminator can be directly used to measure the anomaly of the node, which indicates that the training objective of this method is directly related to anomaly detection and helps normal points and abnormal points form a discriminative node representation.
[0007] The beneficial effects of the present invention are as follows:
[0008] (1) The present invention detects abnormal nodes by explicitly detecting the relationship between a node and its local subgraph, and explicitly discovers the abnormal patterns in the local structure of the node.
[0009] (2) The present invention uses a generative adversarial network to identify abnormal nodes, where the generator generates a large number of potential abnormal instance pairs, greatly improving the generalization ability of the discriminator for unseen abnormal patterns.
[0010] (3) The present invention uses decoupled representation to generate abnormal instance pairs, and simulates the camouflage of abnormal points in the real world by discovering the factors that are consistent between the local subgraph and the randomly initialized noise.
[0011] (4) The training objective of the present invention is directly related to the anomaly detection task, which helps the model form a discriminative representation of abnormal points and non-abnormal points. In addition, by sampling subgraphs in multiple rounds to obtain multiple instance pairs, the model can fully detect the local substructure of the node, and the experimental results prove the effectiveness of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The following further details the specific embodiments of the present invention with reference to the drawings.
[0013] Figure 1 It is a schematic diagram of the framework of the inductive attribute graph anomaly detection method based on node-local subgraph patterns and generative adversarial networks provided by the present invention.
[0014] Figure 2 It is the sampling method of the node-local subgraph instance pair adopted by the present invention.
[0015] Figure 3 It is the model framework diagram of the generator in the model of the present invention.
[0016] Figure 4 It is the model framework diagram of the discriminator in the model of the present invention. Specific embodiments
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] The present invention focuses on the attribute network anomaly detection task under the inductive setting. For the training attribute network G=(A, X), X is the node attribute matrix, and A represents the connection relationship between nodes, that is, the adjacency matrix. The model aims to output the anomaly score score(v') for each node v' in the test network G'=(X', A'), and then determine whether it is an abnormal node. The overall architecture of the model is as Figure 1 shown.
[0019] (1) Node-local subgraph sampling
[0020] First, traverse each node in the graph dataset, take this node as the target point, sample its 1 to 2-hop neighbor subgraphs around it, and discard some adjacent points according to the random sampling strategy to obtain a subgraph of a specific size. To prevent data leakage, we replace the attribute of the target point in the sampled subgraph with a 0 vector. Combine the target point with the sampled subgraph to obtain the instance pair P i =(v i , G i ), where G i =(A i , X i ), A i is the subgraph adjacency matrix, and X i is the subgraph attribute matrix. For each point in the graph dataset, repeat the above operations to generate node-local subgraph instance pairs for each node.
[0021] (2) Anomaly instance pair generation
[0022] The generator uses decoupled representation to generate anomaly instance pairs, improving the generalization ability of the discriminator to unknown anomaly patterns. By changing the attribute of the node v i =(v i , G i ) in the sampled normal instance pair P to generate the anomaly instance pair i To simulate various abnormal situations that may occur in the real world, first The attributes are initialized to Gaussian noise. However, Gaussian noise may not conform to the shape of the real dataset, and abnormal nodes in the real world may disguise themselves in various ways. For example, abnormal users in a social network disguise themselves by associating with normal users and posting content similar to that of normal users. Merely using Gaussian noise as the attributes of abnormal points cannot generate reasonable abnormal instance pairs. The method of the present invention generates the attributes of abnormal nodes by learning a decoupled representation for Gaussian noise.
[0023] Decoupled representation is a way of representing data features. Different from traditional feature representations, decoupled representation assumes that the generation of data is driven by M different factors. Accordingly, decoupled representation divides the features into M channels, and each channel represents a specific factor. The generation of decoupled representation mainly has the following steps: For each point in the sampled subgraph G i (the subscript i is omitted hereinafter), first, it is mapped to the decoupled representation z o = (z o,1 , z o,2 , …, z o,M ), where z o.m is the representation of node o in the subgraph on the m-th channel, representing the feature information driven by data factor m:
[0024] z o,m = ReLU(W m z o.m + b m ) (1)
[0025] Next, an iterative routing algorithm is used to generate a decoupled representation for the abnormal nodes initialized with noise. The iterative routing algorithm identifies the channels in the noise that are consistent with the information of the sampled subgraph, and then aggregates the neighbor information in that channel and retains the abnormal information of the noise in the remaining channels. Specifically, for the abnormal node v^ and the node u ∈ G in the subgraph, the acquisition of the decoupled representation is shown in formulas (2) and (3).
[0026]
[0027]
[0028] t = 1, 2, 3, … T represents the iteration rounds, represents the possibility that node v^ and node u have information consistency on the m-th channel in the t-th round. After multiple rounds of iteration, the decoupled representation of the abnormal node v^ is finally obtained
[0029] To generate abnormal node attributes consistent with the node dimensions in the graph data, a fully connected layer is used to perform dimensional conversion on the generated decoupled representation, and finally abnormal attributes x are generated. v^:
[0030] x v^ = ReLUCW h h v^ + b h ) (4)
[0031] (3) The discriminator determines whether the input instance pair is a normal instance pair
[0032] The discriminator aims to determine whether the input instance pair is a real instance pair sampled from real data or an abnormal instance pair generated by the generator. It outputs an anomaly score by evaluating the relationship between the nodes and local subgraphs in each input instance pair. Its goal is to make the scores of abnormal instance pairs close to 0 and the scores of normal instance pairs close to 1. The discriminator is composed as follows.
[0033] First, a multi-layer graph neural network (GNN) is used to aggregate the information between subgraph nodes and convert high-dimensional attributes into low-dimensional representations. In the first layer, the GNN transforms the node attributes in the subgraph as shown in formula (5).
[0034]
[0035] The input of the first layer of the multi-layer GNN is the sampled subgraph attribute matrix, and the output of the last layer is the final embedding of the graph nodes. In practice, GCN is used as the GNN module here because GCN has good structure and attribute modeling capabilities;
[0036] To identify the relationship between the nodes and sampled subgraphs in the input instance pair in the same feature space, the weight matrix and activation function of GCN are used to map the node attributes in the input instance pair into the same embedding space. The specific steps are shown in formula (6).
[0037]
[0038] where W l is the parameter matrix of the l-th layer of the GNN, and the input is the attribute of the nodes in the generated abnormal instance pair or sampled normal instance pair. For simplicity, use to represent the output of the L-th layer
[0039] After generating representations for the two elements of the input instance pair, an anomaly score is generated for the instance pair. First, the output matrix of the local subgraph is aggregated to generate a graph-level representation. The representations of all nodes in the local subgraph are aggregated using average pooling, as shown in formula (7).
[0040]
[0041] Among them, n i represents the number of nodes in the subgraph, represents the vector output of the multi-layer GNN at the k-th row of the last layer. To measure the abnormal consistency between the subgraph and the central point, and are concatenated and input into a multi-layer perceptron (MLP) to obtain an anomaly score, as shown in formula (8) specifically.
[0042]
[0043] (4) Joint training of the generator and discriminator
[0044] In the same iteration, the generator G and the discriminator D are trained separately, and the parameters of the discriminator are fixed when training the generator, and the parameters of the generator are fixed when training the discriminator. By receiving the gradient signal of the discriminator, the generator can generate various abnormal node-local subgraph instance pairs that conform to the real data manifold; by distinguishing the various abnormal instance pairs generated by the generator, the discriminator can fully learn various abnormal patterns. After multiple rounds of training, the performance of the discriminator and the generator can be improved simultaneously. The objective function of the generator is as shown in formula (9):
[0045]
[0046] The training objective of the discriminator is to correctly distinguish the abnormal instance pairs generated by the generator from the normal instance pairs sampled from the real data as much as possible, and its objective function is as shown in formula (10).
[0047]
[0048] (5) Calculate the anomaly score for the test node
[0049] After the model converges after multiple rounds of iteration, the discriminator D can distinguish the normal instance pairs sampled from the real data and the abnormal instance pairs generated by the generator. Since the anomaly of a node is usually reflected in its relationship with the local subgraph structure, the prediction score can be directly used to measure the anomaly of the node; since the input of the discriminator is the local substructure of the node and does not include all the adjacent subgraphs of the node, multiple rounds of instance pairs can be sampled for the target node through multiple rounds of sampling, and the anomaly score is calculated for each instance pair, and the average value of multiple rounds of sampling is taken as the final anomaly score of the node. The multiple rounds of sampling instance pairs of node v i can be expressed as (p i,1 , p i,2 , … p i,R ) where R represents the total number of sampling rounds; for the r-th round of sampling, the anomaly score d i,r is calculated according to the formula, and for node v iThe abnormal score is the average score of the R-round sampling instance pairs, as shown in formula (11).
[0050]
[0051] Specifically, to verify the effectiveness of the method of the present invention in a real system, this embodiment conducts experiments on the social network datasets BlogCatalog, Flickr and the paper citation dataset ACM. The specific details are as follows:
[0052] (1) Dataset
[0053] The present invention evaluates the proposed invention using three real-world datasets (including two social networks): BlogCatalog and Flickr of the blog directory, and an ACM citation network dataset. The first two social network datasets are from the blog sharing website BlogCatalog and the photo sharing website Flickr respectively. The nodes represent users, the node attributes are the blog or photo tags published by the users, and the edges represent the follow-up relationships between users. The nodes of ACM are the published articles, and the edges represent the citation relationships between these articles. The dataset statistical information is shown in Table 1.
[0054] Table 1 Dataset statistical information
[0055]
[0056] (2) Model details
[0057] To improve the computational efficiency, the number of local subgraph nodes is set to 5. For the generator, the number of channels M is set to 8, the dimension of the decoupled representation is set to 128, the prior distribution of the noise is set to a Gaussian prior, and the number of iterations T is set to 4. For the discriminator, GCN is used as the graph neural network, and the embedding dimension is set to 256. Regarding the optimization hyperparameters, the Adam optimizer is used, where the learning rate of BlogCatalog and Flickr is set to 0.001, and the learning rate of ACM is set to 0.0005. The exponential learning rate decay with a decay rate of 0.9 is used, and the training model rounds are 10 epochs. The batch size is set to 256. In the inference stage, the sampling round number R is set to 32 to obtain accurate anomaly detection results. For each dataset, the average value of the results of ten runs is taken as the reported result.
[0058] (3) Evaluation metrics
[0059] The present invention uses ROC-AUC as an evaluation metric. ROC-AUC is the area value under the ROC curve. The ROC curve depicts the comparison of the true positive rate (the true anomaly rate identified as an anomaly) and the false positive rate (the rate of normal samples identified as anomalies) based on the anomaly score and the true label at different classification thresholds. ROC-AUC measures the probability that a randomly sampled anomalous node has a higher suspicious score than a normal node. An AUC value close to 1 indicates excellent performance of the method.
[0060] Table 2 Results (AUC) under different training dataset sizes for the inductive setting
[0061]
[0062] Table 3 Results (AUC) of the proposed method and the baseline methods in the transductive setting
[0063]
[0064] (4) Comparison system
[0065] To verify the effectiveness of the present invention, the method of the present invention (Ours) is compared with a variety of current advanced and representative models, and the relevant models are introduced separately here.
[0066] ● (RCAE): RCAE [1] is an autoencoder-based method that uses the reconstruction error to measure the anomaly of samples. It improves the generalization ability to noise by introducing an additional residual term, thereby enhancing the robustness and inductive generalization ability of test samples.
[0067] ● (EGBAD): EGBAD 2] is an unsupervised anomaly detection method based on GAN. It includes a generator G for generating normal samples and an encoder E for mapping the generated samples to the latent space. For a test sample x, it uses the Euclidean distance between x and G(E(x)) as the anomaly scoring function.
[0068] ● (FRAUDURE): FRAUDURE [3] is a supervised method that considers attribute, topology, and relationship inconsistencies. In particular, it uses graph convolutional operations to obtain node representations and uses the scalar output by the linear layer as the anomaly score.
[0069] ● (DOMINANT): DOMINANT [4] uses graph convolutional operations to reconstruct the adjacency matrix and attribute matrix of the network. It uses the weighted sum of the reconstruction error to evaluate the anomaly of nodes. Different from the above three anomaly detection methods for the inductive setting, it can only be used in the transductive setting.
[0070] The experimental results of the above baseline model and the model proposed in the present invention under the inductive setting and the transductive setting are shown in Tables 2 and 3. The performance of the method of the present invention under different numbers of training samples is measured by sequentially changing the size of the training data.
[0071] Compared with the baseline model, even with less training data, the model proposed in the present invention has achieved excellent performance under the inductive setting. A large part of the performance improvement is due to the fact that the generator generates a large number of potential abnormal instance pairs to prompt the discriminator to learn various abnormal patterns from them. The method of the present invention has better performance compared with the autoencoder-based method RACE. There may be two reasons: the method of the present invention mines the rich relationships between each node and its local subgraph and learns various abnormal patterns from them; the training objective of the method proposed in the present invention is highly relevant to anomaly detection, making the model more robust to noise. In addition, the present invention also evaluates the model proposed in the present invention under the transductive setting, and the results are shown in Table 3. The performance of the present invention on the three datasets is better than that of the baseline model. Although the present invention is designed for the inductive anomaly detection scenario, its performance under the transductive setting is also very strong. Compared with DOMINANT that uses GNN to reconstruct the network, the proposed framework has better performance in the three benchmark tests. It can be seen that the present invention has learned better representations under the guidance of the training objective directly related to anomaly detection.
[0072] The above content is intended to schematically illustrate the technical solution of the present invention, and the present invention is not limited to the embodiments described above. Without departing from the spirit and scope protected by the claims of the present invention, those of ordinary skill in the art can also make many specific transformations in various forms under the inspiration of the present invention, and these all belong to the protection scope of the present invention.
[0073] References
[0074] [1]Chalapathy, R., Menon, A.K., Chawla, S., 2017. Robust, deep and inductive anomaly detection, in: Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer.pp.36 - 51.
[0075] [2]Zenati, H., Foo, C.S., Lecouat, B., Manek, G., Chandrasekhar, V.R., 2018. Efficient gan-based anomaly detection. arXiv preprint arXiv:1802.06222.
[0076] [3]Zhang, G., Wu, J., Yang, J., Beheshti, A., Xue, S., Zhou, C., Sheng, Q.Z., 2021. Fraudre: fraud detection dual-resistant to graph inconsistency and imbalance, in: 2021 IEEE International Conference on Data Mining (ICDM), IEEE.pp. 867-876.
[0077] [4]Ding, K., Li, J., Bhanushali, R., Liu, H., 2019a. Deep anomaly detection on attributed networks, in: Proceedings of the 2019 SIAM International Conference on Data Mining, SIAM.pp. 594-602
Claims
1. An inductive attribute graph anomaly detection method based on node-local subgraph patterns and generative adversarial networks, characterized in that Including the following steps: (1) Using the subgraph sampling technique to sample node-local subgraph instance pairs for the nodes in the graph. Since the anomalies of nodes are usually reflected in the relationship between the nodes and their local subgraphs, by sampling node-local subgraph pairs, the model can learn the potential anomaly patterns between the nodes and their surrounding local subgraphs. For non-anomalous nodes, the attributes of the nodes usually show a certain consistency with the attributes and structures of their local subgraphs. For anomalous nodes, there are various potential anomaly patterns between the nodes and their local anomalous subgraphs. (2) The generator generates the corresponding anomalous node attributes for each local subgraph sampled in (1) to generate anomalous instance pairs, that is, anomalous instance pairs composed of the sampled local subgraphs and the generated anomalous nodes. The specific operations are as follows: First, use Gaussian noise to simulate various anomalous attributes that may appear in the real world, and then use a graph neural network to transform the Gaussian noise into node attributes that conform to the data flow manifold of the current sampled subgraph attributes. (3) The discriminator is used to distinguish whether the input instance pair is an anomalous instance pair generated by the generator or a normal instance pair sampled from the local subgraph. Specifically, the discriminator uses a graph neural network to generate potential representations for the subgraph and the node of the input instance pair respectively, and calculates the anomaly score of the instance pair based on this representation. After joint training with the generator, the discriminator can identify various potential anomaly patterns. (4) The generator and the discriminator are jointly trained. By receiving the gradient signal of the discriminator, the generator can generate various anomalous point-subgraph instance pairs that conform to the real data flow manifold; by distinguishing the various anomalous instance pairs generated by the generator, the discriminator can fully learn a variety of anomaly patterns. After multiple rounds of training, the performance of the discriminator and the generator can be improved simultaneously.
2. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the inductive attribute graph anomaly detection method based on node-local subgraph patterns and generative adversarial networks described in claim 1.
3. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the inductive attribute graph anomaly detection method based on node-local subgraph patterns and generative adversarial networks described in claim 1.