Graph data processing method, system and equipment and medium
By using the combination method of the abnormality diagram to be purified and the reference abnormality diagram in graph abnormality detection, the abnormality detection model is trained and the node comparison score is calculated, and the interference edges are identified and deleted, the problem of low detection accuracy in the existing technology is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510159531.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing graph abnormality detection method based on comparison learning reduces the detection accuracy due to the existence of interfering edges, and cannot effectively capture complex relationships in the graph data.
By simultaneously inputting the anomaly diagram to be purified and the reference anomaly diagram, randomly walking and extracting the subgraphs of the nodes, building positive instance pairs and negative instance pairs, using the similarity scores of these instance pairs to train the model, compute the node comparison score, identify and delete interfering edges, and iterate and optimize until the set threshold is reached.
The recognition accuracy of the anomaly detection model for interfering edges is improved, the robustness and detection capabilities of the model are enhanced, and the accuracy of interfering edge detection and deletion is gradually improved.
Smart Images

Figure CN119992134A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of anomaly detection, and specifically relates to a graph data processing method, system, device and medium. Background Art
[0002] Graphs are a fundamental data structure consisting of nodes and edges that play a vital role in representing relationships in diverse disciplines, such as recommender systems, social network analysis, and financial risk assessment. In graph analytics, graph anomaly detection (GAD) has become a key research area, which aims to identify patterns that are significantly different from the majority. The detection of such anomalies reveals potential irregularities in the data and facilitates proactive intervention to protect data integrity. This capability has far-reaching applications, especially in fraud detection, discovery of brain pathology mechanisms, and network intrusion prevention.
[0003] Early methods used shallow mechanisms such as ego-network analysis, residual analysis, and CUR decomposition to detect anomalies. However, these methods cannot capture the complex relationships inherent in graph data, which limits their ability to detect complex anomalies. With the rise of deep learning, some studies have used graph neural networks (GNNs) to reconstruct structures and node features. They use reconstruction errors as the basis for anomaly identification. Although these methods have made progress, they require explosive memory resources. In addition, convolution operations in GNNs may smooth out anomaly signals and reduce the uniqueness of abnormal nodes, thereby affecting detection accuracy. Recently, researchers have adopted contrastive learning for anomaly detection. This method uses restarted random walks (RWR) for subgraph sampling, generates positive instance pairs between nodes and their local subgraphs, and generates negative instance pairs from different subgraphs. The degree of abnormality of nodes is evaluated by examining the similarity difference between positive and negative instance pairs.
[0004] However, existing contrastive learning-based methods have significant flaws: contrastive learning relies on the similarity of positive instance pairs and the similarity of negative instance pairs for effective model training, but the presence of interference edges violates this assumption. This problem stems from the presence of anomalies in the graph - feature anomalies and structural anomalies. Feature anomalies manifest themselves in changes in node features, and these distortions propagate through edges connecting nodes with abnormal features, destroying the semantic consistency of the subgraph. On the other hand, structural anomalies occur when abnormal edges connect previously unconnected nodes, changing the inherent topology of the subgraph. Since these two types of interference edges inevitably disrupt the process of generating positive instance pairs, this distortion impairs the model's ability to effectively learn normal patterns, ultimately reducing its ability to distinguish between normal and abnormal data.
[0005] In summary, the existence of interference edges in existing graphs reduces the accuracy of graph anomaly detection. Summary of the invention
[0006] In order to overcome the above-mentioned deficiencies in the prior art, the present invention provides a graph data processing method, comprising the following steps:
[0007] The anomaly graph to be purified and the benchmark anomaly graph are simultaneously input into the anomaly detection model, and the anomaly detection model is trained, specifically: using the anomaly detection model to perform random walks on the anomaly graph to be purified and the benchmark anomaly graph at the same time, extracting subgraphs of nodes from the anomaly graph to be purified and the benchmark anomaly graph respectively; the node and its own subgraph constitute a first positive instance pair, and the node and the subgraph of other nodes constitute a first negative instance pair, the node and the updated node representation in the subgraph constitute a second positive instance pair, and the node and the updated node representation in other subgraphs constitute a second negative instance pair; the anomaly detection model is trained according to the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the benchmark anomaly graph;
[0008] Using the trained anomaly detection model, obtain the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair of the benchmark anomaly graph and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculate the node comparison score according to the difference;
[0009] The edge interference score of the abnormal graph to be cleaned is calculated by node comparison score, the edge interference scores are sorted, the first K edges after sorting are identified as interference edges, and the interference edges are deleted to obtain a roughly clean graph;
[0010] The rough clean graph and the benchmark anomaly graph are simultaneously input into the anomaly detection model, and the training process of the anomaly detection model, the node comparison score calculation process and the interference edge deletion process are repeatedly iterated until the number of deleted interference edges reaches a set threshold, thereby obtaining a target clean graph.
[0011] Preferably, the trained anomaly detection model is used to obtain the difference between the similarity score of the first negative instance pair of the benchmark anomaly graph and the similarity score of the first positive instance pair, and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and the node comparison score is calculated based on the difference, specifically: the anomaly score of the subgraph and node comparison scale is determined based on the difference between the similarity score of the first negative instance pair of the benchmark anomaly graph and the similarity score of the first positive instance pair, and the anomaly score of the node and node comparison scale is determined based on the difference between the similarity score of the second negative instance pair of the benchmark anomaly graph and the similarity score of the second positive instance pair; the node comparison score is calculated using the anomaly score of the subgraph and node comparison scale and the anomaly score of the node and node comparison scale.
[0012] Preferably, the node comparison score is calculated by using the anomaly score of the subgraph and node comparison scale and the anomaly score of the node and node comparison scale, specifically: weighted summation of the anomaly score of the subgraph and node comparison scale and the anomaly score of the node and node comparison scale is performed to obtain the node comparison score.
[0013] Preferably, the node comparison score is obtained by the following formula:
[0014]
[0015] In the formula, represents the comparison score of the node in round r, represents the mean comparison score of the node, and R represents the total number of rounds.
[0016] Preferably, the abnormal graph to be purified and the reference abnormal graph are simultaneously input into the abnormal detection model to train the abnormal detection model, and the loss function used is:
[0017]
[0018] In the formula, represents the comparison loss between nodes and subgraphs in the baseline anomaly graph, represents the contrast loss between nodes and subgraphs in the rough clean graph, represents the node-to-node comparison loss in the baseline anomaly graph, represents the contrast loss between nodes in the rough and clean graph, α and β represent the balance parameters for different views and scales, respectively.
[0019] The present invention also provides a graph data processing system, comprising:
[0020] A model training module is used to input the anomaly graph to be purified and the benchmark anomaly graph into the anomaly detection model at the same time, and train the anomaly detection model, specifically: using the anomaly detection model to perform random walks on the anomaly graph to be purified and the benchmark anomaly graph at the same time, extracting subgraphs of nodes from the anomaly graph to be purified and the benchmark anomaly graph respectively; forming a first positive instance pair with the node and its own subgraph, and forming a first negative instance pair with the subgraphs of other nodes, forming a second positive instance pair with the node and the updated node representation in the subgraph, and forming a second negative instance pair with the updated node representation in other subgraphs; training the anomaly detection model according to the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the benchmark anomaly graph;
[0021] A node comparison score acquisition module is used to use the trained anomaly detection model to obtain the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair of the benchmark anomaly graph and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculate the node comparison score according to the difference;
[0022] The interference edge deletion module is used to calculate the edge interference score of the abnormal graph to be cleaned through the node comparison score, sort the edge interference scores, identify the first K edges after sorting as interference edges, and delete the interference edges to obtain a roughly clean graph;
[0023] The target clean graph acquisition module is used to input the rough clean graph and the benchmark anomaly graph into the anomaly detection model at the same time, and repeat the iterative operation of the anomaly detection model training process, the node comparison score calculation process and the interference edge deletion process until the number of deleted interference edges reaches a set threshold, thereby obtaining the target clean graph.
[0024] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the graph data processing method.
[0025] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the graph data processing method.
[0026] The graph data processing method, system, device and medium provided by the present invention have the following beneficial effects:
[0027] The present invention simultaneously inputs an anomaly graph to be purified and a reference anomaly graph into an anomaly detection model, and by using the anomaly detection model to perform random walks on the anomaly graph to be purified and the reference anomaly graph at the same time, subgraphs of nodes can be extracted from the anomaly graph to be purified and the reference anomaly graph respectively; by extracting subgraphs from the anomaly graph to be purified, the anomaly detection model can be prevented from being affected by interference edges, the ability of the anomaly detection model to judge interference edges is increased, and the process of deleting interference edges is made more accurate; extracting subgraphs from the reference anomaly graph can enable the model to learn rich information and enhance the robustness of the model; the node and the subgraph constitute a first positive instance pair, the subgraph obtained from other nodes constitute a first negative instance pair, the node and the updated node representation in the subgraph constitute a second positive instance pair, and the node and the updated node representation in the other subgraph constitute a second negative instance pair; the anomaly detection model is trained by the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the reference anomaly graph, and the anomaly detection model can be used to obtain a better detection result. By comparing nodes with subgraphs and nodes with nodes, the anomaly detection model can effectively extract feature anomalies and structural anomaly information in the graph, thereby improving the accuracy of the anomaly detection model in detecting interference edges; the node comparison score is calculated by the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair, and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, which can improve the accuracy of the node comparison score calculation; the edge interference score of the abnormal graph to be purified is calculated by the node comparison score, the edge interference score is sorted, and the first K edges after sorting are deleted, so that the most destructive interference edges can be deleted; the interference edges are deleted by using multiple iterations. In each iteration, after deleting a group of interference edges, the parameters of the anomaly detection model are reset to prevent deviations in the previously deleted edges. At the same time, recalculating the node comparison score can enhance the recognition of high-interference edges, so that the detection and deletion processes of interference edges gradually become more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0029] Figure 1 A flowchart of a graph data processing method according to an embodiment of the present invention;
[0030] Figure 2 It is the balance parameter analysis diagram of scale and view, where: Figure 2 (a) and Figure 2 (b) are the balance parameter analysis diagrams of scale and view in the datasets Citeseer and Citation respectively;
[0031] Figure 3 These are the renderings of different edge deletion ratios;
[0032] Figure 4 This is the analysis chart of the interference edge removal accuracy of different models, where: Figure 4 (a) and Figure 4 (b) shows the edge removal accuracy analysis diagram on the Cora and ACM datasets respectively. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.
[0034] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the technical solutions of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0035] In addition, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance. In the description of the present invention, it should be noted that, unless otherwise clearly specified or limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. In the description of the present invention, unless otherwise specified, "plurality" means two or more, which will not be described in detail here.
[0036] Example
[0037] The present invention provides a graph data processing method, specifically Figure 1 As shown, the following steps are included:
[0038] Step 1: Input the anomaly graph to be purified and the benchmark anomaly graph into the anomaly detection model at the same time, and train the anomaly detection model. Specifically, use the anomaly detection model to perform random walks on the anomaly graph to be purified and the benchmark anomaly graph at the same time, and extract subgraphs of nodes from the anomaly graph to be purified and the benchmark anomaly graph respectively; form a first positive instance pair with the node and its own subgraph, and form a first negative instance pair with the subgraph obtained from other nodes; form a second positive instance pair with the node and the updated node representation in the subgraph, and form a second negative instance pair with the updated node representation in other subgraphs; train the anomaly detection model according to the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the benchmark anomaly graph.
[0039] In the present invention, the anomaly detection model used is ANEMONE, and its training process is as follows:
[0040] (1) Sub-graph sampling.
[0041] The abnormal level of a node is closely related to the structure and characteristics of its local subgraph. Normal nodes usually have a high degree of similarity with adjacent nodes, while abnormal nodes exhibit the opposite behavior. Therefore, the present invention adopts random walk RWR to sample subgraphs Ga and Gc from abnormal graphs and clean graphs respectively, each subgraph includes a target node and adjacent nodes, and the size is fixed to N. In order to prevent the abnormal detection model from easily identifying the target node in the subgraph, the present invention initializes the feature embedding of the target node in the subgraph to a zero vector to hide the information.
[0042] (2) Node-subgraph comparison. The purpose of node-subgraph comparison is to learn the consistency between a node and its entire local subgraph. To this end, the present invention constructs instance pairs by combining nodes with subgraphs. Specifically, the first positive instance pair consists of the node ui and its subgraph Gi, while the first negative instance pair is formed by pairing the node ui with a randomly selected subgraph Gj. In order to extract meaningful features from the graph, the present invention adopts a graph convolutional network GCN, which uses the graph structure to aggregate information and capture local dependencies. Through this process, each node in the subgraph integrates features from its neighboring nodes:
[0043]
[0044] in, and are the representations of the (l+1)th layer and the lth layer respectively, is the inverse square root of the subgraph degree matrix, is the adjacency matrix of the subgraph of node i, is the weight matrix of the lth layer.
[0045] Generate an embedding of a sub-graph by average pooling:
[0046]
[0047] Among them, S i represents the multidimensional subgraph embedding obtained through GCN, p represents the embedding of the pth node in the subgraph, N represents the number of nodes, and h i Represents the sub-graph embedding after pooling.
[0048] To compare subgraph embeddings and node embeddings, we align them to the same dimension. In practice, we use a multi-layer perceptron (MLP) to project node features into a low-dimensional space:
[0049]
[0050] in, and denote the representations of the lth layer and the (l+1th layer), respectively. is the weight matrix of the lth layer.
[0051] Subsequently, a discriminator is used to measure the relationship between node embeddings and subgraph embeddings. The discriminator evaluates the similarity between embeddings in positive and negative instance pairs, producing a score for each instance pair. Specifically, a bilinear function is used to quantify the similarity:
[0052]
[0053] Among them, h i and h j Represents the subgraph embedding of node i and node j respectively, n i represents the node embedding of node i.
[0054] The node-subgraph comparison is trained by pulling the first positive instance pair closer and pushing the first negative instance pair away based on similarity, which can be regarded as a binary classification task. Therefore, the present invention adopts the binary cross entropy (BCE) loss function for training:
[0055]
[0056] Among them, y i Represents the label of the instance pair, s i Represents the instance pair score.
[0057] (3) Node-node comparison. Node-node comparison is more effective in identifying node-level anomalies. Each instance pair consists of the original feature embedding and the updated node embedding. The updated node embedding is aggregated from adjacent nodes within the subgraph. The second positive instance pair is established by combining the original feature embedding with the corresponding updated embedding. On the other hand, the second negative instance pair will be generated by pairing the original node embedding with the updated node embedding from an unrelated subgraph. The aggregation of the subgraph is done through a new GCN:
[0058]
[0059] in, and are the representations of the (l+1)th layer and the lth layer respectively, is the inverse square root of the subgraph degree matrix, is the adjacency matrix of the subgraph of node i, is the weight matrix of the lth layer.
[0060]
[0061] in, and denote the representations of the lth layer and the (l+1th layer), respectively. is the weight matrix of the lth layer.
[0062] Step 2: Use the trained anomaly detection model to obtain the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair of the benchmark anomaly graph, and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculate the node comparison score based on the difference; calculate the edge interference score of the anomaly graph to be purified by the node comparison score, sort the edge interference scores, identify the first K edges after sorting as interference edges, and delete the interference edges to obtain a rough clean graph. The rough clean graph and the benchmark anomaly graph are simultaneously input into the anomaly detection model, and the training process of the anomaly detection model, the node comparison score calculation process, and the interference edge deletion process are repeatedly iterated until the number of deleted interference edges reaches the set threshold, and the target clean graph is obtained.
[0063] (1) Node comparison score. For normal nodes, the similarity of positive instance pairs approaches 1, while the similarity of negative instance pairs approaches 0. On the contrary, for abnormal nodes, the scores of positive and negative instance pairs are relatively close. Therefore, the present invention evaluates the abnormality of a node by checking the similarity difference between positive and negative instance pairs. Specifically, the present invention subtracts the similarity score of the positive and negative instances from the negative instance pair. Both scores come from the baseline anomaly graph. The greater the difference, the higher the possibility of anomaly:
[0064]
[0065] In the formula, represents the first negative instance pair score of the benchmark anomaly map, represents the first positive instance pair fraction of the benchmark anomaly map, represents the second negative instance pair score of the benchmark anomaly map, represents the second positive instance pair fraction of the benchmark anomaly graph, represents the node-subgraph comparison score of the benchmark anomaly graph, Represents the node-node comparison score of the benchmark anomaly graph.
[0066] To comprehensively assess anomalies at multiple scales, we aggregate these scores by weighted summation:
[0067]
[0068] Where β is the balance parameter of different scales.
[0069] (2) Interference edge detection. Interference edges significantly reduce the similarity between the target node and its local subgraph, narrowing the score gap between positive and negative instance pairs. This hinders the ability of the contrastive loss function to distinguish instances and undermines the effective training of the model. To alleviate this problem, the first step of the present invention is to accurately quantify the interference level of each edge.
[0070] Interference edges can be divided into two categories: (1) abnormal links formed between normal nodes that were not previously connected; (2) pre-existing edges connecting nodes with abnormal characteristics. Both categories represent connections between dissimilar entities. Therefore, the reduction in the contrast score of a node can be attributed to its association with dissimilar nodes through interference edges. If an edge connects two nodes with high scores, it indicates that both nodes interact with different neighbors, increasing the possibility of interference. Based on this assumption, the present invention aggregates the scores of connected nodes to quantify the interference level associated with the edge. Through this method, the present invention obtains the detection matrix of interference edges:
[0071]
[0072] Based on the edge interference scores, the present invention applies a ranking mechanism to prioritize the edges that are most disruptive to the contrastive learning process. Specifically, the edges with the highest scores are identified for removal. These edges typically exhibit significant deviations in feature distribution compared to adjacent nodes, indicating that they are likely to become sources of interference. The present invention iteratively removes interference edges. In each iteration, after removing a set of edges, the parameters of the model are reset to prevent influence from previously removed edges. The model then undergoes another round of anomaly detection model training, recalculating node scores to enhance the identification of high-interference edges. By effectively evaluating sources of interference and removing interference edges, the detection process becomes increasingly accurate. By removing these target edges, the present invention obtains a clean graph that reduces the impact of interference edges on subsequent training.
[0073] (3) Loss function. The present invention performs multi-scale anomaly perception on the baseline anomaly image and the rough clean image, and the contrastive learning loss function is defined as follows:
[0074]
[0075] In the formula, represents the comparison loss between nodes and subgraphs in the baseline anomaly graph, represents the contrast loss between nodes and subgraphs in the rough clean graph, represents the node-to-node comparison loss in the baseline anomaly graph, represents the contrast loss between nodes in the rough and clean graph, α and β represent the balance parameters for different views and scales, respectively.
[0076] (4) Multi-scale comparative scoring. Clean graphs can mask true anomaly signals and undermine accurate anomaly detection. Therefore, in the score calculation process, the present invention only relies on the scores calculated from the baseline anomaly graph. In addition, due to the inherent randomness in sampling, structural anomaly nodes may occasionally connect to normal neighbors, resulting in missed anomalies. To solve this problem, the present invention performs multiple sampling iterations to calculate the mean and variance:
[0077]
[0078] In the formula, represents the comparison score of the node in round r, represents the mean comparison score of the node, and R represents the total number of rounds.
[0079] (5) After obtaining the target clean graph, the benchmark anomaly graph is detected using the anomaly detection model, specifically: the detection array is initialized with all elements set to zero. For each node that is the endpoint of the interference edge, the present invention increments the value of the corresponding position in the array. During the scoring period, the detection score is calculated iteratively, and its value is proportional to the frequency of the node being connected to the interference edge. A higher score reflects a stronger correlation between the node and the abnormal behavior. Finally, the total node score is calculated by aggregating the multi-scale comparison score and the detection score, as follows:
[0080]
[0081] In the formula, represents the node comparison score, represents the detection score, γ
[0082] represents the fractional balance parameter.
[0083] The present invention also provides a graph data processing system, comprising:
[0084] A model training module is used to input the anomaly graph to be purified and the benchmark anomaly graph into the anomaly detection model at the same time, and train the anomaly detection model, specifically: using the anomaly detection model to perform random walks on the anomaly graph to be purified and the benchmark anomaly graph at the same time, extracting subgraphs of nodes from the anomaly graph to be purified and the benchmark anomaly graph respectively; forming a first positive instance pair with the node and its own subgraph, and forming a first negative instance pair with the subgraphs of other nodes, forming a second positive instance pair with the node and the updated node representation in the subgraph, and forming a second negative instance pair with the updated node representation in other subgraphs; training the anomaly detection model according to the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the benchmark anomaly graph;
[0085] A node comparison score acquisition module is used to use the trained anomaly detection model to obtain the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair of the benchmark anomaly graph and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculate the node comparison score according to the difference;
[0086] The interference edge deletion module is used to calculate the edge interference score of the abnormal graph to be cleaned through the node comparison score, sort the edge interference scores, identify the first K edges after sorting as interference edges, and delete the interference edges to obtain a roughly clean graph;
[0087] The target clean graph acquisition module is used to input the rough clean graph and the benchmark anomaly graph into the anomaly detection model at the same time, and repeat the iterative operation of the anomaly detection model training process, the node comparison score calculation process and the interference edge deletion process until the number of deleted interference edges reaches a set threshold, thereby obtaining the target clean graph.
[0088] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the graph data processing method. The present invention also provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is suitable for the processor to load to execute the graph data processing method.
[0089] In order to verify the method of the present invention, the present invention has carried out experimental verification, and the specific process is as follows:
[0090] (1) Dataset. The present invention conducts experiments on five benchmark datasets: Cora, Citeseer, ACM, DBLP and Citation. Since these datasets themselves do not contain abnormal information, the present invention manually injects anomalies. The present invention injects structural anomalies and feature anomalies into each dataset. Specifically, in order to inject feature anomalies, the present invention randomly selects n' abnormal nodes. For each abnormal node u, the present invention randomly selects 50 candidate nodes, and calculates the Euclidean distance between u and each candidate node, selects the candidate node v with the largest distance, and replaces the feature xu of node u with xv. For structural anomalies, the present invention randomly selects 15 nodes and connects them two by two to form a cluster, and the process is repeated m' times. Two types of anomalies are injected in equal numbers.
[0091] Table 1 Anomaly detection results of different models
[0092]
[0093] (2) In the present invention, for the convenience of description, the graph data processing method of the present invention is referred to as CVGAD, and the present invention selects eight state-of-the-art models as benchmarks. HCM determines anomalies based on the minimum number of hops between nodes, where nodes with a larger number of hops are more likely to be considered anomalies. CoLA, ANEMONE, GRADATE, and NLGAD are four anomaly detection models based on contrastive learning. Sub-CR and SL-GAD combine attribute reconstruction and contrastive learning to comprehensively evaluate the abnormality level of nodes. GFCN uses a graph convolutional network with skip connections to estimate the possibility of anomalies.
[0094] (3) Evaluation index: In order to evaluate the effectiveness, the present invention uses ROC-AUC as the evaluation index.
[0095] (4) Parameter setting. The present invention adopts a single-layer GCN to aggregate information from subgraphs, and both subgraph embeddings and node embeddings are mapped to 64 dimensions. The size of the subgraph is set to 4, and the learning rate is kept fixed at 0.001. In addition, the present invention sets the γ value to 0.8. Five iterations are performed on all datasets. Specifically, the present invention performs 500 rounds of edge removal training process on Cora and CiteSeer, while on DBLP, Citation and ACM, the present invention performs 1000 epochs. In the optimization training phase, the present invention performs 200 rounds on Cora and CiteSeer, and 400 rounds on DBLP, Citation and ACM. In addition, the present invention implements 300 rounds of score calculation.
[0096] (5) Anomaly detection results. In order to evaluate the anomaly detection performance of the present invention, it is compared with eight baseline models on five datasets. Based on the results shown in Table 1, the present invention can draw the following conclusions: (1) The present invention achieves AUC gains of 1.91%, 0.31%, 1.20% and 1.73% on Cora, Citeseer, Citation and ACM, respectively. In addition, the present invention also achieves the second best performance on DBLP, which proves the effectiveness of the model; (2) The performance of the method based on contrastive learning exceeds that of HCM and GFCN. The main reason is that the similarity between abnormal nodes and local subgraphs is low. Therefore, by comparing the similarity between nodes and subgraphs, the present invention can effectively identify abnormal nodes; (3) Sub-CR shows good performance, mainly because other baseline models implement a single learning strategy, while Sub-CR combines reconstruction and contrastive learning strategies. However, by mitigating the impact of interference edges, the present invention achieves better results than Sub-CR on four datasets.
[0097] Table 2 Abnormal detection effects of the present invention and its variants
[0098]
[0099] (6) Ablation study. In order to verify the effectiveness of different modules, the present invention conducts three types of ablation studies. First, the present invention studies the effect of different edge removal strategies. Next, the present invention tests the impact of different view random walks. Finally, the present invention compares the multi-score integration method with the score based only on contrastive learning. The results are shown in Table 2, and the specific analysis is as follows:
[0100] (a) Edge removal and iteration strategy. CVGAD sim Edge removal is performed based on the similarity score of the original features. CVGAD oreImplement non-iterative edge removal. Experimental results show that the performance of the present invention is better than the other three models. First, it proves that the performance of the present invention is better than that of the similarity-based method. Edges with low similarity are not necessarily interference edges. Incorrectly deleting normal edges will destroy the original topological structure of the graph, thereby reducing the effectiveness of detection. In addition, the research results also verify that iterative edge removal is more effective than single edge removal. This is because single-shot edge removal methods may be susceptible to interference edges, which hinders the identification of abnormal nodes and interference edges. This interference continues to affect the training process of instance pairs and reduces the overall accuracy. The iterative process in the present invention allows a small portion of edges to be removed in each round, gradually refining the graph while retaining normal patterns. This approach creates a synergistic effect between multi-scale anomaly perception and progressive purification, thereby improving the overall performance. (b) Multi-view strategy. CVGAD bano and CVGAD bcla The module performs RWR on two baseline anomaly maps and two roughly clean maps. ocla Only walk on the coarse clean graph. The results show that random walks on both views (the coarse clean graph and the baseline anomaly graph) produce better performance. By performing random walks on the coarse clean graph, the sampling of interference edges is avoided, thereby enhancing the similarity between a node and its subgraph. This improves the effectiveness of the contrastive learning process. On the other hand, performing random walks on the baseline anomaly graph helps the model learn diverse and informative features. (c) Multi-score integration strategy. CVGAD con The score of is determined only by the node comparison score. The results show that the combined score has better performance. This is because nodes connected with interference edges are more likely to be anomalies. The detection score can distinguish the degree of anomalies more finely.
[0101] (7) Parameter Analysis. Balance parameters α and β. The present invention studies the influence of view balance parameter α and scale balance parameter β, such as Figure 2 As shown, Figure 2 (a) and Figure 2 (b) are the scale and view balance parameter analysis diagrams of the Citeseer and Citation datasets. As can be seen from the figure, Figure 2 (a) and Figure 2 The performance of (b) shows an initial upward trend and then a downward trend. The other data sets also show the same phenomenon. This shows that the effect of multi-scale learning and multi-view training is better than other methods. According to the results, the present invention sets the α of Cora, Citeseer, DBLP, Citation and ACM to 0.8, 0.4, 0.4, 0.4 and 0.6. At the same time, the β values of the present invention on these data sets are set to 0.6, 0.6, 0.4, 0.6 and 0.6.
[0102] Figure 3 Indicates the influence of percentage K on detection performance. K ranges from 0.005 to 0.02 with an increment of 0.005. As the edge removal rate increases, the curve first rises and then falls. This shows that removing interference edges can improve the performance of the model, while removing too many interference edges will destroy the original topological structure of the graph, thus having a negative impact on performance. According to the results, the present invention sets K to 0.01 for Cora, Citeseer and Citation, and sets K to 0.015 for DBLP and ACM.
[0103] (8) Accuracy of different edge removal methods. RGSE detects abnormal links. The present invention records the proportion of interference edges removed in each iteration. sim and CVGAD gcn , the present invention selects the edge with the lowest score and records the accuracy. rgse ,The present invention deletes the edges with higher scores and keeps the same proportion as other ,methods. Figure 4 This is the analysis chart of the interference edge removal accuracy of different models, where: Figure 4 (a) and Figure 4 (b) are the edge removal accuracy analysis diagrams of the present invention on the Cora and ACM datasets. It can be seen from the figure that the present invention shows significantly superior edge removal accuracy. The reason is that the present invention effectively integrates feature information and graph structure to identify interference edges. In contrast, CVGAD sim Only focusing on feature-level information, CVGAD gcn Smoothing abnormal signals. RGSE,mainly identifies long path links, hub links and adversarial links.,In the dataset of the present invention, interference edges often appear in the form of clusters,,which limits the ability of RGSE to effectively identify them.
[0104] The above embodiments are only preferred specific implementation methods of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of the technical solutions that can be obviously obtained by any technician familiar with the field within the technical scope disclosed in the present invention belong to the protection scope of the present invention.
Claims
1. A graph data processing method, characterized in that: The steps include: The anomaly graph to be purified and the benchmark anomaly graph are simultaneously input into the anomaly detection model, and the anomaly detection model is trained, specifically: using the anomaly detection model to perform random walks on the anomaly graph to be purified and the benchmark anomaly graph at the same time, extracting subgraphs of nodes from the anomaly graph to be purified and the benchmark anomaly graph respectively; the node and its own subgraph constitute a first positive instance pair, and the node and the subgraph of other nodes constitute a first negative instance pair, the node and the updated node representation in the subgraph constitute a second positive instance pair, and the node and the updated node representation in other subgraphs constitute a second negative instance pair; the anomaly detection model is trained according to the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the benchmark anomaly graph; Using the trained anomaly detection model, obtain the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair of the benchmark anomaly graph and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculate the node comparison score according to the difference; The edge interference score of the abnormal graph to be cleaned is calculated by node comparison score, the edge interference scores are sorted, the first K edges after sorting are identified as interference edges, and the interference edges are deleted to obtain a roughly clean graph; The rough clean graph and the benchmark anomaly graph are simultaneously input into the anomaly detection model, and the training process of the anomaly detection model, the node comparison score calculation process and the interference edge deletion process are repeatedly iterated until the number of deleted interference edges reaches a set threshold, thereby obtaining a target clean graph.
2. The graph data processing method according to claim 1, characterized in that: The method uses the trained anomaly detection model to obtain the difference between the similarity score of the first negative instance pair of the benchmark anomaly graph and the similarity score of the first positive instance pair, and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculates the node comparison score according to the difference, specifically: determining the anomaly score of the subgraph and node comparison scale according to the difference between the similarity score of the first negative instance pair of the benchmark anomaly graph and the similarity score of the first positive instance pair, and determining the anomaly score of the node and node comparison scale according to the difference between the similarity score of the second negative instance pair of the benchmark anomaly graph and the similarity score of the second positive instance pair; and calculating the node comparison score according to the anomaly score of the subgraph and node comparison scale and the anomaly score of the node and node comparison scale.
3. The graph data processing method according to claim 2, characterized in that: The node comparison score is calculated by using the anomaly score of the subgraph and node comparison scale and the anomaly score of the node and node comparison scale, specifically: weighted summing the anomaly score of the subgraph and node comparison scale and the anomaly score of the node and node comparison scale to obtain the node comparison score.
4. The graph data processing method according to claim 1, characterized in that: The node comparison score is obtained by the following formula: In the formula, represents the comparison score of the node in round r, represents the mean comparison score of the node, and R represents the total number of rounds.
5. The graph data processing method according to claim 1, characterized in that: The abnormal graph to be purified and the reference abnormal graph are simultaneously input into the abnormal detection model to train the abnormal detection model. The loss function used is: In the formula, represents the comparison loss between nodes and subgraphs in the baseline anomaly graph, represents the contrast loss between nodes and subgraphs in the rough clean graph, represents the node-to-node comparison loss in the baseline anomaly graph, represents the contrast loss between nodes in the rough and clean graph, α and β represent the balance parameters for different views and scales, respectively.
6. A graph data processing system, characterized in that: include: A model training module is used to input the anomaly graph to be purified and the benchmark anomaly graph into the anomaly detection model at the same time, and train the anomaly detection model, specifically: using the anomaly detection model to perform random walks on the anomaly graph to be purified and the benchmark anomaly graph at the same time, extracting subgraphs of nodes from the anomaly graph to be purified and the benchmark anomaly graph respectively; forming a first positive instance pair with the node and its own subgraph, and forming a first negative instance pair with the subgraphs of other nodes, forming a second positive instance pair with the node and the updated node representation in the subgraph, and forming a second negative instance pair with the updated node representation in other subgraphs; training the anomaly detection model according to the similarity scores of the first positive instance pair, the first negative instance pair, the second positive instance pair, and the second negative instance pair of the anomaly graph to be purified and the benchmark anomaly graph; A node comparison score acquisition module is used to use the trained anomaly detection model to obtain the difference between the similarity score of the first negative instance pair and the similarity score of the first positive instance pair of the benchmark anomaly graph and the difference between the similarity score of the second negative instance pair and the similarity score of the second positive instance pair, and calculate the node comparison score according to the difference; The interference edge deletion module is used to calculate the edge interference score of the abnormal graph to be cleaned through the node comparison score, sort the edge interference scores, identify the first K edges after sorting as interference edges, and delete the interference edges to obtain a roughly clean graph; The target clean graph acquisition module is used to input the rough clean graph and the benchmark anomaly graph into the anomaly detection model at the same time, and repeat the iterative operation of the anomaly detection model training process, the node comparison score calculation process and the interference edge deletion process until the number of deleted interference edges reaches a set threshold, thereby obtaining the target clean graph.
7. A computer device, characterized in that: It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the graph data processing method described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the graph data processing method described in any one of claims 1-5.