GNN Encoder and Outlier Detection Method Based on Graph Context Learning

By introducing edge updater and graph updater into the GNN encoder, combined with graph context comparison learning, the problem that abnormal point detection method in the prior art is difficult to effectively utilize graph context, and a stronger abnormal point detection capability is achieved.

CN113076738BActive Publication Date: 2025-05-30BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110385328.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-09
Publication Date
2025-05-30
Estimated Expiration
2041-04-09

AI Technical Summary

Technical Problem

The existing anomaly point detection methods focus on graph structure feature engineering or learning, and it is difficult to effectively use graph context comparison learning to measure the distance between anomalies and normal nodes and graph context.

Method used

A GNN encoder is proposed to remove suspicious links through an edge updater and update the graph representation through a graph updater. Combined with the graph context comparison learning method, the distance between the node and the graph context is measured to detect abnormal points.

Benefits of technology

The method of using graph context comparison to learn to detect abnormal points is realized. Compared with the traditional method, it is more powerful when distinguishing abnormal nodes from normal nodes, and can effectively perform abnormal point detection in unsupervised scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113076738B_ABST
    Figure CN113076738B_ABST
Patent Text Reader

Abstract

This application proposes an outlier detection method based on graph context learning, which relates to the field of computer network information technology. Among them, the method includes: proposing the CoGCL framework, and using graph context contrast learning to measure the distances between outliers and normal nodes and the graph context. To achieve the purpose of contrast learning, this patent designs a graph encoder, which can, to a certain extent, remove suspicious links while learning the representation of the graph context. To alleviate the impact of the scarcity of labeled data, this patent additionally extends CoGCL to a self-supervised pre-training framework CoGCL-pre that does not require labeled data. Through a graph perturbation strategy, this framework can automatically generate pseudo-labels for self-supervised learning. The CoGCL framework adopting the above solution is significantly superior to various existing contrast methods; its self-supervised version CoGCL-pre that does not require supervised data can achieve an effect equivalent to that of the fully supervised version CoGCL, and solves the impact of the scarcity of labeled data on supervised learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer network information technology, and particularly to a GNN encoder and an outlier detection method based on graph context learning. Background Art

[0002] Outlier detection has a profound impact on preventing malicious activities in real-world applications, such as the detection of malicious comments and error information detection. Since graphs can be used to naturally model the dependencies behind data, graph-based outlier detection methods have become the mainstream of development. Recently, with the development of Graph Neural Networks (GNNs), methods for efficiently detecting outliers using GNNs have emerged in an endless stream. The main idea is to use GNNs to learn the representations of nodes, and then distinguish normal or abnormal nodes based on a classifier. Summary of the Invention

[0003] This application aims to solve at least one of the technical problems in the related art to some extent.

[0004] To this end, the first object of this application is to propose a GNN encoder. Different from existing GNN models, the GNN encoder additionally adds an edge updater to remove suspicious links between nodes, and a graph updater to update the graph representation.

[0005] The second object of this application is to propose an outlier detection method based on graph context learning, which solves the problem that existing outlier detection methods focus on graph structure feature engineering or learning, and realizes using graph context contrast learning to measure the distance between abnormal and normal nodes and the graph context.

[0006] To achieve the above object, an embodiment of the first aspect of this application proposes a GNN encoder, including:

[0007] An edge updater, which estimates the suspicious probability of each edge at the beginning of each layer of GNN encoding, and removes suspicious links according to the suspicious probability of each edge to adjust the adjacency matrix;

[0008] A node representation updater, which aggregates the neighbor information of nodes according to the adjusted adjacency matrix, updates the vector representation of the current node, and obtains an updated node vector;

[0009] A graph representation updater, which updates the current graph representation according to the updated node vector and the graph representation of the previous layer.

[0010] Optionally, in the embodiment of this application, in the edge updater, the formula for updating each edge is:

[0011] A l = f edge (H (l-1),A (l-1) ,q (l-1) )

[0012] Among them, A l is the updated adjacency matrix, A (l-1) is the adjacency matrix of the previous layer, H (l-1) is the node vector representation matrix of the previous layer, q (l-1) is the graph representation;

[0013] In the node representation updater, the formula for updating the node representation is:

[0014] H l = f node (H (l-1) , A l )

[0015] Among them, H (l-1) is the node vector representation matrix of the previous time, H (l) is the node vector representation matrix of the updated current layer, A l is the adjacency matrix of the current layer, f node is the update function of the node representation;

[0016] In the graph representation updater, the formula for updating the graph representation is:

[0017] q (l) = f graph (H (l) , q (l-1) )

[0018] Among them, q (l-1) is the representation of the upper layer; q (l) is the graph representation of the updated current layer.

[0019] Optionally, in the embodiments of the present application, the global information of the graph is introduced into the edge updater, and the global information of the graph is the distance between the node and the graph; among them,

[0020] The distance between the node and the graph is used as a potential label to assist in estimating the probability of the suspicious link through the potential label.

[0021] Optionally, in the embodiments of the present application, the method for updating the edge updater includes the following steps:

[0022] First, a link prediction module is constructed based on the graph context, and the link prediction module is constructed by the following formula:

[0023]

[0024] Among them, is a vector connection operator, MLP is a fully connected layer, (h i (l-1) -q (l-1) ) is the global information of the graph, h i (l -1) is the vector representation of node i in the (l - 1)-th layer, is the vector representation of node i after combining the global information of the graph;

[0025] Secondly, calculate the credibility score p ij of each edge A ij ,

[0026]

[0027] where ReLU is a non-linear activation function that maps the score to [0, 1], is the edge between node i and node j, is the vector representation of node i after combining the global information of the graph;

[0028] Thirdly, make the process of discretely removing edges differentiable through the Gumbel-Softmax reparameterization technique;

[0029] where, for the credibility score of each edge, sample a noise ε ∈ Gumbel(0, 1) from the Gumbel distribution, add it to and take the logarithm,

[0030] Finally, call the Sigmoid activation function to map it to between [0, 1], and the formula is:

[0031]

[0032] where λ represents a hyperparameter, the square brackets represent floor function, ε is a noise sampled from the Gumbel distribution ε ∈ Gumbel(0, 1), λ is a pre-set hyperparameter, is the updated edge between node i and node j, taking values of 0 or 1, 0 represents deleting the edge, 1 represents retaining the edge, is the edge between node i and node j,

[0033] Optionally, in the embodiments of the present application, a cross-entropy loss function for link prediction is introduced into the edge updater, and the fitting of the edge updater is accelerated through the cross-entropy loss function;

[0034] where, the formula of the cross-entropy loss function is:

[0035]

[0036] Among them, is the loss function, is the edge of the upper layer (the (l - 1)-th layer), is the confidence score of the edge between node i and node j, where i and j represent node i and node j respectively.

[0037] Optionally, in the embodiments of the present application, in the update of the node representation updater, the formula for aggregating the neighbor information is:

[0038]

[0039] Among them, is the aggregated neighbor vector of node i, and AGGREGATION is the aggregation function of the neighbor vector.

[0040] The vector representation of the current node is updated through a feature function to obtain the updated node vector, and the representation formula of the node vector is:

[0041]

[0042] Among them, COMBINE is the aggregation function, is the vector representation of the updated node i.

[0043] Optionally, in the embodiments of the present application, a memory cache is introduced in the graph representation updater, and the graph representation of the previous layer is recorded through the memory cache, and the memory cache is used as a guide to calculate the importance of each node vector in this layer;

[0044] The recording of the graph representation of the previous layer through the memory cache and the use of the memory cache as a guide to calculate the importance of each node vector in this layer include the following steps:

[0045] Use the graph representation q (l-1) of the previous layer as the memory m, and calculate the importance of each node vector through the following formula:

[0046]

[0047]

[0048] Subsequently, the new graph representation is Among them, the new graph representation is added to the memory cache m = q (l) for the calculation of the next layer;

[0049] Among them, Denote the importance score of the \(i\)-th node in the \(l\)-th layer. Is the normalized representation of the importance score of node \(i\). Is the vector representation of node \(i\), where \(m\) represents the graph representation \(q\) of the previous layer. (l-1) Memory, \(q\) (l) Is the graph representation of the \(l\)-th layer; \(N\) is the number of nodes.

[0050] To achieve the above object, an embodiment of the first aspect of the present application proposes an outlier detection method based on graph context learning, including:

[0051] Obtain a graph network \(G=(V, X, A, Y)\) with node labels, where \(V\) is the set of nodes, \(X\) is the corresponding node feature matrix, \(A\) is the adjacency matrix, and \(A\in R\) N×N , and \(Y\) is the label of the node;

[0052] Learn the distance between the nodes and the graph context in the graph network through the CoGCL outlier detection framework. When the distance between the node and the graph context is greater than a preset value, the node is an outlier node; otherwise, it is a normal node. Among them,

[0053] The outlier detection framework includes the above-mentioned GNN encoder and contrast loss function. The node vector and graph vector of each node are obtained through the GNN encoder, and graph contrast learning is performed on the node vector and graph vector of each node through the contrast loss function.

[0054] Optionally, in the embodiment of the present application, the edge between the outlier node and the normal node is a suspicious link.

[0055] Optionally, in the embodiment of the present application, the formula for performing graph contrast learning on the node vector and graph vector of each node through the contrast loss function is:

[0056]

[0057] where \(h\) i is the vector representation of each node, \(\tau\) is a hyperparameter, \(q\) is the graph representation, \(X\) is the corresponding node feature matrix, \(A\) is the node adjacency matrix, is the loss function of graph contrast learning.

[0058] Optionally, in the embodiment of the present application, the CoGCL outlier detection framework further includes outlier prediction, calculates the cosine similarity score of the node vector and graph vector, and determines whether the node is an outlier node through the cosine similarity score.

[0059] Optionally, in the embodiments of the present application, by adding a graph perturbation strategy on the basis of the framework CoGCL, a self-supervised pre-training-free framework CoGCL-pre can be obtained, including:

[0060] The self-supervised pre-training-free framework CoGCL-pre adds a graph perturbation strategy on the basis of the CoGCL outlier detection framework, injects nodes outside the original graph into the original graph as perturbations, and uses these perturbations as pseudo-outlier nodes for the context of the current original graph, so as to obtain pseudo-label data for pre-training;

[0061] The CoGCL outlier detection framework includes the outlier detection method described above, and regards the pseudo-outlier nodes as the outlier nodes through the CoGCL outlier detection framework;

[0062] The outlier detection framework includes the GNN encoder and the contrast loss function. The node vectors and graph vectors of each node are obtained through the GNN encoder, and graph contrast learning is performed on the node vectors and graph vectors of each node through the contrast loss function.

[0063] Optionally, in the embodiments of the present application, pseudo-outlier nodes are injected into the CoGCL outlier detection framework, and the context of the original graph is destroyed through the pseudo-outlier nodes to construct pseudo-labels.

[0064] Optionally, in the embodiments of the present application, in the graph perturbation strategy, there are many methods for splitting graphs, including:

[0065] When there are multiple graphs for outlier detection, there is naturally a segmentation of multiple graphs at this time, and the multiple graphs can perturb each other;

[0066] When performing segmentation on a large graph, a graph clustering method is called to cluster a large graph into multiple subgraphs, and at this time, the subgraphs can perturb each other.

[0067] Optionally, in the embodiments of the present application, injecting nodes outside the original graph into the original graph as perturbations includes:

[0068] Given a graph G=(V, X, A), we use a certain strategy to divide it into several subgraphs

[0069] For each subgraph G i , we inject a node set belonging to other subgraphs into it Thus, a new perturbed graph is obtained

[0070] where the nodes adjacency matrix is a slice of the overall adjacency matrix A.

[0071] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:

[0073] Figure 1 is a visual representation of mapping the initial vector of each paper to a two-dimensional vector space using the t-SNE method in an embodiment of the present application;

[0074] Figure 2 are the original input features in an embodiment of the present application;

[0075] Figure 3 is the distance distribution between the node vectors and the graph context after being processed by the GCNs model in an embodiment of the present application;

[0076] Figure 4 is the distance distribution between the node vectors and the graph context after being processed by the CoGCL model in an embodiment of the present application;

[0077] Figure 5 is the CoGCL model framework in an embodiment of the present application;

[0078] Figure 6 are the pre-training experimental results in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0079] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0080] The multi-parameter synchronous detection method and device for a multi-membrane electrode of a fuel cell stack according to an embodiment of the present application will be described below with reference to the accompanying drawings.

[0081] The first object of the present application is to propose a GNN encoder. Different from existing GNN models, the GNN encoder additionally adds an edge updater to remove suspicious links between nodes and a graph updater to update the graph representation.

[0082] The second objective of this application is to propose an outlier detection method based on graph context learning, which solves the problem that existing outlier detection methods focus on graph structure feature engineering or learning, and realizes the use of graph context contrast learning to measure the distances between outliers and normal nodes and the graph context.

[0083] To achieve the above objective, an embodiment of the first aspect of this application proposes a GNN encoder, including:

[0084] An edge updater, at the beginning of each layer of GNN encoding, estimates the suspicious probability of each edge, and removes suspicious links according to the suspicious probability of each edge to adjust the adjacency matrix;

[0085] A node representation updater, aggregates the neighbor information of nodes according to the adjusted adjacency matrix, and updates the vector representation of the current node to obtain an updated node vector;

[0086] A graph representation updater, updates the current graph representation according to the updated node vector and the graph representation of the previous layer. Optionally, in the embodiment of this application, in the edge updater, the formula for updating each edge is:

[0087] A l = f edge (H (l-1) , A (l-1) , q (l-1) )

[0088] where, A l is the updated adjacency matrix, A (l-1) is the adjacency matrix of the previous layer, H (l-1) is the node vector representation matrix of the previous layer, q (l-1) is the graph representation;

[0089] In the node representation updater, the formula for updating the node representation is:

[0090] H l = f node (H (l-1) , A l )

[0091] where, H (l-1) is the node vector representation matrix of the previous time, H (l) is the node vector representation matrix of the updated current layer, A l is the adjacency matrix of the current layer, f node is the update function of the node representation;

[0092] In the graph representation updater, the formula for updating the graph representation is:

[0093] q(l) = f graph (H (l) , q (l-1) )

[0094] where q (l-1) is the representation of the upper layer; q (l) is the representation of the updated graph of this layer.

[0095] Optionally, in the embodiments of the present application, global information of the graph is introduced into the edge updater, and the global information of the graph is the distance between a node and the graph; where

[0096] the distance between the node and the graph is used as a potential label, and the probability estimation of the suspicious link is assisted by the potential label.

[0097] Optionally, in the embodiments of the present application, the method for updating the edge updater includes the following steps:

[0098] First, a link prediction module is constructed based on the graph context, and the link prediction module is constructed by the following formula:

[0099]

[0100] where is a vector connection operator, MLP is a fully connected layer, (h i (l-1) - q (l-1) ) is the global information of the graph, h i (l -1) is the vector representation of node i in the (l - 1)th layer, is the vector representation of node i after combining the global information of the graph;

[0101] Secondly, calculate the credibility score p ij of each edge A ij ,

[0102]

[0103] where ReLU is a non - linear activation function that maps the score to [0, 1], is the edge between node i and node j 's credibility score, is the vector representation of node i after combining the global information of the graph;

[0104] Thirdly, the Gumbel - Softmax reparameterization technique is used to make the process of discretely removing edges differentiable;

[0105] where, for the credibility score of each edge Sample a noise ε ∈ Gumbel(0, 1) from the Gumbel distribution, and add it to and take the logarithm,

[0106] Finally, call the Sigmoid activation function to map it to the range [0, 1]. The formula is:

[0107]

[0108] where λ represents the hyperparameter, the square brackets represent floor function, ε is a noise sampled from the Gumbel distribution ε ∈ Gumbel(0, 1), and λ is a pre-set hyperparameter. is the updated edge between node i and node j, which takes values of 0 or 1. 0 means deleting the edge, and 1 means retaining the edge. is the edge between node i and node j 's credibility score.

[0109] Optionally, in the embodiments of the present application, a cross-entropy loss function for link prediction is introduced into the edge updater, and the fitting of the edge updater is accelerated through the cross-entropy loss function;

[0110] where the formula of the cross-entropy loss function is:

[0111]

[0112] where, is the loss function, is the edge of the upper layer (the (l - 1)-th layer), is the edge between node i and node j 's credibility score, and i, j represent node i and node j respectively.

[0113] Optionally, in the embodiments of the present application, in the update of the node representation updater, the formula for aggregating the neighbor information is:

[0114]

[0115] where, is the aggregated neighbor vector of node i, and AGGREGATION is the aggregation function of the neighbor vector;

[0116] Update the vector representation of the current node through a feature function to obtain the updated node vector. The representation formula of the node vector is:

[0117]

[0118] where COMBINE is the aggregation function, It is the vector representation of the updated node i.

[0119] Optionally, in the embodiment of the present application, a memory cache is introduced in the graph representation updater, and the graph representation of the previous layer is recorded through the memory cache, and the memory cache is used as a guide to calculate the importance of each node vector in this layer;

[0120] Recording the graph representation of the previous layer through the memory cache and using the memory cache as a guide to calculate the importance of each node vector in this layer includes the following steps:

[0121] Using the graph representation q of the previous layer (l-1) As the memory m, the importance of each node vector is calculated through the following formula:

[0122]

[0123]

[0124] Subsequently, the new graph representation is Among them, the new graph representation is added to the memory cache m = q (l) For the calculation of the next layer;

[0125] Among them, Represents the importance score of the i-th node in the l-th layer, Is the normalized representation of the importance score of node i, Is the vector representation of node i, m represents the graph representation q of the previous layer (l-1) Memory, q (l) Is the graph representation of the l-th layer; N is the number of nodes.

[0126] To achieve the above object, the second aspect embodiment of the present application proposes an outlier detection method based on graph context learning, including:

[0127] Obtain a graph network G=(V, X, A, Y) with node labels, where V is the set of nodes, X is the corresponding node feature matrix, A is the adjacency matrix, and A∈R N×N , Y is the label of the node;

[0128] Learn the distance between the nodes and the graph context in the graph network through the CoGCL outlier detection framework. When the distance between the node and the graph context is greater than the preset value, the node is an outlier node, otherwise it is a normal node; among them,

[0129] The outlier detection framework includes the GNN encoder and the contrast loss function. The node vector and graph vector of each node are obtained through the GNN encoder, and graph contrast learning is performed on the node vector and graph vector of each node through the contrast loss function.

[0130] Optionally, in the embodiments of the present application, the edge between the abnormal node and the normal node is a suspicious link.

[0131] Optionally, in the embodiments of the present application, by using the contrast loss function for the node vectors

[0132] and the graph vectors for graph contrast learning, the formula is:

[0133]

[0134] where h i is the vector representation of each node, τ is a hyperparameter, q is the graph representation, X is the corresponding node feature matrix, A is the node adjacency matrix, is the loss function for graph contrast learning.

[0135] Optionally, in the embodiments of the present application, the CoGCL anomaly detection framework further includes anomaly prediction, calculating the cosine similarity score of the node vectors and the graph vectors, and determining whether a node is an abnormal node through the cosine similarity score.

[0136] Optionally, in the embodiments of the present application, based on the CoGCL framework, adding a graph perturbation strategy can obtain a self-supervised pre-training-free framework CoGCL-pre, including:

[0137] The self-supervised pre-training-free framework CoGCL-pre injects nodes outside the original graph into the original graph as perturbations on the basis of the CoGCL anomaly detection framework, and uses these perturbations as pseudo-abnormal nodes for the context of the current original graph, so as to obtain pseudo-label data for pre-training;

[0138] The CoGCL anomaly detection framework includes the above-mentioned anomaly detection method, and regards the pseudo-abnormal nodes as the abnormal nodes through the CoGCL anomaly detection framework;

[0139] The anomaly detection framework includes the GNN encoder and the contrast loss function. The node vectors and graph vectors of each node are obtained through the GNN encoder, and graph contrast learning is performed on the node vectors and graph vectors of each node through the contrast loss function.

[0140] Optionally, in the embodiments of the present application, pseudo-abnormal nodes are injected into the CoGCL anomaly detection framework, and the context of the original graph is destroyed by the pseudo-abnormal nodes to construct pseudo-labels.

[0141] Optionally, in the embodiments of the present application, in the graph perturbation strategy, there are many ways to divide the graph, including:

[0142] When there are multiple graphs for which outlier detection is required, there is naturally a division of multiple graphs at this time, and the multiple graphs can perturb each other;

[0143] When performing segmentation on a large graph, the graph clustering method is called to cluster a large graph into multiple subgraphs, and at this time, the subgraphs can perturb each other.

[0144] Optionally, in the embodiments of the present application, injecting nodes outside the original graph into the original graph as perturbations includes:

[0145] Given a graph G=(V,X,A), we use a certain strategy to divide it into several subgraphs

[0146] For each subgraph G i , we inject a node set belonging to other subgraphs into it Thus, a new perturbed graph is obtained

[0147] where the nodes adjacency matrix is a slice of the overall adjacency matrix A.

[0148] To enable those skilled in the art to better understand the present application, an outlier detection method based on graph context learning is taken as an example;

[0149] The method of the embodiments of the present application is described below with reference to the accompanying drawings:

[0150] In the embodiments of the present application, in order to further understand the behavior of outliers, we use the Google Scholar paper collection of the scholar "Jun Lu" from Yale University for visual analysis. If two papers have the same co-authors or are published in the same conference, then there is an edge between the pair of papers. In the following figures, we explore how to distinguish abnormal (misassigned) papers from normal papers.

[0151] Figure 1 Shows a visual representation of mapping the initial vector of each paper to a two-dimensional vector space using the t-SNE method;

[0152] Among them, small dots represent normal papers and large dots represent abnormal papers.

[0153] Furthermore, the initial vector of the paper is obtained by inputting the title and keywords of the paper into the BERT model. As can be seen from the figure, whether it is a normal paper or an abnormal paper, their absolute feature distributions are very different, and the difference of abnormal papers is even greater. This makes the previous optimization methods based on node classification unable to well solve the problem of feature distribution difference. Subsequently, we quantified this observation by calculating the distance (cosine similarity) between each node and the graph context (the average value of all node features). Figure 2 It is clearly shown that although the distance distributions of normal papers and abnormal papers from the graph context are very similar (y ∈ [0.97, 1]), normal papers and abnormal papers can still be clearly distinguished. Based on this observation, we can obtain a more general hypothesis, that is, outliers are more different from the distribution of the whole graph compared with normal points. Based on this hypothesis, we propose the CoGCL model, using graph contrast learning as the optimization objective to compare the distances between abnormal and normal nodes and the graph context. Figure 3 and Figure 4 respectively represent the distance distributions between the node vectors processed by GCNs and the CoGCL model and the graph context, demonstrating the powerful ability of CoGCL compared with traditional GCNs-based classification models in distinguishing abnormal nodes from normal nodes.

[0154] Furthermore, previous methods are all based on supervised learning, so they are all affected by the scarcity of labeled data. Especially in some fields, it is often impossible to obtain sufficient and effective labeled data. Therefore, we propose the CoGCL-pre framework based on CoGCL. This framework uses a graph context perturbation strategy, that is, injecting foreign nodes (outliers) into the original graph to disrupt the context of the original graph, so as to construct pseudo-labels and self-supervised training tasks, thereby getting rid of the dependence on supervised data to a certain extent.

[0155] Specifically, in the embodiments of the present application, an outlier detection framework CoGCL based on graph context contrast learning is proposed, which can better solve the problem of data distribution difference compared with traditional GNNs-based classification methods; at the same time, a self-supervised learning method CoGCL-pre is proposed, which solves the problem of the model's dependence on labeled data to a certain extent.

[0156] Furthermore, the outlier detection problem can be defined as inputting a graph network G=(V, X, A, Y) with node labels, where V represents a set of N nodes, and A ∈ R N×N represents the adjacency matrix. Without loss of generality, we define G as an undirected, unweighted, and single-node relationship graph, that is, if there is an edge between nodes v i and v j then A ij = 1, otherwise Aij = 0. X is the corresponding node feature matrix, where x i ∈R d represents the node v i has a d-dimensional feature vector. Y represents the labels of each node, where y i = 1 represents that the node v i is an abnormal node, otherwise it is a normal node. The purpose of anomaly detection is to learn the function g: R d → {0, 1} to predict whether a given node is a normal node (0) or an abnormal node (1).

[0157] Furthermore, in order to accurately detect anomalies, this patent proposes CoGCL, an anomaly detection framework based on graph context contrast. CoGCL is based on an observation of the actual situation, that is, there are significant differences in the distances between normal nodes and abnormal nodes and the graph context. Therefore, the farther a node is from most nodes in the feature space, the more likely it is to be an abnormal node. Therefore, the optimization strategy of graph context contrast learning is defined as follows: Given a graph G, we first use a GNN encoder of features to obtain the vector representation h i of each node v i and the graph representation q, that is, (H, q) = f GNN (X, A), where Subsequently, we regard q as a query, regard the vectors of all normal nodes as positive example values, regard abnormal nodes as negative examples, and use the infoNCE loss function to implement graph contrast learning, which is specifically defined as follows:

[0158]

[0159] Among them, this loss function narrows the distance between normal nodes and the graph context representation while pushing the distance between abnormal nodes and the context farther, so as to achieve the effect of identifying abnormal points.

[0160] Among them, h i is the vector representation of each node, τ is a hyperparameter, q is the graph representation, X is the corresponding node feature matrix, A is the node adjacency matrix, is the loss function of graph contrast learning.

[0161] In the embodiments of this application, in order to achieve we define the GNN encoder f GNN . Different from ordinary GNN models, we additionally design a node-edge updater to remove suspicious links, that is, the edges between abnormal nodes and normal nodes, and a graph representation updater to update the graph context representation in real time, which are specifically defined as follows:

[0162] Edge Updater: Estimate the suspicious probability of each edge at the beginning of each layer of GNN encoding, and adjust the adjacency matrix based on this to remove suspicious links as much as possible, that is

[0163] A l =f edge (H (l-1) ,A (l-1) ,q (l-1) )

[0164] where A l is the updated adjacency matrix, A (l-1) is the adjacency matrix of the previous layer, H (l-1) is the node vector representation matrix of the previous layer, and q (l-1) is the graph representation;

[0165] Node Representation Updater: Aggregate the neighbor information of nodes to update the vector representation of the current node according to the adjusted adjacency matrix. This module can be replaced by any GNN encoder, that is

[0166] H l =f node (H (l-1) ,A l )

[0167] where H (l-1) is the node vector representation matrix of the previous time, H (l) is the node vector representation matrix of the updated current layer, A l is the adjacency matrix of the current layer, and f node is the update function of the node representation;

[0168] Graph Representation Updater: Update the current graph representation according to the updated node vector and the graph representation of the previous layer, that is

[0169] q (l) =f graph (H (l) ,q (l-1) )

[0170] where q (l-1) is the representation of the upper layer; q (l) is the graph representation of the updated current layer.

[0171] Furthermore, the schematic diagram of the overall CoGCL framework is as shown in Figure 2 . At each layer, the edge updater, node representation updater, and graph representation updater will act on the graph sequentially. After l layers of convolution, the loss function will act on the finally obtained graph representation and node representation to calculate the loss, and then call the gradient descent algorithm to optimize the overall module.

[0172] Anomaly Point Prediction: After obtaining the node vector representation and the graph vector representation, it is different from the traditional method of directly predicting the labels of nodes through node vectors. CoGCL calculates the cosine similarity score between the node vector and the graph vector. The lower the score, the more likely the node is an abnormal node. This scoring mechanism can be more flexibly applied in different scenarios.

[0173] In the embodiments of this application, in order to implement CoGCL based on self-supervised learning, specifically, the effect of the anomaly detection model based on supervised learning is greatly affected by the quality of label data. To obtain an ideal detection effect, a large amount of high-quality supervised data is required, and the label data of anomaly points is often difficult to obtain in some fields. In recent years, graph-based self-supervised learning methods have achieved good results in graph pre-training. For example, GAE preserves the structure information of the graph by reconstructing the adjacency matrix; GPT-GNN additionally preserves the attribute relationships of nodes by predicting the attributes of nodes; DGI preserves graph information by maximizing the mutual information between the graph context and the node vectors; GraphCL preserves graph information by maximizing two augmented views of a graph. Therefore, we propose a self-supervised pre-training strategy without labels to solve the anomaly detection problem in graphs. Inspired by graph contrastive learning, we propose to construct pseudo-labels by disrupting the original graph. Specifically, we inject nodes outside the original graph into the original graph as perturbations, so these perturbations can act as pseudo-anomaly nodes for the current original graph context, that is,

[0174] Original Graph Perturbation: Given a graph G=(V, X, A), we use a certain strategy to divide it into several subgraphs For each subgraph G i , we inject a node set belonging to other subgraphs into it Thus, a new perturbed graph is obtained where the nodes of adjacency matrix is a slice of the overall adjacency matrix A. The schematic diagram of graph perturbation is as Figure 5 shown.

[0175] In the embodiments of this application, the implementation details of CoGCL will be introduced in detail below. We first introduce the implementation method of the GNN encoder with three updaters,

[0176] Edge Updater: The suspicious links are the edges between normal nodes and abnormal nodes. These edges violate the homogeneity assumption of traditional GNN encoders, that is, adjacent neighbor nodes have similar attribute representations. Therefore, these suspicious links directly affect the performance of GNN encoders. Most previous methods directly estimate the probability of a link being suspicious based on the vector representations of the nodes connected by the link itself. However, we additionally introduce the global information of the graph, that is, the distance between the nodes and the graph. This distance information can be regarded as a potential label to assist in estimating the probability of suspicious links. Specifically, if the distance gap between two nodes connected by an edge and the graph context is larger, the probability that this edge is a suspicious link is relatively high. Specifically, we first define a link prediction module based on the graph context,

[0177]

[0178] wherein, is the vector concatenation operator, and MLP is the fully connected layer. (h i (l-1) -q (l-1) ) is the added global information of the graph, h i (l-1) is the vector representation of node i at the (l - 1)-th layer, is the vector representation of node i after incorporating the global information of the graph.

[0179] Subsequently, we calculate the credibility score p ij of edge A ij ,

[0180]

[0181] wherein, ReLU is the non-linear activation function that maps the score to [0, 1], is the credibility score of the edge between node i and node j, is the vector representation of node i after incorporating the global information of the graph;

[0182] Furthermore, we use the Gumbel-Softmax reparameterization technique to make the discrete edge deletion process differentiable.

[0183] Specifically, for the credibility score of each edge, we sample a noise ε ∈ Gumbel(0, 1) from the Gumbel distribution, add it to and take the logarithm. Finally, we apply the Sigmoid activation function to map it to between [0, 1]. The formula is as follows:

[0184]

[0185] where λ represents a hyperparameter, the square brackets represent floor function, ε is a noise sampled from the Gumbel distribution ε ∈ Gumbel(0, 1), and λ is a pre-set hyperparameter. is the updated edge between node i and node j, taking values of 0 or 1, where 0 means deleting the edge and 1 means keeping the edge. is the edge between node i and node j and its credibility score.

[0186] We additionally impose a cross-entropy loss function for link prediction to accelerate the fitting of the edge updater.

[0187] where is the loss function, is the edge of the upper layer (layer l - 1), is the edge between node i and node j and its credibility score, where i and j represent node i and node j respectively.

[0188] In the embodiments of the present application, the node represents an updater: updating the vector representation of the current node can be specifically divided into the following two steps. First, we aggregate the information of neighbors according to the modified adjacency matrix, that is

[0189]

[0190] where is the aggregated neighbor vector of node i, and AGGREGATION is the aggregation function of neighbor vectors;

[0191] Furthermore, we combine the aggregated neighbor information with the vector representation of the node itself through a combination function to obtain a new vector representation.

[0192]

[0193] In the implementation process, we use the aggregation and combination functions of GIN [Xu, 2018].

[0194] where COMBINE is the aggregation function, is the vector representation of the updated node i.

[0195] In the embodiments of the present application, the graph represents an updater: after obtaining the updated node vectors, we can call traditional pooling methods such as summation, averaging, and maximizing to update the graph representation. However, traditional pooling methods do not distinguish between normal nodes and abnormal nodes. To solve this problem, we introduce a memory cache to record the graph representation of the previous layer and use it as a guide to calculate the importance of each node vector in this layer. Specifically, we first use the graph representation q of the previous layer (l-1)Calculate the importance of each node vector as memory m.

[0196]

[0197] Furthermore, the new graph is represented as At the same time, add it to the memory cache m = q (l) For the calculation of the next layer.

[0198] Among them, represents the importance score of the i-th node in the l-th layer, is the normalized representation of the importance score of node i, is the vector representation of node i, m represents the graph representation q of the previous layer (l-1) memory, q (l) is the graph representation of the l-th layer; N is the number of nodes.

[0199] In the embodiment of the present application, in the graph perturbation strategy, there are many ways to divide the graph. In this patent, if there are already multiple graphs that need to be detected for outliers, there is naturally a division of multiple graphs at this time. For example, in an academic knowledge system, the papers of each author can be regarded as a graph. Therefore, detecting abnormal papers requires detection on different scholar graphs. Thus, for each scholar, the papers of another scholar are perturbations; if a large graph is divided, we call the graph clustering method to cluster a large graph into multiple subgraphs, and at this time, the subgraphs can perturb each other.

[0200] In the embodiment of the present application, the present application has conducted sufficient experiments on the dataset AMiner of the academic knowledge graph.

[0201] Dataset AMiner1: It is a free online academic search and mining system, which has collected more than 100 million expert scholars and 260 million paper collections. We extracted the papers owned by 1,104 experts from AMiner, and regarded each paper as a node in the graph. If there are the same co-authors, work institutions or published in the same conference between any two papers, an edge is added to these two papers. The true label of the wrong papers in each expert profile is manually annotated.

[0202] Evaluation metrics: We use two metrics, Area Under ROC Curve (AUC) and Mean Average Precision (MAP), to comprehensively evaluate the effect of outlier detection: AUC is a comprehensive classification metric widely used in the field of outlier detection; MAP is a ranking metric, and when used in the field of outlier detection, it more emphasizes the relative ranking of outliers.

[0203] Furthermore, the evaluation of CoGCL in the supervised scenario:

[0204] Training settings: Among the papers owned by 1,104 experts in AMiner, we selected approximately 70% of the experts as the training set and the remaining 30% as the test set. The initial vector of the paper is obtained by inputting the title and keywords of the paper into BERT.

[0205] Comparison methods: We compared two classic graph neural network models, GCN and GIN; in addition, we also compared two state-of-the-art GNN-based outlier detection models, GraphConsis and CARE-GNN.

[0206] Experimental results: The experimental results are shown in Table 1. On the AMiner dataset, the outlier detection effect of CoGCL is far better than the state-of-the-art comparison methods: it is 11.70 - 20.45% higher in the AUC metric and 19.58 - 28.19% higher in the MAP metric. The experimental results fully demonstrate the superiority of the optimization framework based on graph context contrast learning.

[0207] The following table shows the experimental results of outlier detection:

[0208]

[0209]

[0210] Furthermore, the evaluation of CoGCL-pre in the unsupervised scenario:

[0211] Training settings: In the AMiner system, we additionally extracted 4,800 expert papers and used the graph perturbation strategy to perturb the original expert paper graph to obtain pseudo-labeled data for pre-training. At the same time, the same test set as CoGCL was used to evaluate the effect of CoGCL-pre.

[0212] Comparison methods: We compared four state-of-the-art graph self-supervised pre-training framework models, GAE, GPT-GNN, DGI, and GraphCL.

[0213] Experimental results: The experimental results are as Figure 6 shown. The following three observations can be obtained from the figure: 1. CoGCL-pre can achieve the same effect as CoGCL based on supervised learning without using supervised data; 2. When CoGCL-pre is fine-tuned using all supervised data, its effect exceeds CoGCL by approximately 1.96% in the MAP metric; 3. The effect of CoGCL-pre with any percentage of supervised data is significantly better than the other comparison methods. The above three experimental results fully demonstrate the effectiveness of the self-supervised pre-training model framework we proposed.

[0214] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0215] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0216] Any process or method description in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0217] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be embodied specifically in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0218] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0219] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0220] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0221] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. An outlier detection method based on graph context learning, characterized in that, applied to paper classification, the outlier detection method includes the following steps: Obtain a graph network G=(V, X, A, Y) with node labels, where V is the set of nodes, X is the corresponding node feature matrix, A is the adjacency matrix, and A∈R N×N , and Y is the label of the nodes; Learning the distance between nodes and graph context in the graph network through the CoGCL outlier detection framework. When the distance between a node and the graph context is greater than a preset value, the node is an outlier node; otherwise, it is a normal node. Among them, The CoGCL outlier detection framework includes a GNN encoder and a contrast loss function. The node vectors and graph vectors of each node are obtained through the GNN encoder, and graph contrast learning is performed on the node vectors and graph vectors of each node through the contrast loss function; The GNN encoder includes: an edge updater, a node representation updater, and a graph representation updater. At the beginning of each layer of GNN encoding, the edge updater estimates the suspicious probability of each edge and removes suspicious links according to the suspicious probability of each edge to adjust the adjacency matrix. The node representation updater aggregates the neighbor information of the node according to the adjusted adjacency matrix and updates the vector representation of the current node to obtain an updated node vector. The graph representation updater updates the current graph representation according to the updated node vector and the graph representation of the previous layer. Among them, the t-SNE method is used to map the initial vector of each paper to a visual representation in a two-dimensional vector space. If two papers have the same collaborators or are published in the same conference, then there is an edge between this pair of papers. The initial vector of the paper is obtained by inputting the title and keywords of the paper into the BERT model.

2. The outlier detection method according to claim 1, characterized in that, introducing the global information of the graph into the edge updater, and the global information of the graph is the distance between the node and the graph. Among them, taking the distance between the node and the graph as a potential label, and using the potential label to assist in estimating the probability of the suspicious link; introducing a memory cache into the graph representation updater, recording the graph representation of the previous layer through the memory cache, and using the memory cache as a guide to calculate the importance of each node vector in this layer; The step of recording the graph representation of the previous layer through the memory cache and using the memory cache as a guide to calculate the importance of each node vector in this layer includes the following steps: Use the graph of the upper layer to represent q (l-1) As memory m, calculate the importance of each node vector through the following formula: Subsequently, the new graph is represented as wherein, the new graph representation is added to the memory cache m = q (l) for the calculation of the next layer; Among them, represents the importance score of the i-th node in the l-th layer, is the normalized representation of the importance score of node i, is the vector representation of node i, m represents the graph representation q of the previous layer (l-1) memory, q (l) is the graph representation of the l-th layer; N is the number of nodes.

3. The outlier detection method according to claim 2, characterized in that, The method for updating the edge updater includes the following steps: First, construct a link prediction module based on the graph context, and the link prediction module is constructed by the following formula: Among them, is a vector connection operator, and MLP is a fully connected layer. is the global information of the graph. is the vector representation of node i in the (l - 1)-th layer. is the vector representation of node i after combining the global information of the graph. Secondly, calculate the credibility score p of each of the said edges ij ,​ where ReLU is a non - linear activation function that maps the score to [0, 1], is the edge between node i and node j of the credibility score, is the vector representation of node i after incorporating the global information of the graph; Secondly, use the Gumbel-Softmax reparameterization technique to make the process of discretely removing edges differentiable; Among them, for the confidence score of each edge Sample a noise ε ∈ Gumbel(0, 1) from the Gumbel distribution and add it to Add and take the logarithm Finally, call the Sigmoid activation function to map it to between [0,1], and the formula is: Among them, λ represents a hyperparameter, the square brackets represent floor function, ε is a noise sampled from the Gumbel distribution ε ∈ Gumbel(0, 1), and λ is a pre-set hyperparameter. is the edge between the updated node i and node j, taking values of 0 or 1, where 0 means deleting the edge and 1 means retaining the edge. is the edge between node i and node j 's credibility score.

4. The outlier detection method according to claim 3, characterized in that, introducing a cross-entropy loss function for link prediction into the edge updater, and accelerating the fitting of the edge updater through the cross-entropy loss function; Among them, the formula of the cross-entropy loss function is: Among them, is the loss function, is the edge of the upper (l - 1)-th layer, is the edge between node i and node j and the credibility score of it, where i and j respectively represent node i and node j.

5. The outlier detection method according to claim 1, characterized in that, in the update of the node representation updater, the formula for aggregating the neighbor information is: Among them, is the aggregated neighbor vector of node i, and AGGREGATION is the aggregation function of the neighbor vector; The vector representation of the current node is updated through a feature function to obtain an updated node vector, and the representation formula of the node vector is: where COMBINE is an aggregation function, is the vector representation of the updated node i.

6. The outlier detection method according to claim 1, characterized in that, the CoGCL outlier detection framework, further includes outlier prediction, calculating the cosine similarity score between the node vector and the graph vector, and determining whether the node is an outlier node through the cosine similarity score; wherein, the edge between the outlier node and the normal node is a suspicious link.

7. The outlier detection method according to claim 1, characterized in that, the formula for performing graph contrast learning on the node vector and the graph vector of each node through the contrast loss function is: where h i is the vector representation of each node, τ is a hyperparameter, q is the graph representation, X is the corresponding node feature matrix, A is the node adjacency matrix, is the loss function for graph contrastive learning.

8. The outlier detection method according to claim 1, characterized in that, includes: adding a graph perturbation strategy to the CoGCL outlier detection framework, the graph perturbation strategy is to inject foreign nodes into the original graph, the foreign nodes act as pseudo-outlier nodes for the context of the original graph, and the context of the original graph is destroyed through the pseudo-outlier nodes to construct pseudo-labels.

9. The outlier detection method according to claim 8, characterized in that, The method of segmenting the graph in the graph perturbation strategy includes: dividing the original graph into I subgraphs by calling graph segmentation methods such as clustering and I is a positive integer; For each of the subgraphs G i , inject the node set of subgraph G i into the subgraph G j to obtain the perturbed graph ​ Among them nodes adjacency matrix is a slice of the overall adjacency matrix A, and v j represents the nodes of subgraph j.

Citation Information

Patent Citations

  • Method and system for enhancing anti-attack capability of graph model

    CN111309975A

  • Social network abnormal account detection method and system

    CN111767472A