Dynamic graph anomaly detection method based on graph embedding and diffusion sampling

By using graph embedding and diffusion sampling methods in dynamic graph abnormality detection, combined with GCN, GRU and self-attention mechanisms, the shortcomings in dynamic graph abnormality detection in the prior art are solved, and the accuracy and performance of detection are improved.

CN119939450AActive Publication Date: 2025-05-06YUNNAN UNIV

Patent Information

Application Number
CN202411858808.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The existing dynamic graph anomaly detection methods have shortcomings in capturing the dynamic nature of graph evolution, the accuracy of obtaining the information on the nodes, and the integration of structural information and time information, resulting in a degradation of detection performance.

Method used

A dynamic graph anomaly detection method based on graph embedding and diffusion sampling is adopted to obtain the most influential node subgraphs through the diffusion matrix, feature extraction and temporal information fusion are combined with GCN and GRU models, and a self-attention mechanism is used to reduce the impact of noise.

Benefits of technology

It improves the accuracy of abnormal edge detection, reduces the impact of useless neighbor noise on node embedding, more effectively integrates structural and time information, and improves the performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939450A_ABST
    Figure CN119939450A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic graph anomaly detection method based on graph embedding and diffusion sampling, and aims to improve the accuracy of abnormal edge detection. The method comprises the following steps: acquiring a dynamic graph data set, segmenting the dynamic graph data set into an adjacent matrix set, and randomly initializing node features; generating an adjacent matrix in the time window, calculating a diffusion matrix, and selecting a node with the maximum influence to form a subgraph; inputting the sub-graph into a node spatio-temporal information fusion module, and extracting features through a two-layer GCN model and a GRU unit; carrying out negative sampling on abnormal edges, and sending node features into an anomaly detector to obtain anomaly scores; and calculating a cross entropy loss function based on the abnormal score and training the model. According to the method, neighbor noise influence is reduced through diffusion sampling, structure and time information are fused, and dynamic graph anomaly detection performance is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining and complex network research, and in particular relates to a dynamic graph anomaly detection method based on graph embedding and diffusion sampling. Background Art

[0002] In recent years, the study of graphs has received great attention and development in the fields of social networks, knowledge graphs, e-commerce, and transportation networks. In the real world, the properties and connections of graphs are constantly changing over time. Taking social networks as an example, users are constantly evolving through online activities on social platforms, resulting in constant changes in their properties and connections with other users. Therefore, unlike static graphs, dynamic graphs need to be mined over time to capture these changes. In the field of dynamic graph technology, the detection of abnormal edges is a key task. For example, in a social network, spam or other harassing information may appear, and even false information may be spread. Detecting such abnormal messages is crucial to maintaining a healthy environment for online communities.

[0003] Existing dynamic graph anomaly detection methods can be roughly classified into two categories: methods based on autoencoder structures and end-to-end methods. Methods based on autoencoder structures usually obtain the initial encoding through the initial feature encoding method, and then use the encoder and decoder to reconstruct the initial encoding and calculate the error. The appropriate encoder parameters are obtained through training to obtain the final potential representation vector. The end-to-end method is a model that directly inputs from the original dynamic graph to the anomaly detection output, usually based on supervised or semi-supervised learning. It does not need to be divided into intermediate steps such as feature extraction and reconstruction, but directly learns the mapping from input to output.

[0004] End-to-end methods can better capture the changes of dynamic graphs over time, which is very beneficial for tasks that require dynamic tracking and anomaly identification. However, its performance usually depends on high-quality labeled data. Without sufficient anomaly labels, the detection performance of the model may decline. Most end-to-end methods first obtain the embedding representation of the graph nodes, and then obtain the anomaly score based on the embedding. However, existing methods do not take into account the importance of nodes, resulting in unnecessary node noise.

[0005] In addition, the prior art has the following three problems that have not been well solved:

[0006] (1) Capturing the dynamics of graph evolution. Current research focuses on static graphs, and there are relatively few methods for processing dynamic graphs. Existing methods only use recurrent neural network models or attention mechanisms to capture the dynamics of graph evolution, but lack the idea of ​​obtaining long-term dynamic evolution.

[0007] (2) The obtained node context information is not accurate. Existing methods extract the h-hop neighbors of the target node or use them as context. However, this strategy can lead to performance degradation and inefficiency under the power-law distribution of real-world datasets. For popular nodes with higher degrees, the number of h-hop neighbors is explosive, causing the node noise information in the sampled context to affect the accuracy of model anomaly detection. Secondly, sampling h-hop neighbors ignores the different importance of nodes in the graph structure. The importance of neighbors between nodes varies depending on their type. For example, closer neighbors will have a greater impact on nodes. However, this simple strategy can only view shared neighbors and exclusive neighbors equally when sampling context nodes.

[0008] (3) Fusion of structural information and temporal information. The existing method is to model the topological structure and temporal dynamic characteristics separately. The residual structural information may not be fully utilized, resulting in information loss and reduced model performance. Therefore, designing a method that can better integrate structural and temporal information is also an important challenge for abnormal edge detection;

[0009] Therefore, it is very necessary to provide a dynamic graph anomaly detection method based on graph embedding and diffusion sampling to overcome the shortcomings of existing methods. Summary of the invention

[0010] The purpose of the present invention is to provide a dynamic graph anomaly detection method based on graph embedding and diffusion sampling, which can improve the accuracy of abnormal edge detection and provide a new research idea for end-to-end dynamic graph anomaly detection method.

[0011] The technical solution adopted by the present invention is a dynamic graph anomaly detection method based on graph embedding and diffusion sampling, which is characterized by comprising the following steps:

[0012] Step S1: Get a dynamic graph dataset, and divide the edge indexes of all moments into an adjacency matrix set {A0…A t}, randomly initialize each node feature to X; where, represents the adjacency matrix at time t, represents a set of real numbers, n represents the number of nodes;

[0013] Step S2: within the time window, generate an adjacency matrix according to the edge index of each timestamp, and calculate all diffusion matrices within the window time t;

[0014] Step S3: Based on the diffusion matrix S at time t obtained in step S2 t , for the diffusion matrix S t Select the k most influential nodes in each row of , and get the node set of all node subgraphs at time t represents a set of real numbers, n represents the number of nodes;

[0015] Step S4: The sub-image window obtained in S3 is represented as Input into the node spatiotemporal information fusion module to obtain the node feature representation at each moment The spatiotemporal information fusion module includes a feature extraction module and a time information extraction module;

[0016] Step S5: negatively sample abnormal edges, and represent the abnormal edges, normal edges, and node features obtained in S4 at each moment for abnormal edge detection. Send it to the anomaly detector to get the anomaly scores of all edges;

[0017] Step S6: Based on the edge anomaly score obtained in S5, the loss function is calculated and the loss is minimized to perform model training.

[0018] Furthermore, in step S2, the original graph is used for diffusion operation to quantify the importance of nodes and nodes, and the k nodes with the greatest influence on the target node are obtained through the diffusion matrix. For a given static graph, the adjacency matrix in, represents a set of real numbers, n represents the number of nodes; the graph diffusion matrix is ​​defined as:

[0019]

[0020] Where S represents the graph diffusion matrix, T m represents the mth power of the generalized transfer matrix T, θ m Represents the weight coefficient that determines the ratio of global to local information; represents a set of real numbers, and n represents the number of nodes.

[0021] Furthermore, in step S3, the node set of all node subgraphs at time t The formula is as follows:

[0022]

[0023] in, represents the node set of the subgraph at time t, S t represents the diffusion matrix at time t, topk(·) represents the diffusion matrix S t Select k nodes with the greatest influence from each row of ;

[0024] After calculating the subgraph nodes corresponding to all window moments, the obtained subgraph windows are represented as Among them, t represents the current time, w represents the size of the time window, represents the node set of the subgraph at time t, and v represents the time in the time window.

[0025] Furthermore, in step S4, the feature extraction module uses a two-layer GCN model to obtain the node features of each timestamp, and randomly initializes the node features. The initial features are represented by X, and the formula is as follows:

[0026]

[0027] in, represents the sum of the adjacency matrix and the identity matrix, W represents the trainable weight matrix, express The diagonal matrix, Z (0) Represents the feature value after a layer of GCN, W (0) Represents the trainable weight parameters of the first layer of GCN, Z (1) Represents the feature value after two layers of GCN, W (1) represents the trainable weight parameter of the second layer GCN, and Relu(·) represents the activation function Relu;

[0028] Use GCN to obtain node features at each moment and time window in, represents a real number set, n represents the number of nodes, d str represents the dimension after extracting structural information, Represents the structural feature representation of the node at time v in the time window, Represents the set of node structure features represented by the time window of size w at time t, where the range of v is [tw,t].

[0029] Furthermore, in step S4, the time information extraction module includes a GRU module and a self-attention mechanism module. First, the time information is extracted by the GRU module, and the formula is as follows:

[0030]

[0031]

[0032] h t =(1-z t )⊙h t-1 +z t ⊙h' t (9)

[0033] Among them, ⊙ represents the element-by-element product operation, W Z , U z , b z , W r , Ur , b r , W h , U h , b h is a trainable parameter in the GRU unit, W Z represents the weight matrix of the update gate, U z represents the recursive weight matrix of the update gate for the hidden state at the current time step, b z represents the bias vector of the update gate, W r represents the weight matrix of the reset gate, U r represents the recursive weight matrix of the reset gate for the hidden state at the current time step, b r represents the bias vector of the reset gate, W h The weight matrix representing the candidate hidden state, U h Represents the recursive weight matrix of candidate hidden states for the hidden state at the current time step, b h represents the bias vector of the candidate hidden state, It represents the node structure feature representation after two layers of GCN at time t, h t-1 represents the state at time t-1, z t represents the update gate, r t represents the reset gate, h' t represents the candidate hidden state, h t From its previous state h t-1 Calculated, represents the state passed to time t, σ(·) represents the sigmoid activation function, and tanh(·) represents the tanh function; the GRU network takes the features of each timestamp as input and inputs the output of the current timestamp to the next timestamp.

[0034] Furthermore, in step S4, after extracting the time information through the GRU module, an adaptive self-attention mechanism module is added to obtain the time information. Get the final node embedding representation Among them, h v represents the spatiotemporal feature representation of the node at time v in the time window, H t Represents the node features in a time window of size w at time t, where the range of v is [tw,t];

[0035] The formula is as follows:

[0036]

[0037] Among them, W Q , W K , W V represents the trainable parameters, d embrepresents the spatiotemporal feature representation dimension of the node; Q, K, V represent the three matrices of the attention mechanism, which are used to calculate the similarity between vectors; softmax(·) represents the softmax function, Pooling(·) represents the average pooling operation, represents the structural feature representation of the node at time t, H t represents the node features at time t in the time window, K T Represents the transposed matrix of K, and finally gets represents the final node feature representation, represents a real number set, n represents the number of nodes, d emb The spatiotemporal features of the nodes represent the dimension.

[0038] Furthermore, in step S5, the node feature representation of abnormal edge detection in, represents a real number set, n represents the number of nodes, d emb The spatiotemporal features of the nodes represent the dimension. The anomaly detector uses a neural network with a fully connected layer. The formula for the edge anomaly score between nodes i and j is as follows:

[0039]

[0040] Where mlp(·) represents a two-layer fully connected neural network, score(i,j) represents the abnormal score of the edge between nodes i and j, It is the final feature representation of node i, j for abnormal edge detection.

[0041] Furthermore, in step S6, the loss function adopts a cross entropy loss function, and its formula is as follows:

[0042]

[0043] Among them, N represents the sum of negative sampling and positive sampling edges, y represents the current sequence number of the traversed edge sum, represents the loss function, and They represent the anomaly scores of the positive sample edges and negative sample edges between nodes i and j respectively.

[0044] Furthermore, in step S2, the method for calculating the diffusion matrix is ​​modified according to the Laplace operator, and the diffusion matrix S at time t is t The formula is as follows:

[0045]

[0046] Among them, S trepresents the diffusion matrix at time t, D represents the diagonal matrix, α∈(0,1), α represents the propagation probability, A represents the adjacency matrix of the graph, I n Represents the identity matrix with n nodes; in, represents a set of real numbers, and n represents the number of nodes.

[0047] The beneficial effects of the present invention are as follows: the present application provides a dynamic graph anomaly detection method based on graph embedding and diffusion sampling, which introduces diffusion sampling to sample an exclusive subgraph for each node to obtain node features, thereby reducing the influence of useless neighbor noise on node embedding representation, further introducing attention mechanism and gated recurrent neural network and adding structural information residual connection, thereby more effectively integrating structure and time information, obtaining the final representation and calculating the anomaly score of the edge. The present invention is an innovative exploration of dynamic graph anomaly detection methods, which is conducive to the exploration of complex network structures and can promote the development of complex network research fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1 It is a flow chart of an embodiment of the method of the present invention.

[0050] Figure 2 Schematic diagram of abnormal edge detection according to the method of the present invention. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] Example 1

[0053] The embodiment of the present invention provides a dynamic graph anomaly detection method based on graph embedding and diffusion sampling, such as Figures 1-2 As shown, the following steps are included:

[0054] Step S1: Get a dynamic graph dataset, and divide the edge indexes of all moments into an adjacency matrix set {A0…A t}, randomly initialize each node feature to X; where, represents the adjacency matrix at time t, represents a set of real numbers, and n represents the number of nodes.

[0055] Step S2: In the time window, an adjacency matrix is ​​generated according to the edge index of each timestamp, and then all diffusion matrices within the window are calculated. Different from the neighbor sampling of the traditional method, the present invention uses the original graph for diffusion operation. Through graph diffusion, the importance between nodes is quantified. For example, the importance of node i to node j is the i row and j column of the diffusion matrix. The larger the value, the higher the importance. Thus, the k nodes with the greatest influence on the target node are obtained through the diffusion matrix. This reduces the noise impact of useless neighbor nodes and improves the accuracy of anomaly detection. For a given static graph adjacency matrix in, represents a set of real numbers, and n represents the number of nodes. The graph diffusion matrix is ​​defined as:

[0056]

[0057] Where S represents the graph diffusion matrix, T m represents the mth power of the generalized transfer matrix T, θ m Represents the weight coefficient that determines the ratio of global to local information; represents a set of real numbers, and n represents the number of nodes.

[0058] For the diffusion operation adopted in this embodiment, in order to avoid multiple iterations, the method for calculating the diffusion matrix is ​​modified according to the Laplace operator. The diffusion matrix S at time t is t The formula is as follows:

[0059]

[0060] Among them, S t represents the diffusion matrix at time t, D represents the diagonal matrix, α∈(0,1), α represents the propagation probability, A represents the adjacency matrix of the graph, I n Represents the identity matrix with n nodes; in, represents a set of real numbers, and n represents the number of nodes.

[0061] Step S3: Based on the diffusion matrix S at time t obtained in step S2 t , for the diffusion matrix S t Select the k most influential nodes in each row of , and get the node set of all node subgraphs at time t represents a real number set, and n represents the number of nodes. The formula is as follows:

[0062]

[0063] in, represents the node set of the subgraph at time t, S t represents the diffusion matrix at time t, topk(·) represents the diffusion matrix S t For each row of , k nodes with the greatest influence are selected.

[0064] After calculating the subgraph nodes corresponding to all window moments, the obtained subgraph windows are represented as Among them, t represents the current time, w represents the size of the time window, represents the node set of the subgraph at time t, and v represents the time in the time window.

[0065] Step S4: The sub-image window obtained in S3 is represented as The information is input into the node spatiotemporal information fusion module to obtain the feature representation of each moment. The spatiotemporal information fusion module includes a feature extraction module and a time information extraction module.

[0066] The feature extraction module uses a two-layer GCN model to obtain the node features of each timestamp, randomly initializes the node features, and the initial features are represented by X. The formula is as follows:

[0067]

[0068] in, represents the sum of the adjacency matrix and the identity matrix, W represents the trainable weight matrix, express The diagonal matrix, Z (0) Represents the feature value after a layer of GCN, W (0) Represents the trainable weight parameters of the first layer of GCN, Z (1) Represents the feature value after two layers of GCN, W (1) represents the trainable weight parameter of the second layer GCN, and Relu(·) represents the activation function Relu.

[0069] Using graph convolutional layers, each node can aggregate embeddings from its neighbors. By stacking graph convolutional layers in a neural network to extract structural information, the target node embedding is extracted separately in the subgraph sampled from each node to obtain the node features at each moment. and time window in, represents a real number set, n represents the number of nodes, d str represents the dimension after extracting structural information, Represents the structural feature representation of the node at time v in the time window, Represents the set of node structure features represented by the time window of size w at time t, where the range of v is [tw,t].

[0070] The temporal information extraction module includes a GRU (Gated Recurrent Unit) module and a self-attention mechanism module. First, the temporal information is captured through the GRU module and the gradient vanishing and exploding problems are alleviated. The formula is as follows:

[0071]

[0072] h t =(1-z t )⊙h t-1 +z t ⊙h' t (9)

[0073] Among them, ⊙ represents the element-by-element product operation, W Z , U z , b z , W r , U r , b r , W h , U h , b h is a trainable parameter in the GRU unit, W Z represents the weight matrix of the update gate, U z represents the recursive weight matrix of the update gate for the hidden state at the current time step, b z represents the bias vector of the update gate, W r represents the weight matrix of the reset gate, U r represents the recursive weight matrix of the reset gate for the hidden state at the current time step, b r represents the bias vector of the reset gate, W h The weight matrix representing the candidate hidden state, U h Represents the recursive weight matrix of candidate hidden states for the hidden state at the current time step, b h Bias vector representing the candidate hidden states. It represents the node structure feature representation after two layers of GCN at time t, h t-1 represents the state at time t-1, z t represents the update gate, r t represents the reset gate, h' t represents the candidate hidden state, h t From its previous state h t-1Calculated from, represents the state passed to time t, σ(·) represents the sigmoid activation function, through which the data can be converted to numbers in the range of [0,1], tanh(·) represents the tanh function, through which the data can be converted to values ​​in the range of [-1,1]. The GRU network takes the features of each timestamp as input and inputs the output of the current timestamp to the next timestamp.

[0074] Secondly, in order to fully consider the influence of historical information and the periodicity of user behavior, an adaptive self-attention mechanism module is added after the GRU module. Get the final node embedding representation Among them, h v represents the spatiotemporal feature representation of the node at time v in the time window, H t Represents the node features in a time window of size w at time t, where the range of v is [tw,t].

[0075] The formula is as follows:

[0076]

[0077] Among them, W Q , W K , W V represents the trainable parameters, d emb represents the spatiotemporal feature representation dimension of the node; Q, K, V represent the three matrices of the attention mechanism, which are used to calculate the similarity between vectors; softmax(·) represents the softmax function, Pooling(·) represents the average pooling operation, represents the structural feature representation of the node at time t, H t represents the node features at time t in the time window, K T represents the transposed matrix of K. Finally, we get represents the final node feature representation, represents a real number set, n represents the number of nodes, d emb The spatiotemporal features of the nodes represent the dimension.

[0078] Step S5: negatively sample abnormal edges, and then represent the abnormal edges, normal edges, and node features obtained in S4 at each moment for abnormal edge detection. The anomaly scores of all edges are obtained by sending them to the anomaly detector, where represents a real number set, n represents the number of nodes, d emb The spatiotemporal feature representation dimension of the node is represented. The anomaly detector adopts a neural network with a fully connected layer. The formula for the edge anomaly score between nodes i and j is as follows:

[0079]

[0080] Where mlp(·) represents a two-layer fully connected neural network, score(i,j) represents the abnormal score of the edge between nodes i and j, It is the final feature representation of node i, j for abnormal edge detection.

[0081] Step S6: Based on the edge anomaly score obtained in S5, the loss function is calculated and the loss is minimized to perform model training. Since the label value is only 0 and 1, the loss function adopts the cross entropy loss function, and its formula is as follows:

[0082]

[0083] Among them, N represents the sum of negative sampling and positive sampling edges, y represents the current sequence number of the traversed edge sum, represents the loss function, and They represent the anomaly scores of positive and negative edges between nodes i and j, respectively.

[0084] Based on the four real-world network datasets in Table 1, a dynamic graph anomaly detection method based on graph embedding and diffusion sampling is used.

[0085] Table 1 Four real-world network datasets

[0086] Dataset Number of nodes Number of edges Average degree UCIMessages 1899 13838 14.57 Email-DNC 1866 39264 42.08 Bitcoin-Alpha 3777 24173 12.80 Bitcoin-OTC 5881 35588 12.10

[0087] The present invention uses different autoencoder models and graph neural network models to compare the results according to different types of data sets. The selected models are: AddGraph (https: / / www.ijcai.org / Proceedings / 2019 / 614), StrGNN (https: / / dl.acm.org / doi / 10.1145 / 3459637.3481955), TADDY (https: / / ieeexplore.ieee.org / document / 9599560 / ), RegraphGAN (https: / / doi.org / 10.1016 / j.neunet.2023.07.026). The experimental results of the present invention and other models are shown in Table 2, where the best value is in bold. AUC (Area Under Curve) is selected as the evaluation index. Compared with previous deep learning methods, the proposed method performs well, especially on the UCI Messages dataset, and only second to RegraphGAN in the 10% abnormal case on Bitcoin-OTC.

[0088] According to the AUC performance comparison in Table 2, compared with the best baseline, the present invention shows an average performance gap of 2.2%. It can be seen that the methods of the present invention are superior to the previous methods.

[0089] Table 2 Comparison of the present invention with other deep learning anomaly detection models

[0090]

[0091] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0092] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A dynamic graph anomaly detection method based on graph embedding and diffusion sampling, characterized in that: The following steps are involved: Step S1: Get a dynamic graph dataset, and divide the edge indexes of all moments into an adjacency matrix set {A0…A t }, randomly initialize each node feature to X; where, represents the adjacency matrix at time t, represents a set of real numbers, n represents the number of nodes; Step S2: within the time window, generate an adjacency matrix according to the edge index of each timestamp, and calculate all diffusion matrices within the window time t; Step S3: Based on the diffusion matrix S at time t obtained in step S2 t , for the diffusion matrix S t Select the k most influential nodes in each row of , and get the node set of all node subgraphs at time t represents a set of real numbers, n represents the number of nodes; Step S4: The sub-image window obtained in S3 is represented as Input into the node spatiotemporal information fusion module to obtain the node feature representation at each moment The spatiotemporal information fusion module includes a feature extraction module and a time information extraction module; Step S5: negatively sample abnormal edges, and represent the abnormal edges, normal edges, and node features obtained in S4 at each moment for abnormal edge detection. Send it to the anomaly detector to get the anomaly scores of all edges; Step S6: Based on the edge anomaly score obtained in S5, the loss function is calculated and the loss is minimized to perform model training.

2. According to claim 1, a dynamic graph anomaly detection method based on graph embedding and diffusion sampling is characterized in that: In step S2, the original graph is used for diffusion operation to quantify the importance of nodes and nodes, and the k nodes with the greatest influence on the target node are obtained through the diffusion matrix. For a given static graph adjacency matrix in, represents a set of real numbers, n represents the number of nodes; the graph diffusion matrix is ​​defined as: Where S represents the graph diffusion matrix, T m represents the mth power of the generalized transfer matrix T, θ m Represents the weight coefficient that determines the ratio of global to local information; represents a set of real numbers, and n represents the number of nodes.

3. The method for dynamic graph anomaly detection based on graph embedding and diffusion sampling according to claim 1, characterized in that: In step S3, the node set of all node subgraphs at time t The formula is as follows: in, represents the node set of the subgraph at time t, S t represents the diffusion matrix at time t, topk(·) represents the diffusion matrix S t Select k nodes with the greatest influence from each row of ; After calculating the subgraph nodes corresponding to all window moments, the obtained subgraph windows are represented as Among them, t represents the current time, w represents the size of the time window, represents the node set of the subgraph at time t, and v represents the time in the time window.

4. The method for dynamic graph anomaly detection based on graph embedding and diffusion sampling according to claim 1, characterized in that: In step S4, the feature extraction module uses a two-layer GCN model to obtain the node features of each timestamp, and randomly initializes the node features. The initial features are represented by X, and the formula is as follows: in, represents the sum of the adjacency matrix and the identity matrix, W represents the trainable weight matrix, express The diagonal matrix, Z (0) Represents the feature value after a layer of GCN, W (0) Represents the trainable weight parameters of the first layer of GCN, Z (1) Represents the feature value after two layers of GCN, W (1) represents the trainable weight parameter of the second layer GCN, and Relu(·) represents the activation function Relu; Use GCN to obtain node features at each moment and time window in, represents a real number set, n represents the number of nodes, d str represents the dimension after extracting structural information, Represents the structural feature representation of the node at time v in the time window, Represents the set of node structure features represented by the time window of size w at time t, where the range of v is [tw,t].

5. A method for detecting anomalies in dynamic graphs based on graph embedding and diffusion sampling according to claim 1, characterized in that: In step S4, the time information extraction module includes a GRU module and a self-attention mechanism module. First, the time information is extracted by the GRU module, and the formula is as follows: h t =(1-z t )⊙h t-1 +z t ⊙h' t (9) Among them, ⊙ represents the element-by-element product operation, W Z , U z , b z , W r , U r , b r , W h , U h , b h is a trainable parameter in the GRU unit, W Z represents the weight matrix of the update gate, U z represents the recursive weight matrix of the update gate for the hidden state at the current time step, b z represents the bias vector of the update gate, W r represents the weight matrix of the reset gate, U r represents the recursive weight matrix of the reset gate for the hidden state at the current time step, b r represents the bias vector of the reset gate, W h The weight matrix representing the candidate hidden state, U h Represents the recursive weight matrix of candidate hidden states for the hidden state at the current time step, b h represents the bias vector of the candidate hidden state, It represents the node structure feature representation after two layers of GCN at time t, h t-1 represents the state at time t-1, z t represents the update gate, r t represents the reset gate, h' t represents the candidate hidden state, h t From its previous state h t-1 Calculated, represents the state passed to time t, σ(·) represents the sigmoid activation function, and tanh(·) represents the tanh function; the GRU network takes the features of each timestamp as input and inputs the output of the current timestamp to the next timestamp.

6. A method for detecting anomalies in dynamic graphs based on graph embedding and diffusion sampling according to claim 5, characterized in that: In step S4, after extracting the time information through the GRU module, an adaptive self-attention mechanism module is added to obtain the Get the final node embedding representation Among them, h v represents the spatiotemporal feature representation of the node at time v in the time window, H t Represents the node features in a time window of size w at time t, where the range of v is [tw,t]; The formula is as follows: Among them, W Q , W K , W V represents the trainable parameters, d emb represents the spatiotemporal feature representation dimension of the node; Q, K, V represent the three matrices of the attention mechanism, which are used to calculate the similarity between vectors; softmax(·) represents the softmax function, Pooling(·) represents the average pooling operation, represents the structural feature representation of the node at time t, H t represents the node features at time t in the time window, K T Represents the transposed matrix of K, and finally gets represents the final node feature representation, represents a real number set, n represents the number of nodes, d emb The spatiotemporal features of the nodes represent the dimension.

7. The method for dynamic graph anomaly detection based on graph embedding and diffusion sampling according to claim 1, characterized in that: In step S5, the node feature representation of abnormal edge detection in, represents a real number set, n represents the number of nodes, d emb The spatiotemporal features of the nodes represent the dimension. The anomaly detector uses a neural network with a fully connected layer. The formula for the edge anomaly score between nodes i and j is as follows: Where mlp(·) represents a two-layer fully connected neural network, score(i,j) represents the abnormal score of the edge between nodes i and j, It is the final feature representation of node i, j for abnormal edge detection.

8. The method for dynamic graph anomaly detection based on graph embedding and diffusion sampling according to claim 1, characterized in that: In step S6, the loss function adopts the cross entropy loss function, and its formula is as follows: Among them, N represents the sum of negative sampling and positive sampling edges, y represents the current sequence number of the traversed edge sum, represents the loss function, and They represent the anomaly scores of the positive sample edges and negative sample edges between nodes i and j respectively.

9. A method for detecting anomalies in dynamic graphs based on graph embedding and diffusion sampling according to claim 1 or 2, characterized in that: In step S2, the method for calculating the diffusion matrix is ​​modified according to the Laplace operator, and the diffusion matrix S at time t is t The formula is as follows: Among them, S t represents the diffusion matrix at time t, D represents the diagonal matrix, α∈(0,1), α represents the propagation probability, A represents the adjacency matrix of the graph, I n Represents the identity matrix with n nodes; in, represents a set of real numbers, and n represents the number of nodes.

Citation Information

Patent Citations

  • Internet of Things time series data anomaly detection method and system based on dynamic graph attention

    CN118094427A

  • Dynamic word embeddings

    US20180157644A1

Cited By

  • Medicine appearance defect detection method and related equipment

    CN121599973A

  • A method for detecting defects in the appearance of a pharmaceutical product and related apparatus

    CN121599973B