A compound anomaly detection method and system enhanced by random substructure features
By adopting multi-scale substructure random sampling and feature fusion methods in compound abnormality detection, the inefficiency problem caused by the lack of node-level characteristics in the prior art is solved, and higher abnormal detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202410655411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-05-24
AI Technical Summary
Existing compound anomaly detection methods are inefficient when lacking node-level characteristics and cannot effectively capture the potential abnormal structure of the compound.
Through random sampling of multiple substructures of different sizes, the neighborhood features (substructure features) of each node are mined, and node representations are generated through feature fusion, activation and multi-hop feature aggregation, and anomaly is finally detected through graph-level representations.
It improves the accuracy of compound abnormality detection, can effectively capture potential abnormal structures, reduce dependence on node-level characteristics, and improve detection efficiency.
Smart Images

Figure CN118538312B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer service technology, and in particular relates to a compound anomaly detection method and system enhanced by random substructure features. Background Art
[0002] Graph anomaly detection is a powerful tool that can be applied in many fields, such as fraudulent user detection in social networks, fraudulent transactions in economics, etc.
[0003] In chemical engineering and biology, many molecular compounds with unknown properties are constantly being synthesized. Whether there is an automated method to preliminarily classify and screen the newly synthesized molecules and determine whether they may be molecules with abnormal properties, such as toxic substances, is a very important research, which can save a lot of energy for researchers in related fields. Compounds can be represented by graphs, for example, atoms and bonds are represented as nodes and edges, which makes related technologies in the field of complex network analysis naturally competent for these tasks, and graph anomaly detection methods are good tools for detecting abnormal compounds.
[0004] In recent years, the development of Graph Neural Network (GNN) has improved the accuracy of anomaly detection and greatly enriched the application scenarios of anomaly detection. Most of the existing methods are based on representation learning, that is, first mining the features of the graph and embedding the graph network in non-Euclidean space into a low-dimensional space to obtain the corresponding vector representation, and then classifying the vector into two categories: normal or abnormal. The current mainstream methods basically improve the ability of the method from two perspectives. One is to improve the ability of graph representation learning, mining more valuable graph features to obtain a low-dimensional representation with better expression ability. On the other hand, enhancing the ability of the classifier will enable the vector representing the graph to be more accurately classified. However, some methods rely on the node-level features provided by the data set for message transmission. When this information is missing, they usually initialize the node-level features through the degree of the node, random vectors, etc., and this method will be greatly reduced in efficiency when used to detect objects such as compounds and molecules. On the one hand, these research objects usually do not have such features, and on the other hand, their topological structure is more important than the node-level features themselves.
[0005] To this end, the present invention proposes a compound anomaly detection method and system enhanced by random substructure features. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention proposes a compound anomaly detection method and system enhanced by random substructure features. Starting from the essence of the anomaly, that is, capturing the abnormal substructure hidden in the graph, specifically, the neighborhood features (substructure features) of each node are mined through random sampling of substructures of multiple different sizes. These substructure features can reflect the potential abnormal structure to a certain extent, providing necessary information for compound anomaly detection, thereby improving the accuracy of anomaly detection.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A method for detecting anomalies in compounds enhanced by random substructure features comprises the following steps:
[0009] Sampling multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths;
[0010] Counting the number of non-repeated nodes in the walking sequence and concatenating them to obtain substructure features of the nodes;
[0011] Standardizing the substructure features using the L2 norm;
[0012] Use LeakyReLU to activate the standardized substructure features;
[0013] Performing feature fusion on the activated substructure features and node-level features;
[0014] Perform matrix multiplication between the adjacency matrix and the fused features to obtain a local feature aggregation, and perform multi-hop feature aggregation on the aggregated features through GraphSage to obtain the representation of all nodes;
[0015] The learned node representations are constrained by the graph structure to obtain the reconstruction loss;
[0016] Based on the node representation, the graph level representation is obtained through the graph readout function;
[0017] Classifying the graph-level representation to obtain anomaly factor scores, and obtaining a label loss based on the anomaly factor scores;
[0018] The reconstruction loss and the label loss are jointly optimized to construct a compound anomaly detection model;
[0019] Based on the compound anomaly detection model, the anomaly detection task of the compound to be tested is achieved.
[0020] Preferably, the method of sampling multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths includes:
[0021] Through the Node2vec method, each node in the graph is used as the starting node to perform multiple biased random walks of different lengths to obtain walk sequences of different lengths.
[0022] Preferably, the method of fusing the activated substructure features with the node-level features comprises:
[0023] If the node has node-level features, where the node-level features include: node attributes and node labels, they are concatenated in the order of [node attributes, node labels, substructure features] and then mapped to the new feature space by the fully connected layer to complete feature fusion.
[0024] Preferably, the method of setting constraints on the learned node representation through the graph structure includes:
[0025]
[0026] In the formula, is the predicted adjacency matrix, and H is the low-dimensional node representation matrix.
[0027] Preferably, the method for classifying the graph-level representation includes:
[0028]
[0029] Where Sigmoid(·) is the activation function, W3 and W4 are both learnable weight matrices, b3 and b4 are bias terms, is the predicted anomaly score, For the picture Graph-level representation of .
[0030] The present invention also provides a compound anomaly detection system enhanced by random substructure features, comprising: a sampling module, a splicing module, a normalization module, an activation module, a fusion module, an aggregation module, a constraint module, a readout module, a scoring module, a construction module and a detection module;
[0031] The sampling module is used to sample multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths;
[0032] The splicing module is used to count the number of non-repeated nodes in the walking sequence and splice them to obtain the substructure characteristics of the nodes;
[0033] The standardization module is used to standardize the substructure features using the L2 norm;
[0034] The activation module is used to activate the standardized substructure features using LeakyReLU;
[0035] The fusion module is used to fuse the activated substructure features with the node-level features;
[0036] The aggregation module is used to perform matrix multiplication on the adjacency matrix and the fused features to obtain a local feature aggregation, and perform multi-hop feature aggregation on the aggregated features through GraphSage to obtain all node representations;
[0037] The constraint module is used to set constraints on the learned node representation through the graph structure to obtain the reconstruction loss;
[0038] The readout module is used to obtain a graph-level representation based on the node representation through a graph readout function;
[0039] The scoring module is used to classify the graph-level representation, obtain anomaly factor scores, and obtain label losses based on the anomaly factor scores;
[0040] The construction module is used to jointly optimize the reconstruction loss and the label loss to construct a compound anomaly detection model;
[0041] The detection module is used to implement anomaly detection tasks of the compounds to be tested based on the compound anomaly detection model.
[0042] Preferably, the process of sampling multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths includes:
[0043] Through the Node2vec method, each node in the graph is used as the starting node to perform multiple biased random walks of different lengths to obtain walk sequences of different lengths.
[0044] Preferably, the process of fusing the activated substructure features with the node-level features comprises:
[0045] If the node has node-level features, where the node-level features include: node attributes and node labels, they are concatenated in the order of [node attributes, node labels, substructure features] and then mapped to the new feature space by the fully connected layer to complete feature fusion.
[0046] Preferably, the process of setting constraints on the learned node representations through the graph structure includes:
[0047]
[0048] In the formula, is the predicted adjacency matrix, and H is the low-dimensional node representation matrix.
[0049] Preferably, the process of classifying the graph-level representation includes:
[0050]
[0051] Where Sigmoid(·) is the activation function, W3 and W4 are both learnable weight matrices, b3 and b4 are bias terms, is the predicted anomaly score, For the picture Graph-level representation of .
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] The present invention starts from the essence of anomalies, that is, capturing abnormal substructures hidden in the graph. Specifically, the neighborhood features (substructure features) of each node are mined through random sampling of substructures of multiple different sizes. These substructure features can reflect the potential abnormal structure to a certain extent, providing necessary information for compound anomaly detection, thereby improving the accuracy of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0055] Figure 1 It is a flow chart of a compound anomaly detection method enhanced by random substructure features according to an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] Embodiment 1
[0059] like Figure 1As shown, the present invention provides a compound anomaly detection method enhanced by random substructure features. The method starts from the nature of the anomaly and captures the abnormal substructure hidden in the compound structure graph network. For example, the mutation of the p53 protein may cause it to have abnormal biological functions or participate in pathological processes, and its ability to bind to DNA is affected, so that it cannot effectively activate downstream gene expression, and ultimately leads to abnormal cell proliferation and cancer. The neighborhood features (substructure features) of each node are mined through random sampling of substructures of multiple different scales. These substructure features can reflect the potential abnormal structure to a certain extent and provide necessary information for compound anomaly detection. The following steps are included:
[0060] Multi-scale substructure sampling. Through Node2vec, each node in the graph is used as the starting node to perform multiple biased random walks of different lengths to obtain walk sequences of different lengths. Then, the number of non-repeating nodes in each sequence is counted. The number of non-repeating nodes in multiple walk sequences of each node is concatenated as the substructure feature of the node. Because it is a substructure obtained based on random walk sampling, this feature is called a random substructure feature.
[0061] Since different nodes have different neighborhood characteristics, the numerical differences in different dimensions of the features obtained in the above multi-scale substructure sampling process may be large. Therefore, the feature is standardized with the help of the L2 norm.
[0062] Subsequently, LeakyReLU is used to activate the standardized attributes, and then feature fusion is performed. If the node has node-level attributes (including the hot code representation of the node label), it is concatenated with the activated standardized substructure representation, and then mapped to the new feature space by the fully connected layer to complete the feature fusion. Then, with the help of the adjacency matrix, matrix multiplication is performed between the adjacency matrix and the fused features to complete a local feature aggregation, and the feature is aggregated through multi-hop features through GraphSage to obtain a low-dimensional node representation matrix.
[0063] The model is trained through two downstream tasks: graph reconstruction and anomaly detection.
[0064] Graph reconstruction refers to constraining the learned low-dimensional representation through the graph structure, performing matrix multiplication on it and its own transpose, hoping that it will be close to the adjacency matrix.
[0065] Anomaly detection is also the final output of the present invention. First, a graph-level representation is obtained through a graph readout function, and then the representation is classified to determine whether the original image is abnormal or normal.
[0066] The loss functions of the above two downstream tasks are merged through a joint optimization strategy, and the model training is completed.
[0067] The specific implementation steps of the method are as follows:
[0068] Step S1: Multi-scale substructure sampling. Define this process as g(A, L, s)→H str , where A is the adjacency matrix of a graph, L = {l1, l2, ..., l m} is a sequence of sampling scales, and s is the number of repeated samplings at each sampling scale. Based on the Node2vec method, a total of s×|L| walk sequences of each element with a length of L can be obtained for each node.
[0069] Step S2: Count the number of non-repeated nodes in each walking sequence. Thus, each node can obtain a vector with a dimension of S as the initial substructure feature. For node i, this feature is represented by X i .
[0070] Step S3: Since different nodes have different neighborhood features, the numerical differences of the features obtained in step S2 in different dimensions may be large. Therefore, the features are normalized using the L2 norm:
[0071]
[0072]
[0073] where X′ i That is, the substructure feature of node i after standardization. For all nodes in a graph, a substructure feature matrix X can be obtained. str ∈R |V|×|s| , |V| is the number of nodes in the graph.
[0074] Step S4: Use LeakyReLU to activate the standardized substructure features, which is expressed in the matrix as follows:
[0075] H str =LeakyReLU(W1X str′ +b1),
[0076] Where W1 and b1 are the learnable weight matrix and bias term respectively.
[0077] Step S5, feature fusion: If the node has node-level attributes (including the hot code representation of the node label), it is combined with the activated substructure representation H obtained in step S4 str Splice to get H c , and then get H through the full connection layer and activation function c′ Then, a local feature aggregation is performed with the help of the adjacency matrix to obtain H a , the above process is expressed by the following formula:
[0078] H c′ =LeakyReLU(W2Hc+b2),
[0079] H a =A·H c′ ,
[0080] Step S6, message passing and feature aggregation: perform feature aggregation through GraphSage and obtain a low-dimensional node representation matrix form H.
[0081] Step S7, graph reconstruction constraints: impose certain constraints on the learned H through the graph structure, and specifically adopt the graph reconstruction method:
[0082]
[0083] in is the predicted adjacency matrix, which is compared with the actual adjacency matrix A. The difference between the two can be measured by MSELoss. The loss function is defined as the reconstruction loss
[0084]
[0085] Step S8: In step S6, for the image The representation matrix of all its nodes has been obtained Where |V i | indicates The number of nodes is now converted into a 1×d vector by reading out the function γ(·), that is, Graph-level representation of
[0086]
[0087] Step S9, abnormal factor scoring: First, the graph-level representation is classified by MLP, and the process is expressed as follows:
[0088]
[0089] Among them, Sigmoid(·) is the activation function, W3 and W4 are both learnable weight matrices, b3 and b4 are bias terms, Compound Graph Network The predicted anomaly score indicates the probability that the compound is considered anomaly. Further, it is compared with the actual anomalies in the dataset, and the information loss can be calculated by biased BCELoss. The loss function is defined as label loss
[0090]
[0091] In the above formula, is the number of graphs, w p represents the weight applied to normal samples.
[0092] Step S10, joint optimization: The reconstruction loss and label loss are added together, where Apply a weight greater than 0 and less than or equal to 1. Get the final optimization goal:
[0093]
[0094] Where α is the weight, is the L2 regularization term to avoid overfitting.
[0095] Step S11: After completing the model training process of S1-S10, all parameters are fixed. These parameters constitute a compound anomaly detection model, denoted as model GRSN (·), for a new input The anomaly detection result can be obtained based on the following formula:
[0096]
[0097] The calculation method is the same as that in S9 are consistent.
[0098] Embodiment 2
[0099] The data of this embodiment are selected from TUDataset, and six data sets including HSE, DHFR, p53, AIDS, Proteins_full, and MMP are selected. They are all bioinformatics data sets, and their molecular structures are modeled by graphs. In addition, in order to further evaluate the generalization ability of the present invention, two social network data sets are also taken into account, namely REDDIT (online community comment interaction data set) and IMDB (movie collaboration data set). In this embodiment, the area under the ROC curve and the coordinate axis (Area Under Curve, AUC) is an evaluation index. The higher the value, the better. Each data set runs five independent experiments and counts the mean and standard deviation of the corresponding AUC, expressed as a percentage.
[0100] The anomaly detection evaluation results of GRSN on these eight datasets are shown in Table 1.
[0101] Table 1 AUC performance evaluation results of GRSN
[0102]
[0103] In order to discuss the performance level of the methods, we selected ten methods for comparison, namely GraphSage, GAT, GCN, iGAD, HimNet, TUAF, GLocalKD, MssGAD, OCGIN, and OCGTL. Table 2 shows the evaluation results of all methods on eight datasets.
[0104] Table 2 Performance evaluation results of all methods
[0105]
[0106] GRSN surpasses all methods on DHFR, IMDB-BINARY, p53, REDDIT-BINARY, Proteins_full, and MMP datasets, with AUC averages of 74.8, 74.9, 77.4, 86.6, 79.7, and 85.8, respectively. Although GRSN is not the best on the AIDS and HSE datasets, it still shows the performance level of the first echelon. On the AIDS dataset, GRSN's AUC index is 0.1 lower than MssGAD, but its standard deviation is only 0.1, which shows that GRSN is very stable and has high robustness; on the HSE dataset, iGAD's AUC is 2.1 higher than GRSN.
[0107] It is worth mentioning that the three datasets of HSE, p53, and MMP do not contain node-level features, but each node has label information. In practical applications, various methods usually use one-hot encoding to embed node labels as initial node-level features for message transmission; the two datasets of IMDB and REDDIT do not contain node labels, which means that no node-level features can be directly obtained, so some methods cannot run on these datasets (indicated by N / A), and some methods use the degree of the node to initialize the node-level features (the method of HimNet). GRSN obtains the neighborhood information of each node through multi-scale random substructure sampling, and uses this information as part of the node-level features, which can be used to initialize node-level features and as a supplement to the original features. This ability makes GRSN perform well on these datasets (except the HSE dataset), and has excellent recognition capabilities for abnormal graphs and abnormal compounds.
[0108] If the multi-scale random substructure sampling is removed from GRSN (denoted as GRSN-P), and then replaced with random equal-dimensional vectors (denoted as GRSN-R), the same experimental process as above can be used to verify the effectiveness of multi-scale random substructure sampling. The results are shown in Table 3.
[0109] Table 3 Ablation experiment results on multi-scale random substructure sampling
[0110]
[0111] Obviously, both GRSN-P and GRSN-R have different degrees of performance degradation compared with GRSN. Among them, GRSN-P has a performance degradation of at most 6.7% compared with GRSN (MMP dataset), and GRSN-R has a performance degradation of at most 21.4% compared with GRSN.
[0112] In summary, the multi-scale random substructure sampling introduced in the present invention is effective. It enables the present invention to not rely on node-level features, and its comprehensive performance in compound anomaly detection surpasses existing cutting-edge methods.
[0113] Embodiment 3
[0114] The present invention also provides a compound anomaly detection system enhanced by random substructure features, comprising: a sampling module, a splicing module, a normalization module, an activation module, a fusion module, an aggregation module, a constraint module, a readout module, a scoring module, a construction module and a detection module;
[0115] The sampling module is used to sample multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths;
[0116] The splicing module is used to count the number of non-repeated nodes in the walking sequence and splice them to obtain the substructure characteristics of the nodes;
[0117] The standardization module is used to standardize the substructure features using the L2 norm;
[0118] The activation module is used to activate the standardized substructure features using LeakyReLU;
[0119] The fusion module is used to fuse the activated substructure features with the node-level features;
[0120] The aggregation module is used to perform matrix multiplication between the adjacency matrix and the fused features to obtain a local feature aggregation, and then perform multi-hop feature aggregation on the aggregated features through GraphSage to obtain the representation of all nodes;
[0121] The constraint module is used to set constraints on the learned node representation through the graph structure to obtain the reconstruction loss;
[0122] The readout module is used to obtain graph-level representation based on node representation through graph readout function;
[0123] The scoring module is used to classify the graph-level representation, obtain anomaly factor scores, and obtain label losses based on the anomaly factor scores;
[0124] The construction module is used to jointly optimize the reconstruction loss and the label loss to construct a compound anomaly detection model;
[0125] The detection module is used to implement the anomaly detection task of the compound to be tested based on the compound anomaly detection model.
[0126] In this embodiment, the process of sampling multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths includes:
[0127] Through the Node2vec method, its key parameters p and q need to be fine-tuned according to the characteristics of the data set. By default, it is recommended to set them to 2 and 0.6 respectively. Each node in the graph is used as the starting node to perform multiple biased random walks of different lengths to obtain walk sequences of different lengths.
[0128] "Multi-scale", by specifying multiple walk lengths, each node will generate multiple walk sequences of different lengths.
[0129] In this embodiment, the process of fusing the activated substructure features with the node-level features includes:
[0130] If the node has node-level features, including node attributes and node labels, they are concatenated in the order of [node attributes, node labels (one-hot encoding), substructure features], and then mapped to the new feature space by the fully connected layer to complete feature fusion. If not, feature fusion is not performed and substructure features are used directly.
[0131] In this embodiment, the process of setting constraints on the learned node representations through the graph structure includes:
[0132]
[0133] In the formula, is the predicted adjacency matrix, and H is the low-dimensional node representation matrix.
[0134] In this embodiment, the process of classifying the graph-level representation includes:
[0135]
[0136] Where Sigmoid(·) is the activation function, W3 and W4 are both learnable weight matrices, b3 and b4 are bias terms, is the predicted anomaly score, For the picture Graph-level representation of .
[0137] The output of the detection module indicates whether a compound is abnormal or not (for example, 0 for abnormal and 1 for normal, which actually depends on the definition of abnormal labels in different data sets).
[0138] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A compound anomaly detection method enhanced by random substructure features, characterized in that: The following steps are involved: Sampling multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths; Counting the number of non-repeated nodes in the walking sequence and concatenating them to obtain substructure features of the nodes; Standardizing the substructure features using the L2 norm; Use LeakyReLU to activate the standardized substructure features; Performing feature fusion on the activated substructure features and node-level features; Perform matrix multiplication between the adjacency matrix and the fused features to obtain a local feature aggregation, and perform multi-hop feature aggregation on the aggregated features through GraphSage to obtain the representation of all nodes; The learned node representations are constrained by the graph structure to obtain the reconstruction loss; Based on the node representation, the graph level representation is obtained through the graph readout function; Classifying the graph-level representation to obtain anomaly factor scores, and obtaining a label loss based on the anomaly factor scores; The reconstruction loss and the label loss are jointly optimized to construct a compound anomaly detection model; Based on the compound anomaly detection model, an anomaly detection task of the compound to be tested is realized; Methods for sampling multi-scale substructures in a compound structure graph network and obtaining walk sequences of different lengths include: Through the Node2vec method, each node in the graph is used as the starting node to perform multiple biased random walks of different lengths to obtain walk sequences of different lengths.
2. The compound anomaly detection method enhanced by random substructure features according to claim 1, characterized in that: The method for fusing the activated substructure feature with the node-level feature includes: Node-level features include: node attributes and node labels, which are concatenated in the order of [node attributes, node labels, substructure features] and then mapped to the new feature space by the fully connected layer to complete feature fusion.
3. The compound anomaly detection method enhanced by random substructure features according to claim 1, characterized in that: Methods for setting constraints on learned node representations through graph structures include: In the formula, is the predicted adjacency matrix, and H is the low-dimensional node representation matrix.
4. The compound anomaly detection method enhanced by random substructure features according to claim 1, characterized in that: Methods for classifying the graph-level representation include: Where Sigmoid(·) is the activation function, W3 and W4 are both learnable weight matrices, b3 and b4 are bias terms, is the predicted anomaly score, For the picture Graph-level representation of .
5. A compound anomaly detection system enhanced by random substructure features, characterized in that: include: Sampling module, splicing module, normalization module, activation module, fusion module, aggregation module, constraint module, readout module, scoring module, construction module and detection module; The sampling module is used to sample multi-scale substructures in the compound structure graph network to obtain walk sequences of different lengths; The splicing module is used to count the number of non-repeated nodes in the walking sequence and splice them to obtain the substructure characteristics of the nodes; The standardization module is used to standardize the substructure features using the L2 norm; The activation module is used to activate the standardized substructure features using LeakyReLU; The fusion module is used to fuse the activated substructure features with the node-level features; The aggregation module is used to perform matrix multiplication on the adjacency matrix and the fused features to obtain a local feature aggregation, and perform multi-hop feature aggregation on the aggregated features through GraphSage to obtain all node representations; The constraint module is used to set constraints on the learned node representation through the graph structure to obtain the reconstruction loss; The readout module is used to obtain a graph-level representation based on the node representation through a graph readout function; The scoring module is used to classify the graph-level representation, obtain anomaly factor scores, and obtain label losses based on the anomaly factor scores; The construction module is used to jointly optimize the reconstruction loss and the label loss to construct a compound anomaly detection model; The detection module is used to implement anomaly detection tasks of the compounds to be tested based on the compound anomaly detection model; The process of sampling multi-scale substructures in the compound structure graph network and obtaining walk sequences of different lengths includes: Through the Node2vec method, each node in the graph is used as the starting node to perform multiple biased random walks of different lengths to obtain walk sequences of different lengths.
6. The compound anomaly detection system enhanced by random substructure features according to claim 5, characterized in that: The process of fusing the activated substructure features with the node-level features includes: Node-level features include: node attributes and node labels, which are concatenated in the order of [node attributes, node labels, substructure features] and then mapped to the new feature space by the fully connected layer to complete feature fusion.
7. The compound anomaly detection system enhanced by random substructure features according to claim 5, characterized in that: The process of setting constraints on the learned node representation through the graph structure includes: In the formula, is the predicted adjacency matrix, and H is the low-dimensional node representation matrix.
8. The compound anomaly detection system enhanced by random substructure features according to claim 5, characterized in that: The process of classifying the graph-level representation includes: Where Sigmoid(·) is the activation function, W3 and W4 are both learnable weight matrices, b3 and b4 are bias terms, is the predicted anomaly score, For the picture Graph-level representation of .
Citation Information
Patent Citations
Method for constructing complex attribute network representation model based on path aggregation
CN112395512A
Event detection method based on multi-scale heterogeneous graph embedding algorithm
CN114528479A