Abnormal node detection method and system based on network enhancement joint sparse canonical correlation analysis
By calculating kNN graphs in attribute networks and extracting features of combined graph neural networks, calculating reconstruction errors and sparse typical vectors, and comprehensively detecting abnormal nodes, the problems of insufficient feature extraction capabilities and failure to fully utilize network topology in the existing technology are solved, and more accurate detection of abnormal nodes is achieved.
Patent Information
- Application Number
- CN202411961028.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-13
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems in the detection of attribute network abnormal nodes such as insufficient feature extraction capabilities and failure to fully utilize the network topology structure, resulting in poor detection results.
Using a method based on network enhancement combined sparse typical correlation analysis, the kNN graph of each node is calculated, the network structure characteristics are enhanced, and the features are extracted in combination with the graph neural network, reconstruction error and sparse typical vector are calculated, and abnormal nodes are detected in a comprehensive anomaly score.
The detection performance of the model is improved, the information dissemination ability of discrete nodes and subgraphs is enhanced, the generalization ability is improved, and abnormal nodes can be detected more accurately.
Smart Images

Figure CN119945928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet detection technology, and in particular to an abnormal node detection method and system based on network enhanced joint sparse canonical correlation analysis. Background Art
[0002] Attribute networks refer to networks in which nodes contain attribute information. With the rapid growth of the Internet and social media, the number of nodes in the attribute networks of the real world, such as social media networks, transportation networks, financial transaction networks, and communication networks, has become larger and larger, the relationships have become more complex, and the attribute information carried by the nodes themselves has become richer. At the same time, these attribute networks in the real world are also facing a variety of potential threats and risks, such as the spread of false information on social media, abnormal phenomena of traffic network congestion, financial fraud, and telecommunications fraud in communications. Attribute network abnormal node detection technology aims to detect abnormal nodes that deviate significantly from most normal nodes, which can effectively deal with the threats faced by the attribute networks in the real world.
[0003] Traditional abnormal node detection methods use specific detection rules, clustering, machine learning, etc. based on the experience of professionals to detect abnormal nodes in attribute networks. However, traditional methods have problems such as insufficient feature extraction capabilities and failure to fully utilize the topological structure of the network, resulting in poor detection results.
[0004] The development of graph neural networks (GNNs) in recent years has provided a new approach to solving the problem of abnormal node detection in attributed networks. According to the research of Ding et al. [Ding K, Li J, Bhanushali R, et al. Deep anomaly detection on attributed networks [C] / / Proceedings of the 2019 SIAM International Conference on Data Mining 2019: 594-602.], using GNN to extract features and detect abnormal nodes can make full use of the structure and attribute information of attribute networks and improve detection performance. Duan et al. [J.Duan, B.Xiao, S.Wang, H.Zhou, and X.Liu, “ARISE: Graph anomaly detection on attributed networks via substructureawareness, IEEE Trans.Neural Netw.Learn Syst., early access, Sep.22, 2023.] proposed ARISE to detect dense patterns in the network as suspicious areas where abnormal nodes exist, and then calculated the similarity between node pairs, and detected abnormal nodes based on the average similarity. Wasim et al. [Khan W, Ishrat M, KhanAN, etal. DetectingAnomaliesinAttributedNetworks through SparseCanonicalCorrelationAnalysiscombinedwithRandomMaskingandPadding[J].IEEEAccess, 2024.] proposed PROPOSED to detect abnormal nodes based on sparse canonical analysis and random feature padding.
[0005] The ability of graph neural network models to extract features mainly relies on information propagation. However, in the real world, data quality is low, and network data obtained through data collection is usually incomplete, with discrete nodes and subgraphs. Information propagation cannot complete the information of discrete nodes and subgraphs, resulting in inaccurate learned structure and node attribute features. In addition, methods based on graph neural networks capture network information and detect abnormal nodes by constructing complex models. The focus of such methods is to capture network information through complex models rather than detecting abnormal nodes. This also leads to the risk of overfitting, and the abnormal nodes identified on new data sets are not effective enough.
[0006] In order to solve the above problems, people have been seeking an ideal technical solution. Summary of the invention
[0007] The purpose of the present invention is to address the deficiencies of the prior art and thus provide an abnormal node detection method and system based on network enhanced joint sparse canonical correlation analysis.
[0008] The method of the present invention first calculates the kNN (k-nearest neighbor) graph of each node based on the features. The characteristics of the kNN graph can ensure that each node has at least k neighbors, thereby ensuring the connectivity between nodes and enhancing the structural characteristics of the attribute network. For discrete points, this connectivity can effectively promote information propagation between nodes and greatly improve the model performance; at the same time, random features are added to enhance the generalization ability of the model. Then, the structural characteristics of the attribute network and the node attribute characteristics are extracted through the graph neural network, and the reconstruction error and sparse typical vector are calculated. The reconstruction error can ensure that the model detects most of the anomalies in the network, and the sparse typical vector can process high-dimensional sparse data and detect specific anomalies. Finally, the reconstruction error of each node and the correlation of the typical sparse vector are comprehensively considered to calculate the anomaly score and detect abnormal nodes.
[0009] Specifically, the first aspect of the present invention provides an abnormal node detection method based on network enhanced joint sparse canonical correlation analysis, comprising:
[0010] Obtain the graph network data of the attribute network G to be detected, wherein the graph network data includes the node v i , node attribute feature X, structural feature A;
[0011] Use node attribute feature X and structural feature A as training input data to train the abnormal node detection model;
[0012] The abnormal node detection model includes a network structure enhancement part, a network feature learning part, a reconstruction error and typical vector calculation part, and an abnormal node detection part;
[0013] The network structure enhancement part is configured as follows: based on the node attribute feature X, an enhanced graph structure of the node is obtained; the structural information of the attribute network is completed according to the enhanced graph structure of each node; at the same time, a Gaussian distribution is used to add random features to the node to obtain an enhanced attribute network G'=(A', X'), wherein A' is the structural feature of the enhanced attribute network, and X' is the node attribute feature of the enhanced attribute network;
[0014] The network feature learning part is configured to: input the enhanced attribute network G'=(A', X') into the graph neural network model for training, learn the structural features and node attribute features of the enhanced attribute network G', and obtain the learned structural features Hstructure and node attribute feature H attribute ;
[0015] The reconstruction error and typical vector calculation part is: according to the learned structural feature H structure and node attribute feature H attribute As well as the node attribute feature X and structural feature A of the attribute network G to be detected, the structural reconstruction error of the attribute network is calculated and attribute reconstruction error At the same time, based on the learned structural features H structure and node attribute feature H attribute , use the sparse canonical correlation analysis method to calculate the structural typical vector U and the attribute typical vector V; based on the obtained structural typical vector U and attribute typical vector V, calculate the deviation e of the typical vector of each node i ;
[0016] The abnormal node detection part is configured to reconstruct the error according to the attributes of each node Structural reconstruction error and the typical vector deviation e i Calculate the node abnormality score i ; According to the abnormal value score s of each node i Ranking detection of abnormal nodes.
[0017] Based on the above, the method to obtain the enhanced attribute network is:
[0018] First, based on the node attribute feature X, the K-nearest neighbor algorithm is used to calculate the i The k most similar nodes;
[0019] Next, compare each node v i Among the neighbor nodes in the attribute network to be detected G and the newly generated k neighbor nodes, if the number of neighbor nodes in G is greater than or equal to k, no new node is added; if it is less than k, the target node v is given i Add k new neighbor nodes; that is:
[0020]
[0021] Among them, s(X i ,X j ) represents node v i and any node v j The size of the similarity between top k Represents the target node v i and node v j The similarity between them is small, for the first k names, |v i | represents node v iThe number of neighbors;
[0022] At the same time, random noise drawn from Gaussian distribution is added to reduce overfitting of nodes.
[0023] Based on the above, the structural feature A' and the node attribute feature X' are used as the input of the graph neural network model, and the two share the weight W l ;
[0024] The graph neural network model uses a two-layer GIN encoder, and the process of capturing network features is divided into the following steps:
[0025] First, the node attribute feature X' is processed by two layers of graph neural network to obtain the final embedding representation H' attribute ,make
[0026]
[0027] in, is the normalized adjacency matrix, usually D is the degree matrix of the attribute network G, W 0 and W 1 are the weight matrices of the first and second layers respectively; the final output after the second layer It is the encoding representation of node attributes, denoted by H attribute ;
[0028] Next, we use the same GIN layer with shared weights to encode pure structural information and initialize
[0029]
[0030] in, is the normalized adjacency matrix, usually D is the degree matrix of the attribute network G, W 0 and W 1 The weight matrices of the first and second layers are respectively, and the final output after the second layer It is the encoding representation of the node structure, denoted by H structure .
[0031] Based on the above, the method for calculating the structural reconstruction error and attribute reconstruction error of the attribute network is:
[0032] A dual-channel decoder is used to perform attribute reconstruction and structure reconstruction respectively, including an attribute decoder and a structure decoder;
[0033] The attribute reconstruction process of the attribute decoder is:
[0034] A fully connected neural network is used for feature conversion to obtain the reconstructed node attribute features; namely:
[0035]
[0036] Where X 1 =H structure , represents the reconstructed node attribute features, l∈{1,2,3…} represents the number of decoder layers;
[0037] The structural reconstruction process of the structural decoder is:
[0038] The low-dimensional embedding similarity of each node pair is calculated through the inner product to reconstruct the edges of the network topology structure. The similarity of the node pairs is calculated through the low-dimensional features of the network attribute information of the attribute network to obtain the reconstructed structural features;
[0039]
[0040] Among them, Z represents the low-dimensional node attribute characteristics of the network, Z = H attribute ,use(.) T Represents the transpose of the matrix and uses the Sigmoid activation function for normalization, that is, Sigmoid(.);
[0041] Use the reconstructed node attribute features The difference between the node attribute feature X of the original attribute network structure is taken as the attribute reconstruction error R S ;
[0042] Use the reconstructed structural features The difference between the structural feature A of the original attribute network structure is taken as the structural reconstruction error R A .
[0043] Based on the above, the method for calculating the typical vector of the structure and the typical vector of the attribute is:
[0044] Use KL divergence to align node attribute features H attribute and structural features H structure The potential representation distribution of
[0045] The structural features H learned for the given two sets of variables structure , node attribute characteristics H attribute , using sparse canonical correlation analysis (SCCA); the goal of SCCA is to identify two groups of canonical variables, Where q represents the potential space dimension of node attributes and graph structure, maximizing the node attribute H attribute And the graph structure H structure correlation, and add sparsity constraints;
[0046] The optimization problem in the SCCA calculation process is described as:
[0047]
[0048] The following constraints need to be met:
[0049] |W attribute |2≤1,|W structure |2≤1#(3.10)
[0050] |W attribute |1≤c attribute ,|W structure |1≤c structure #(3.11)
[0051] Among them, the L1 norm constraint |W attribute |1≤c attribute, |W structure |1;
[0052] In the SCCA calculation process, the least squares method ALS was used for parameter optimization;
[0053] After SCCA calculation, the output is attribute sparse typical variables U and attribute sparse typical variables V;
[0054] U=H attribute W attribute
[0055] V=H structure W structure
[0056] Then, the correlation between node attributes and structures is calculated using sparse canonical correlation variables, and the deviation between node attributes and structures is used as the basis for detecting abnormal nodes.
[0057]
[0058] Among them, U i and V i is node v i The sparse canonical variable score, e i Represents node v i Bias in structure and attributes.
[0059] Based on the above, the abnormal node detection model adopts an overall training method, and the loss function is set to:
[0060]
[0061] Among them, γ, η, μ, λ are weight parameters used to balance the contribution of each part in the loss function;
[0062] is the reconstruction loss for training graph neural networks, D KL (P|Q) is used to align the attributes and structures KL divergence, |W attribute |1-(|W structure |1 is the L1 regularization of SCCA, constraining sparsity.
[0063] In a second aspect, the present application provides an abnormal node detection system based on network enhanced joint sparse canonical correlation analysis, comprising: a data acquisition module and an abnormality detection module, wherein:
[0064] The data acquisition module is used to acquire the graph network data of the attribute network G to be detected, wherein the graph network data includes the node v i , node attribute feature X, structural feature A;
[0065] Anomaly detection module, used to train an abnormal node detection model by taking node attribute feature X and structural feature A as training input data;
[0066] The abnormal node detection model includes a network structure enhancement part, a network feature learning part, a reconstruction error and typical vector calculation part, and an abnormal node detection part;
[0067] The network structure enhancement part is configured as follows: based on the node attribute feature X, an enhanced graph structure of the node is obtained; the structural information of the attribute network is completed according to the enhanced graph structure of each node; at the same time, a Gaussian distribution is used to add random features to the node to obtain an enhanced attribute network G'=(A', X'), wherein A' is the structural feature of the enhanced attribute network, and X' is the node attribute feature of the enhanced attribute network;
[0068] The network feature learning part is configured to: input the enhanced attribute network G'=(A', X') into the graph neural network model for training, learn the structural features and node attribute features of the enhanced attribute network G', and obtain the learned structural features H structure and node attribute feature H attribute ;
[0069] The reconstruction error and typical vector calculation part is: according to the learned structural feature H structure and node attribute feature H attribute As well as the node attribute feature X and structural feature A of the attribute network G to be detected, the structural reconstruction error of the attribute network is calculated and attribute reconstruction error At the same time, based on the learned structural features H structure and node attribute feature H attribute, use the sparse canonical correlation analysis method to calculate the structural typical vector U and the attribute typical vector V; based on the obtained structural typical vector U and attribute typical vector V, calculate the deviation e of the typical vector of each node i ;
[0070] The abnormal node detection part is configured to reconstruct the error according to the attributes of each node Structural reconstruction error and the typical vector deviation e i Calculate the node abnormality score i ; According to the abnormal value score s of each node i Ranking detection of abnormal nodes.
[0071] In a third aspect, the present application provides an electronic device, including:
[0072] at least one processor, and a memory coupled to the at least one processor;
[0073] The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis as described.
[0074] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed, the abnormal node detection method based on network-enhanced joint sparse canonical correlation analysis as described above can be implemented.
[0075] In a fifth aspect, the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis as described.
[0076] The present invention has the following beneficial effects and advantages:
[0077] 1) The present invention proposes a KNN graph-based attribute network enhancement method. The KNN graph is calculated based on the features of each node, ensuring that each node has at least k neighbor nodes, completing the discrete nodes and discrete subgraphs in the attribute network, enhancing the structural characteristics of the attribute network, facilitating information propagation and feature extraction of the graph neural network, and improving the detection performance of the model.
[0078] 2) This paper proposes a method for calculating anomaly scores by combining reconstruction error with sparse typical vector correlation. By detecting abnormal nodes based on reconstruction error and calculating the correlation between node attribute features and structural features, abnormal nodes with high similarity to normal nodes are detected, which makes up for the shortcomings of the reconstruction error method. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 This is the abnormal node detection model diagram proposed in this application.
[0080] Figure 2 It is the influence diagram of the GNN encoder result.
[0081] Figure 3 This is a comparison chart of network enhancement experimental results.
[0082] Figure 4 A schematic diagram of an electronic device provided in Example 3 of the present application.
[0083] Figure 5 A schematic diagram of a computer-readable storage medium provided for Example 3 of the present application. DETAILED DESCRIPTION
[0084] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0085] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may include steps or units that are not listed.
[0086] To facilitate understanding of the technical solutions provided by the present application, the technical terms involved in the embodiments of the present application are explained below.
[0087] Attribute network, generally defined as G = (V, A, X), where V = {v1, v2, v3 ... v n}, (|V|=n) represents the nodes of the attribute network; is the structural feature of the attribute network and also the adjacency matrix of the attribute network. ij = 0, indicating that node v i and node v j There is no edge connection between them. ij =1 means there is an edge connection between the two; is the node attribute feature of the attribute network, which is also the attribute matrix of the attribute network. The vector of the i-th row represents the i-th node v of the attribute network. i Attribute information of the .
[0088] SCCA, sparse canonical correlation analysis, is a method for studying the correlation between two high-dimensional data sets. The core idea of SCCA is to project the two data sets into a common low-dimensional space so that their correlation is maximized.
[0089] The following is a detailed description of the methods, systems, devices, media and products provided in the present application through some embodiments and their application scenarios in conjunction with the accompanying drawings.
[0090] Example 1
[0091] like Figure 1 As shown, this embodiment provides an abnormal node detection method based on network enhanced joint sparse canonical correlation analysis, including:
[0092] Obtain the graph network data of the attribute network G to be detected, wherein the graph network data includes the node v i , node attribute feature X, structural feature A;
[0093] The node attribute feature X and the structural feature A are used as training input data to train the abnormal node detection model (NEAnomSCCA);
[0094] The abnormal node detection model includes a network structure enhancement part, a network feature learning part, a reconstruction error and typical vector calculation part, and an abnormal node detection part;
[0095] The network structure enhancement part is configured as follows: based on the node attribute feature X, an enhanced graph structure of the node is obtained; the structural information of the attribute network is completed according to the enhanced graph structure of each node; at the same time, a Gaussian distribution is used to add random features to the node to obtain an enhanced attribute network G'=(A', X'), wherein A' is the structural feature of the enhanced attribute network, and X' is the node attribute feature of the enhanced attribute network;
[0096] The network feature learning part is configured to: input the enhanced attribute network G'=(A', X') into the graph neural network model for training, learn the structural features and node attribute features of the enhanced attribute network G', and obtain the learned structural features H structure and node attribute feature H attribute ;
[0097] The reconstruction error and typical vector calculation part is: according to the learned structural feature H structure and node attribute feature H attribute As well as the node attribute feature X and structural feature A of the attribute network G to be detected, the structural reconstruction error of the attribute network is calculated and attribute reconstruction error At the same time, based on the learned structural features H structure and node attribute feature Hattribute , use the sparse canonical correlation analysis method to calculate the structural typical vector U and the attribute typical vector V; based on the obtained structural typical vector U and attribute typical vector V, calculate the deviation e of the typical vector of each node i ;
[0098] The abnormal node detection part is configured to reconstruct the error according to the attributes of each node Structural reconstruction error and the typical vector deviation e i Calculate the node abnormality score i ; According to the abnormal value score s of each node i Ranking detection of abnormal nodes.
[0099] Network Enhancement
[0100] Discrete nodes hinder the information propagation of attribute networks. In order to solve this problem, a relatively effective approach is to introduce a new graph structure to ensure that all nodes are interconnected. However, directly replacing the original graph structure with a new graph structure will inevitably cause information loss. Therefore, the present invention proposes a new enhanced graph structure generated by the K-nearest neighbor algorithm based on the original attribute network structure, referred to as knn structure enhancement, and the enhanced graph is called a knn graph. The abnormal node detection model uses the knn graph for information propagation during training. The specific process is as follows:
[0101] First, based on the node attribute feature X, the K-nearest neighbor algorithm is used to calculate the i The k most similar nodes;
[0102] Next, compare each node v i Among the neighbor nodes in the attribute network to be detected G and the newly generated k neighbor nodes, if the number of neighbor nodes in G is greater than or equal to k, no new node is added; if the number is less than k, the target node v is given i Add k new neighbor nodes; that is:
[0103]
[0104] Among them, s(X i ,X j ) represents node v i and any node v j The size of the similarity between top k Represents the target node v i and node v j The similarity between them is small, for the first k, |v i | represents node v i The number of neighbors of the target node vi and node v j The similarity size is the top k And v i The number of nodes is less than k, so let A' ij Add a variable between them, that is, A' ij =1;
[0105] At the same time, random noise is added to reduce overfitting of nodes. Random noise is extracted from Gaussian distribution to ensure that the model does not overly rely on any single feature, thereby improving the generalization ability of the model.
[0106] Feature Learning
[0107] Graph neural network is a deep learning model specially designed for processing graph structure data. It can extract features from graphs and perform well in a variety of tasks. This application uses graph neural network, namely GNN network, to learn the features of attribute network.
[0108] The input of the graph neural network is: node features and graph structure. In this application, the enhanced structural features A' and node attribute features X are used as the input of the graph neural network, and the two share the parameter W l , shared weights enable the model to learn a unified set of parameters that can effectively encode the attribute characteristics and structural features of the network.
[0109] This application uses a two-layer GIN encoder because the two-layer network can not only capture the direct neighbors of each node (first order), but also capture the neighbors of the neighbors (second order), capturing richer and more complex network features. This process can be divided into the following steps:
[0110] First, the attribute matrix X of the attribute network is processed by two layers of GIN network to obtain the final embedding representation H' attribute ,make
[0111]
[0112] in, is the normalized adjacency matrix, usually D is the degree matrix of the attribute network G, W 0 and W 1 The weight matrices of the first and second layers are respectively. The final output after the third layer It is the encoding representation of node attribute features, denoted by H attribute .
[0113] Next, to encode pure structural information, this application uses the same GIN layer with shared weights, but this time the input will be different, mainly learning structural features, initialization
[0114]
[0115] in, is the normalized adjacency matrix, usually D is the degree matrix of the attribute network G, W 0 and W 1 The weight matrices of the first and second layers are respectively. The final output after the third layer It is the encoding representation of the structural features, denoted by H structure .
[0116] Reconstruction error calculation
[0117] In order to reconstruct the attribute network structure and node attributes from the low-dimensional features of the structural information and attribute information, this application uses a dual-channel decoder. The reconstruction process includes the following steps:
[0118] First, the attribute decoder uses a fully connected neural network to perform feature conversion and reconstruct node attributes. The specific formula is shown in formula (6);
[0119]
[0120] Where X 1 =H structure , represents the reconstructed attribute features, and l∈{1,2,3…} represents the number of decoder layers.
[0121] Next, the low-dimensional embedding similarity of each node pair is calculated through the inner product to reconstruct the edges of the attribute network topology. The similarity of node pairs is calculated using the low-dimensional features of the network attribute information of the attribute network. See formula (7) for details;
[0122]
[0123] Among them, Z represents the low-dimensional attribute features of the attribute network, Z = H attribute . Use (.)T to represent the transpose of the matrix, and use the Sigmoid activation function for normalization, that is, Sigmoid(.).
[0124] Then use the reconstructed attribute features The attribute reconstruction error is recorded as R S Similarly, the structural reconstruction error is calculated and recorded as R A .
[0125] Typical eigenvector calculation
[0126] Step 1: KL divergence alignment;
[0127] This application uses KL divergence to align the node attribute features H attribute and structural features H structure The main purpose is to minimize the latent space representation of node attribute features and structural features and optimize the alignment distribution of anomaly detection;
[0128] L KL =D KL (P(H structure )|P(H attribute )) (8)
[0129] P(H attribute ) and P(H structure ) are the probability distributions of the latent space representations of node attribute features and structural features, respectively.
[0130] Since both node attribute features and structural features use the same shared parameters for feature compression, the low-dimensional representations of node attribute attributes and structures are in the same feature space, and the comparability between the two must be ensured in the same feature space. The next stage of the process, SCCA, requires the results of KL divergence alignment.
[0131] Step 2: Sparse typical vector calculation;
[0132] The main basis for abnormal node detection is the node attribute feature H attribute and structural features H structure SCCA can obtain the typical vectors of the two, and then calculate the correlation based on the typical vectors. Given two sets of variables H attribute and H structure , both are in the same feature space with dimension n×q, where n is the number of nodes and q is the spatial dimension of the potential node attributes and structure. The goal of SCCA is to identify two sets of typical variables, Where q represents the potential space dimension of node attributes and structures, maximizing the node attribute feature H attribute and structural features H structure The correlation is improved while adding sparsity constraints.
[0133] The optimization problem in the SCCA calculation process is described as:
[0134]
[0135] The following constraints need to be met:
[0136] |W attribute |2≤1,|W structure |2≤1 (10)
[0137] |W attribute |1≤c attribute,|W structure 11≤c structure (11)
[0138] Among them, the L1 norm constraint |W attribute |1≤c attribute, |W structure |1. Constraining sparsity in typical vectors ensures that each vector only utilizes a limited number of important features in its respective feature set. The parameter c attribute and c structure Constrained sparsity. Due to the sparsity constraint, the optimization problem in SCCA can be challenging. This application uses the least squares method (ALS) for parameter optimization, which iteratively optimizes one variable while keeping other variables fixed. The output of SCCA includes attribute sparse canonical variables U and attribute sparse canonical variables V, see formulas (12) and (13).
[0139] U=H attribute W attribute (12)
[0140] V=H structure W structure (13)
[0141] Then, the sparse canonical correlation variables are used to calculate the correlation between node attributes and structures, and the deviation between node attributes and structures is used as the basis for detecting abnormal nodes. The specific calculation is shown in formula (14).
[0142]
[0143] Among them, U i and V i is node v i The sparse canonical variable score of U is: i Defined as the element value of the i-th row of the column vector U, let: Then the sparse canonical variable score U1 of v1 is a1, v i The sparse canonical variate score U i for a i ; V i Similarly i ;
[0144] If the correlation between the structure and node attributes is small, that is, the deviation between the two is too large, the possibility of abnormality in the node is greater. i Represents node v i Deviations in structure and properties, e i The larger it is, the greater the deviation.
[0145] Abnormal node detection
[0146] The node outlier score is mainly calculated by the deviation of reconstruction error and sparse typical vector correlation, see formula (15) for details.
[0147]
[0148] in, Represents node v i The structural reconstruction error is v represents the node i The attribute reconstruction error, e i For node v i α and β are used to adjust the deviation of the correlation of sparse typical vectors and the ratio of reconstruction error, α, β∈(0,1). If α is too large, the generalization of the model is stronger, and if β is too large, the model pays more attention to the detection of the current data set. Finally, the abnormal value score S of each node is determined based on the two. i , according to S i The size of the attribute can be used to detect abnormal nodes in the attribute network.
[0149] Loss Function
[0150] The loss function is the key to model training optimization and is an important part of the optimization model. The abnormal node detection model of this application adopts overall training, see formula (16).
[0151]
[0152] Among them, γ, η, μ, and λ are weight parameters used to balance the contribution of each part in the loss function. is the reconstruction loss for training graph neural networks, D KL (P|Q) is used to align the attributes and structures KL divergence, |W attribute |1-(|W structure |1, is the L1 regularization of SCCA, constraining sparsity.
[0153] Comparative experiment
[0154] 1. Dataset Introduction
[0155] This experiment was conducted on four widely used attribute network anomaly detection benchmark datasets, namely BlogCatalog, Cora, Citeseer, and Pubmed.
[0156] BlogCatalog: A blog sharing website, where nodes represent website users and links represent the follow-up relationship between users. The attributes of nodes are composed of the user's personalized content, such as the tags of the user's blog posts, shared pictures, and customized tags.
[0157] Cora: This dataset consists of 2708 papers, represented by Node. There are 5429 citation links, represented by Edges. The attributes of each paper indicate whether the paper has a certain keyword, and the dimension of the attribute is 1433.
[0158] Citeseer: This dataset consists of 3327 scientific papers, represented by Nodes. There are 4732 citation links, represented by Edges. The dimension of the attributes of each scientific paper is 3703.
[0159] Pubmed: This dataset consists of 19,717 publications with 44,338 citation links between nodes. The dimension of the attributes of each publication is 500.
[0160] 2. Experimental Dataset Construction
[0161] The node attributes of the above four datasets are vectorized using the bag-of-words model. The attributes of each node are represented by a vector, and the size of the vector dimension is determined by the dictionary size of the bag-of-words model.
[0162] Since the above four data sets do not have annotated abnormal nodes and there is currently no unified abnormal node annotation standard, this experiment adopts the currently widely used structural perturbation and attribute perturbation anomaly injection methods.
[0163] Inject structurally abnormal nodes by disturbing the network topology. The principle behind this method is that in many real-world scenarios, very few nodes in a subgraph are fully connected, so fully connected structures are considered abnormal. First, randomly select m subgraphs in the network, each with n nodes; then, fully connect the n nodes in the m subgraphs to form a fully connected structure. The fully connected nodes in these subgraphs are considered structurally abnormal nodes.
[0164] Generate attribute abnormal nodes by perturbing node attributes. Use the node with the largest Euclidean distance to the target node in the network to perturb the attribute of the target node. First, randomly select m×n nodes from the network as candidate attribute abnormal nodes; then, for each target node in the candidate nodes, randomly select k nodes from the data except the candidate nodes, calculate the Euclidean distance between the target node and k nodes, and find the node j with the largest Euclidean distance to the target node; then replace the attribute of the target node with the attribute of node j; finally, generate m×n nodes as attribute abnormal nodes.
[0165] According to the above injection method and according to the size of the network, each data set is injected with abnormal nodes of about 5% of the network size. Finally, the perturbed network is obtained, and the total number of anomalies is listed in the last row of Table 4.1. In this experiment, all category labels are deleted, and only abnormal labels can be seen in the inference stage.
[0166] Experimental indicators
[0167] In this experiment, AUC value (Area Under the Curve) is used as the evaluation indicator to evaluate the performance of the model.
[0168] AUC is a common metric used to evaluate the performance of a binary classification model. It represents the area under the ROC curve. Its physical meaning is: given random positive and negative examples, the probability that the model predicts a positive example as a positive example is higher than the probability that the model predicts a negative example as a positive example. Therefore, the larger the AUC value, the better the model performance, that is, the more accurate the model is in distinguishing positive and negative examples.
[0169] This experiment uses Windows 10, Python 3.8, PyTorch 1.1 as the operating environment, and an NVIDIA GeForce graphics card with 16GB of video memory.
[0170] In order to ensure the fairness of the experimental results, this abnormal node detection model and the comparative experiment use the same environment and settings: the optimizer uses Adam, the learning rate is 0.01, the final node embedding dimension is set to 64, the number of layers of the graph neural network is set to 2, and the dropout value is 0.3. In addition to the above common parameters, the other parameters of the comparative experiment remain the same as the original settings.
[0171] GNN encoder selection
[0172] This section analyzes the impact of different GNN encoders on the NEAnomSCCA model, mainly exploring the impact of common graph neural network encoders GCN, GAT, and GIN on model performance. Experimental verification was conducted on four datasets, and the experimental results are as follows: Figure 2 As shown:
[0173] According to the experimental results, GIN is the best encoder for the entire model, so GIN is selected as the final encoder for the model. In the anomaly detection task, GIN more accurately distinguishes normal nodes from abnormal nodes. Although GAT is not as good as GIN in the abnormal node detection task, it is still improved compared to GCN. This is because GAT considers the weight contribution of different neighbors when aggregating neighbor information. In particular, for the data imbalance problem in the abnormal node detection task, GAT slows down the destruction of the homogeneity of the network by abnormal nodes in the process of aggregating neighbor information. GCN performed generally in this task. The main reason is that there are inconsistently connected edges in the data set, that is, the normal nodes and abnormal nodes are inconsistently connected. GCN introduces noise data in the aggregation process, resulting in the interference of the trained graph representation by the noise data.
[0174] Model performance comparison experiment
[0175] The experimental results of the NEAnomSCCA model are compared with the other six baseline models. The relevant comparison models are introduced below.
[0176] LOF: Attribute-based anomaly detection method. It determines whether a node is an abnormal node by observing the attribute values of the same community. It only considers the attribute information of the node and ignores the structural anomaly of the attribute network.
[0177] DOMINANT: Anomaly detection architecture based on graph neural network. It uses graph convolution and autoencoder to jointly reconstruct the adjacency matrix and attribute matrix. The degree of abnormality of each node is evaluated by calculating the reconstruction error. This method has the problem of over-smoothing. When GCN aggregates neighbor information, it will smooth the abnormal information.
[0178] AnomalyDAE: GAT is applied to the embedding of attribute networks to learn the importance between different neighbor nodes, thereby improving the effect of anomaly detection.
[0179] JAANE: The target node and its neighbor information are considered simultaneously during the aggregation process, and the feature vector of the node is learned by fusing the two. The hypersphere of normal nodes is learned by reconstructing the error and regularization modules, and nodes outside the hypersphere are regarded as abnormal nodes.
[0180] ARISE: This method mainly detects anomalies by identifying specific patterns in the network. It uses a region proposal module to detect dense patterns in the network as suspicious areas where abnormal nodes exist. It then calculates the similarity between node pairs and detects abnormal nodes based on the average similarity. In addition, it uses graph contrast learning to detect node attribute anomalies.
[0181] PROPOSED: An abnormal node detection model based on sparse typical analysis and random feature filling first adds random features to each node to enhance randomness and prevent overfitting. Then it extracts structure and attribute features and aligns them to the same latent space through KL divergence. Then it calculates the sparse typical vectors of attributes and structures through SCCA, calculates the anomaly score based on the typical vector, and detects abnormal nodes.
[0182] Table 1 Comparison of experimental results
[0183]
[0184] From Table 1, we can conclude that the NEAnomSCCA model has achieved good results on all four datasets. Specifically, NEAnomSCCA has achieved certain improvements on BlogCatalog, Cora, Flickr, and Pubmed, respectively, with the best AUC values of the comparison model increasing by 3.46%, 1.04%, and 1.69%, respectively. This shows that NEAnomSCCA enhances the robustness of the anomaly score through SCCA sparsity analysis, making up for the deficiency that the reconstruction error cannot detect edge abnormal nodes, and at the same time, the network structure enhancement enhances the model's information propagation ability. The traditional abnormal node detection method is obviously inferior to the method based on graph neural network, which shows that the traditional mechanism cannot capture the attribute and structural information of the network at the same time, and the ability to process the network is limited.
[0185] Ablation analysis
[0186] In order to gain a deeper understanding of the role of each module, an ablation experiment was conducted on the model. The main focus was on exploring how the following modules affect the performance of the NEAnomSCCA model: (1) the impact of network structure enhancement; (2) the impact of SCCA sparsity analysis.
[0187] The impact of network structure enhancement
[0188] In this study, the impact of network structure enhancement on the performance of the NEAnomSCCA model is analyzed. This section compares the model without the network structure enhancement module with the NEAnomSCCA model. The model without the network structure enhancement module is denoted as AnomSCCA. The experimental results of the original model and the model without the network structure enhancement module on four data sets are shown in the figure. Figure 3 shown.
[0189] It can be clearly seen from the data that NEAnomSCCA has the best AUC performance on all four data sets. Further analysis shows that before NEAnomSCCA performs information propagation, it first completes the structure of discrete nodes in the network through network structure enhancement. In this way, when the target network learns node features through information propagation of the graph neural network model, it can enhance the feature learning of discrete nodes, promote network propagation, better learn network features, and enhance detection effects. However, NEAnomSCCA removes the network structure enhancement module. When learning node features during information propagation, discrete nodes hinder the information propagation of the network, resulting in inaccurate learned node features and affecting the accuracy of abnormal node detection.
[0190] Impact of SCCA Sparsity Analysis
[0191] In order to verify that the SCCA sparsity analysis of the proposed NEAnomSCCA model is helpful for detecting abnormal nodes in attribute networks, an ablation experiment is designed for analysis and verification. The model without SCCA sparsity analysis is compared with the NEAnomSCCA model. The new model without SCCA sparsity analysis is denoted as NEAnom. The experimental results of the two models on four datasets are shown in Table 2.
[0192] From the experimental results, it can be seen that after adding the SCCA sparsity analysis module, the experimental results of the NEAnomSCCA model on the four data sets have been improved to a certain extent, and the AUC values have increased by 3.46%, 8.92%, and 4.47% respectively. The experiment fully proves the necessity of SCCA sparsity analysis in the entire model; adding the SCCA sparsity analysis module can reduce the possibility of model overfitting, and when detecting abnormal nodes, it can calculate the anomaly score based on the correlation between node attributes and structure; if only the reconstruction error detection is used to detect abnormal nodes, the edge abnormal nodes cannot be identified, and the SCCA sparsity analysis makes up for this defect well.
[0193] Table 2 SCCA ablation results
[0194]
[0195] Example 2
[0196] This embodiment provides an abnormal node detection system based on network enhanced joint sparse canonical correlation analysis, including: a data acquisition module and an abnormality detection module, wherein:
[0197] The data acquisition module is used to acquire the graph network data of the attribute network G to be detected, wherein the graph network data includes the node v i , node attribute feature X, structural feature A;
[0198] Anomaly detection module, used to train an abnormal node detection model by taking node attribute feature X and structural feature A as training input data;
[0199] The abnormal node detection model includes a network structure enhancement part, a network feature learning part, a reconstruction error and typical vector calculation part, and an abnormal node detection part;
[0200] The network structure enhancement part is configured as follows: based on the node attribute feature X, an enhanced graph structure of the node is obtained; the structural information of the attribute network is completed according to the enhanced graph structure of each node; at the same time, a Gaussian distribution is used to add random features to the node to obtain an enhanced attribute network G'=(A', X'), wherein A' is the structural feature of the enhanced attribute network, and X' is the node attribute feature of the enhanced attribute network;
[0201] The network feature learning part is configured to: input the enhanced attribute network G'=(A', X') into the graph neural network model for training, learn the structural features and node attribute features of the enhanced attribute network G', and obtain the learned structural features H structure and node attribute feature H attribute ;
[0202] The reconstruction error and typical vector calculation part is: according to the learned structural feature H structure and node attribute feature H attribute As well as the node attribute feature X and structural feature A of the attribute network G to be detected, the structural reconstruction error of the attribute network is calculated and attribute reconstruction error At the same time, based on the learned structural features H structure and node attribute feature H attribute , use the sparse canonical correlation analysis method to calculate the structural typical vector U and the attribute typical vector V; based on the obtained structural typical vector U and attribute typical vector V, calculate the deviation e of the typical vector of each node i ;
[0203] The abnormal node detection part is configured to reconstruct the error according to the attributes of each node Structural reconstruction error and the typical vector deviation e i Calculate the node abnormality score i ; According to the abnormal value score s of each node i Ranking detection of abnormal nodes.
[0204] It should be noted that the system embodiment of this embodiment is similar to the method embodiment of Embodiment 1, so the description is relatively simple, and the relevant parts can refer to the method of Embodiment 1.
[0205] Example 3
[0206] The present application also provides an electronic device, referring to Figure 4 , Figure 4 Schematic diagram of an electronic device proposed in an embodiment of the present application. Figure 4 As shown, the electronic device includes: a memory and a processor, the memory and the processor are connected via a bus communication, the memory stores a computer program, and the computer program can be run on the processor, thereby implementing the steps in the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis disclosed in Example 1 of the present application.
[0207] The present application also provides a computer-readable storage medium. Figure 5 , Figure 5 Schematic diagram of a computer-readable storage medium proposed in an embodiment of the present application. Figure 5 As shown, a computer program / instruction is stored on a computer-readable storage medium, and when the computer program / instruction is executed by a processor, the steps in the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis disclosed in Example 1 of the present application are implemented.
[0208] The embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps in the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis as disclosed in Example 1 of the present application.
[0209] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0210] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0211] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, systems, devices, storage media, and program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0212] The methods, systems, devices, media and products provided by the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the methods and core ideas of the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An abnormal node detection method based on network enhanced joint sparse canonical correlation analysis, characterized in that: include: Obtain the graph network data of the attribute network G to be detected, wherein the graph network data includes the node v i , node attribute feature X, structural feature A; Use node attribute feature X and structural feature A as training input data to train the abnormal node detection model; The abnormal node detection model includes a network structure enhancement part, a network feature learning part, a reconstruction error and typical vector calculation part, and an abnormal node detection part; The network structure enhancement part is configured as follows: based on the node attribute feature X, an enhanced graph structure of the node is obtained; the structural information of the attribute network is completed according to the enhanced graph structure of each node; at the same time, a Gaussian distribution is used to add random features to the node to obtain an enhanced attribute network G ’ =(A ’ , X ’ ), where A ’ To enhance the structural characteristics of the attribute network, X ’ To enhance the node attribute characteristics of the attribute network; The network feature learning part is configured to: enhance the attribute network G ’ =(A ’ , X ’ ) is input into the graph neural network model for training and learning to enhance the attribute network G ’ The structural features and node attribute features of the learned structural features H structure and node attribute feature H attribute ; The reconstruction error and typical vector calculation part is: according to the learned structural feature H structure and node attribute feature H attribute As well as the node attribute feature X and structural feature A of the attribute network G to be detected, the structural reconstruction error of the attribute network is calculated and attribute reconstruction error At the same time, based on the learned structural features H structure and node attribute feature H attribute , use the sparse canonical correlation analysis method to calculate the structural typical vector U and the attribute typical vector V; based on the obtained structural typical vector U and attribute typical vector V, calculate the deviation e of the typical vector of each node i ; The abnormal node detection part is configured to reconstruct the error according to the attributes of each node Structural reconstruction error and the typical vector deviation e i Calculate the node abnormality score i ; According to the abnormal value score s of each node i Ranking detection of abnormal nodes.
2. The abnormal node detection method based on network enhanced joint sparse canonical correlation analysis according to claim 1 is characterized in that: The method to obtain the enhanced attribute network is: First, based on the node attribute feature X, the K-nearest neighbor algorithm is used to calculate the i The k most similar nodes; Next, compare each node v i Among the neighbor nodes in the attribute network to be detected G and the newly generated k neighbor nodes, if the number of neighbor nodes in G is greater than or equal to k, no new node is added; if it is less than k, the target node v is given i Add k new neighbor nodes; that is: Among them, s(X i ,X j ) represents node v i and any node v j The size of the similarity between top k Represents the target node v i and node v j The similarity between them is small, for the first k, |v i | represents node v i The number of neighbors; At the same time, random noise drawn from Gaussian distribution is added to reduce overfitting of nodes.
3. The abnormal node detection method based on network enhanced joint sparse canonical correlation analysis according to claim 1 is characterized by: The structural feature A' and the node attribute feature X ’ As the input of the graph neural network model, the two share the weight W l ; The graph neural network model uses a two-layer GIN encoder, and the process of capturing network features is divided into the following steps: First, the node attribute feature X ’ After being processed by two layers of graph neural network, the final embedding representation H' is obtained attribute ,make in, is the normalized adjacency matrix, usually D is the degree matrix of the attribute network G, W 0 and W 1 are the weight matrices of the first and second layers respectively; the final output after the second layer It is the encoding representation of node attributes, denoted by H attribute ; Next, we use the same GIN layer with shared weights to encode pure structural information and initialize in, is the normalized adjacency matrix, usually D is the degree matrix of the attribute network G, W 0 and W 1 The weight matrices of the first and second layers are respectively, and the final output after the second layer It is the encoding representation of the node structure, denoted by H struchure .
4. The abnormal node detection method based on network enhanced joint sparse canonical correlation analysis according to claim 1 is characterized in that: The method for calculating the structural reconstruction error and attribute reconstruction error of the attribute network is: A dual-channel decoder is used to perform attribute reconstruction and structure reconstruction respectively, including an attribute decoder and a structure decoder; The attribute reconstruction process of the attribute decoder is: A fully connected neural network is used for feature conversion to obtain the reconstructed node attribute features; namely: Where X 1 =H structure , represents the reconstructed node attribute features, l∈{1,2,3…} represents the number of decoder layers; The structural reconstruction process of the structural decoder is: The low-dimensional embedding similarity of each node pair is calculated through the inner product to reconstruct the edges of the network topology structure. The similarity of the node pairs is calculated through the low-dimensional features of the network attribute information of the attribute network to obtain the reconstructed structural features; Among them, Z represents the low-dimensional node attribute characteristics of the network, Z = H attribute ,use(.) T Represents the transpose of the matrix and uses the Sigmoid activation function for normalization, that is, Sigmoid(.); Use the reconstructed node attribute features The difference between the node attribute feature X of the original attribute network structure is taken as the attribute reconstruction error R S ; Use the reconstructed structural features The difference between the structural feature A of the original attribute network structure is taken as the structural reconstruction error R A .
5. The abnormal node detection method based on network enhanced joint sparse canonical correlation analysis according to claim 1 is characterized in that: The method for calculating the typical vector of the structure and the typical vector of the attribute is: Use KL divergence to align node attribute features H attribute and structural features H structure The potential representation distribution of The structural features H learned for the given two sets of variables structure , node attribute characteristics H attribute , using sparse canonical correlation analysis (SCCA); the goal of SCCA is to identify two groups of canonical variables, Where q represents the potential space dimension of node attributes and graph structure, maximizing the node attribute H attribute And the graph structure H structure The correlation of , and add sparsity constraints; The optimization problem in the SCCA calculation process is described as: The following constraints need to be met: |In attribute |2≤1,|In structure |2≤1#(3.10) |In attribute |1≤c attribute ,|In structure |1≤c structure #(3.11) Among them, the L1 norm constraint |W attribute |1≤c attribute ,|W structure |1; In the SCCA calculation process, the least squares method ALS was used for parameter optimization; After SCCA calculation, the output is attribute sparse typical variables U and attribute sparse typical variables V; U=H attribute W attribute V=H stucture W structure Then, the correlation between node attributes and structures is calculated using sparse canonical correlation variables, and the deviation between node attributes and structures is used as the basis for detecting abnormal nodes. Among them, U i and V i is node v i The sparse canonical variable score, e i Represents node v i Bias in structure and attributes.
6. The abnormal node detection method based on network enhanced joint sparse canonical correlation analysis according to claim 1 is characterized in that: The abnormal node detection model adopts an overall training method, and the loss function is set as: Among them, γ, η, μ, λ are weight parameters used to balance the contribution of each part in the loss function; is the reconstruction loss for training graph neural networks, D KL (P|Q) is used to align the attributes and structures KL divergence, |W attribute |1-(|W structure |1 is the L1 regularization of SCCA, constraining sparsity.
7. An abnormal node detection system based on network enhanced joint sparse canonical correlation analysis, characterized in that: Contains: data acquisition module and anomaly detection module, among which, The data acquisition module is used to acquire the graph network data of the attribute network G to be detected, wherein the graph network data includes the node v i , node attribute feature X, structural feature A; Anomaly detection module, used to train an abnormal node detection model by taking node attribute feature X and structural feature A as training input data; The abnormal node detection model includes a network structure enhancement part, a network feature learning part, a reconstruction error and typical vector calculation part, and an abnormal node detection part; The network structure enhancement part is configured as follows: based on the node attribute feature X, an enhanced graph structure of the node is obtained; the structural information of the attribute network is completed according to the enhanced graph structure of each node; at the same time, a Gaussian distribution is used to add random features to the node to obtain an enhanced attribute network G ’ =(A ’ , X ’ ), where A ’ To enhance the structural characteristics of the attribute network, X ’ To enhance the node attribute characteristics of the attribute network; The network feature learning part is configured to: enhance the attribute network G ’ =(A ’ , X ’ ) is input into the graph neural network model for training and learning to enhance the attribute network G ’ The structural features and node attribute features of the learned structural features H structure and node attribute feature H attribute ; The reconstruction error and typical vector calculation part is: according to the learned structural feature H structure and node attribute feature H attribute As well as the node attribute feature X and structural feature A of the attribute network G to be detected, the structural reconstruction error of the attribute network is calculated and attribute reconstruction error At the same time, based on the learned structural features H structure and node attribute feature H attribute , use the sparse canonical correlation analysis method to calculate the structural typical vector U and the attribute typical vector V; based on the obtained structural typical vector U and attribute typical vector V, calculate the deviation e of the typical vector of each node i ; The abnormal node detection part is configured to reconstruct the error according to the attributes of each node Structural reconstruction error and the typical vector deviation e i Calculate the node abnormality score i ; According to the abnormal value score s of each node i Ranking detection of abnormal nodes.
8. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, it can implement the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the abnormal node detection method based on network enhanced joint sparse canonical correlation analysis as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Abnormal data identification method and device, electronic equipment and storage medium
CN117034159A
Abnormality detection method and device, electronic equipment and storage medium
CN117725548A
Network node full-granularity anomaly detection method and system based on attribute enhanced sampling
CN118074958A
Network embedding method based on attention mechanism
CN118094252A
Attribute network abnormal node detection method and system based on iterative filtering
CN118573418A