Literature reference network anomaly detection method based on graph enhanced multi-scale contrast learning

Through graph-enhanced multi-scale comparison learning framework, the denoising diffusion model and meta-path generation enhancement graphs are solved, and the problem of inconsistency in the node-level and sub-graph-level representations in the literature cited network is improved, and the robustness and effectiveness of anomaly detection are improved.

CN120372497APending Publication Date: 2025-07-25HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419153.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing literature cites network anomaly detection method when calculating node-level representations, ignores the structure and context information of the target node, resulting in insufficient representation, and inconsistent node-level and sub-graph-level representations of heterogeneous nodes, affecting the robustness and detection effect of the model.

Method used

The graph-enhanced multi-scale comparison learning framework is adopted to inject the features of the target node into some neighbors through the denoising diffusion probability model, and the enhanced graph is generated, and the metapath and multi-scale comparison module are used to improve the consistency of node-level and sub-graph-level representations, and a multi-scale comparison loss is constructed to calculate the abnormal score.

Benefits of technology

It effectively alleviates the inconsistency problem of node-level and sub-graph-level representations, improves the robustness and abnormal detection performance of the model, and significantly improves the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372497A_ABST
    Figure CN120372497A_ABST
Patent Text Reader

Abstract

The invention discloses a literature citation network anomaly detection method based on graph enhanced multi-scale contrast learning. The method comprises the following steps: constructing an attribute network of a literature citation network; constructing a graph enhancement multi-scale contrast learning framework used for document reference network anomaly detection, wherein the graph enhancement multi-scale contrast learning framework comprises a graph enhancement module and a multi-scale contrast module; the graph enhancement module is used for injecting features of a target node in an attribute network into part of neighbors of the target node based on a denoising diffusion probability model so as to generate an enhanced graph; the multi-scale comparison module firstly extracts a meta-path and a sub-graph from the enhanced graph, then calculates a node-level representation and a sub-graph-level representation to construct multi-scale comparison loss, and finally calculates an anomaly score based on the similarity of the node-level representation and the sub-graph-level representation; and performing literature reference network anomaly detection based on the trained graph enhanced multi-scale contrast learning framework. According to the method, the hidden abnormal modes in the literature reference network can be well revealed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal detection of literature citation networks, and in particular to a method for abnormal detection of literature citation networks based on graph-enhanced multi-scale contrast learning. Background Art

[0002] A literature citation network is a complex network composed of various academic literatures (such as papers, monographs, patents, etc.) and their mutual citation relationships, and is a common attributed graph network. The academic literature citation behavior is an important way for the accumulation, continuation and inheritance of scientific knowledge, and is also the cornerstone for promoting academic development and innovation. However, abnormal citation behaviors such as irrelevant self-citation and forced induced citation that occur for deceptive or unethical motives will damage the rigor and fairness of scientific research, harm the scientific research environment of academic integrity, and seriously hinder scientific progress and academic development. Detecting abnormal academic literature citation behaviors, effectively supervising and regulating academic research behaviors, has important practical significance for the scientific and technological innovation development of our country.

[0003] In a literature citation network, academic literatures and their citation relationships can be respectively modeled as nodes and edges in an attributed graph network, so as to transform the literature citation network into graph-structured data. Based on this, the problem of abnormal citation detection in a literature citation network can be formalized as a graph anomaly detection task. Graph anomaly detection technology uses graph mining algorithms to identify deviations in citation patterns, providing a comprehensive solution for dealing with complex literature citation networks.

[0004] In a literature citation network, due to the difficulty of obtaining a large amount of labeled data, self-supervised learning, which is essentially unsupervised learning, has been applied in graph anomaly detection. The basic concept behind self-supervised learning is to obtain supervision signals from the data itself and obtain high-quality representations with strong generalization ability for downstream tasks. Among various self-supervised methods, contrastive learning has recently received great attention. Contrastive learning aims to create positive and negative sample pairs for comparison, with the goal of maximizing the consistency within positive sample pairs and minimizing the consistency within negative sample pairs.

[0005] Due to the complexity of graph data, anomalies are often hidden at various scales (such as node and subgraph levels). Existing studies usually use node-level and subgraph-level representations to perform multi-scale contrast. They regard the node-level and subgraph-level representations from the same target node as positive pairs, and the representations from different target nodes as negative pairs. Although these methods have achieved success, there are still two additional problems. First, when calculating node-level representations, the mainstream contrastive learning methods for anomaly detection ignore the structural and contextual information of the target node and produce insufficient representations. Existing studies only use node attributes to calculate node-level representations, completely ignoring neighbor and structural information. As a result, the calculated representations cannot capture the information of structural and contextual anomalies, leading to suboptimal performance.

[0006] Secondly, two instances with different class labels can be considered as positive pairs (referred to as inferior pairs for simplicity), and their consistency is maximized during model training. This may have an adverse effect on the robustness of the model and reduce the detection results. Due to the highly imbalanced distribution in the number of abnormal nodes and normal nodes, there are some nodes with heterophilic dominant neighbors, that is, most neighbors have class labels different from the target node (abbreviated as heterophilic nodes). Correspondingly, for a heterophilic node, its node-level and subgraph-level representations (from its attributes and neighbors respectively) often have different class labels. However, general graph contrastive learning methods regard them as positive pairs and try to force them to be consistent. This may have a negative impact on the effectiveness of the model. Figure 1 Figure (a) in [reference] gives an example of a heterophilic node, where the target node is abnormal while most of its neighbors are normal. From Figure 1 the node-level representation derived from the node attributes in Figure (b) in [reference] often has an abnormal label. However, from Figure 1 the subgraph at the subgraph level derived from the neighbors in Figure (c) in [reference] often has a normal label. Pulling them closer may reduce the model performance. Summary of the Invention

[0007] In view of the above problems, the present invention proposes an abnormal detection method for literature citation networks based on graph augmentation multi-scale contrastive learning.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention proposes a graph augmentation multi-scale contrastive learning framework (GCLAD) for abnormal detection of literature citation networks, which can adaptively operate on the neighbors of the target node to alleviate the problem of sub-optimal positive samples and explore meta-paths to enhance the node-level representation. This framework mainly consists of two parts: a graph augmentation module and a multi-scale contrast module. In the graph augmentation module, the present invention designs a diffusion node transformer based on the denoising diffusion probability model (DDPM), which can inject the features of the target node into some of its neighbors to generate an enhanced graph. This injection process is dynamically adjusted by an adaptive weight, which is calculated according to the attribute curvature designed by the present invention. The attribute curvature is a mixture of graph curvature and node attributes, which can reflect the intensity of structural overlap between a pair of nodes considering the node attributes. This injection strategy can alleviate the inconsistency problem between the node-level and subgraph-level representations of heterogeneous nodes. This is because for heterogeneous nodes, the injection operation enables their features to be integrated into adjacent nodes. Therefore, the subgraph-level representation derived from its neighbors will be closer to the node-level representation obtained from the node itself. As a result, the two representations become more consistent. Therefore, the enhanced graph can improve the subgraph-level representation and alleviate the inconsistency problem, ultimately benefiting the model performance.

[0010] In the multi-scale contrast module, based on the enhanced graph, the present invention constructs a multi-scale contrast learning network, which conducts contrast learning at the node and subgraph levels to discover hidden abnormal patterns in complex graphs. For each target node, the present invention introduces meta-paths to utilize some relevant neighbors to enhance the node-level representation and calculates the subgraph-level representation based on the enhanced subgraph. Then, in addition to node-subgraph and node-node contrasts, the present invention further explores enhanced graph-based subgraph-subgraph contrasts to construct a multi-scale contrast network, the goal of which is to maximize the consistency between representations from the same target node. Therefore, the representations obtained from this framework are used to calculate anomaly scores to identify abnormal nodes.

[0011] Specifically, the present invention proposes a method for anomaly detection in a literature citation network based on graph augmentation multi-scale contrast learning, including:

[0012] Step 1: Construct an attribute network of the literature citation network. In the attribute network, nodes represent literature, edges represent citation relationships, and the attributes of nodes correspond to the attributes of literature.

[0013] Step 2: Construct a graph augmentation multi-scale contrast learning framework for anomaly detection in the literature citation network. The graph augmentation multi-scale contrast learning framework includes a graph augmentation module and a multi-scale contrast module. The graph augmentation module, based on the denoising diffusion probability model, injects the features of the target node in the attribute network into some of the target node's neighbors to generate an enhanced graph. The multi-scale contrast module first extracts meta-paths and subgraphs from the enhanced graph, then calculates the node-level representation and the subgraph-level representation to construct a multi-scale contrast loss. Finally, the anomaly score is calculated based on the similarity of the node-level and subgraph-level representations.

[0014] Step 3: Conduct anomaly detection in the literature citation network based on the trained graph augmentation multi-scale contrast learning framework.

[0015] Furthermore, in the graph augmentation module, a node transformer is constructed based on the denoising diffusion probability model. For a given neighbor node of the target node, the node transformer takes the prior from this neighbor node as input and the target node as a condition, and these conditions act on the reverse diffusion process of the denoising diffusion probability model to generate a transformed node. The node transformer first adds noise to the given neighbor node using the forward process to form a prior, and then inputs the prior into the reverse diffusion process to generate a clean embedding through step-by-step denoising, and replaces a proportion of the target node's neighbors with the generated embedding to construct the enhancement Figure 1 and replaces another proportion of the target node's neighbors with the generated embedding to construct the enhancement Figure 2 .

[0016] Further, during the denoising process, the features of the target nodes are extracted, and the target nodes are iteratively injected into the latent variables using adaptive weights.

[0017] Further, the adaptive weights are calculated based on the attribute curvature, which is a mixture of graph curvature and node attributes.

[0018] Further, the attribute curvature is calculated as follows:

[0019]

[0020] where C(i,j) is the attribute curvature of the node pair (v i ,v j ), d(i,j) is the graph distance between v i and v j , W(,) refers to the Wasserstein distance between two probability distributions; m i (j) is the probability distribution vector defined by node v i relative to node v j considering node attributes; m j (i) is the probability distribution vector defined by node v j relative to node v i considering node attributes; In the formula:

[0021]

[0022] where x = 1,..., N represents the index of m i (j), δ is a coefficient with a value in [0,1], represents the set of common neighbors of v i and v j , refers to the difference set of the neighbor sets of v i and v j , S i ′ j = S ij / (1 - δ) represents the normalized similarity, and S ij represents the attribute similarity between v i and v j .

[0023] Further, the average attribute curvature is used as the adaptive weight.

[0024] Further, in the multi-scale contrast module, for a given target node, multiple meta-paths are extracted, and the neighbors of the target node induced by the meta-paths are screened. PathSim is used to measure the similarity between nodes during the screening process. The screened meta-paths are input into a shared graph convolutional network to learn the node-level representation, and the average pooling function is used as the readout module to obtain the node-level representation.

[0025] Further, in the multi-scale contrast module, the random walk with restart algorithm is used to sample the subgraph around the target node. The subgraph is input into a new shared graph convolutional network to learn the subgraph-level representation, and the attribute information of the target node in the subgraph is masked. Then, the readout module is used to obtain the subgraph-level representation.

[0026] Further, in the multi-scale contrast module, the multi-scale contrast loss is constructed in the following way:

[0027]

[0028] In the formula:

[0029]

[0030]

[0031] Among them, γ1, γ2, and γ3 are respectively and the balance parameters of the three losses, and are the node-subgraph contrast loss, the node-node contrast loss, and the subgraph-subgraph contrast loss respectively; λ is a trade-off parameter used to balance and the importance of the two losses; and are respectively the node-subgraph losses corresponding to enhancing Figure 1 and enhancing Figure 2 ; represents the similarity between the node-level representation corresponding to enhancing Figure 1 and the subgraph-level representation, W s is a learnable matrix, represents the Sigmoid activation function, represents the subgraph-level representation of enhancing Figure 1 ; represents and the similarity between them, and represent the node-level representations of enhancing Figure 1 and enhancing Figure 2 respectively. In the positive sample pair, yi equals 1, and in the negative sample pair, y i equals 0, represents and the similarity between

[0032] Furthermore, in the multi-scale contrast module, the anomaly score of the target node is calculated in the following manner:

[0033]

[0034] In the formula:

[0035]

[0036] where R is the number of rounds of anomaly detection, represents the anomaly score of the r-th round of anomaly detection, AS avg represents the average value of the multi-round anomaly scores, and represent the similarities of the negative sample pair and the positive sample pair respectively.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] (1) The present invention proposes a graph-enhanced multi-scale contrast learning framework for anomaly detection in literature citation networks, which explores subgraph-subgraph contrast based on the enhanced graph to reveal hidden anomaly patterns in complex graphs.

[0039] (2) The present invention designs a diffusion-based graph enhancement module to adaptively manipulate neighbor nodes to generate an enhanced graph. This can effectively improve the representation ability at the subgraph level and alleviate the problem of suboptimal positive sample pairs.

[0040] (3) The present invention uses some relevant neighbors on the meta-path to enhance the node-level representation. A large number of experiments conducted on five datasets demonstrate the superiority of the graph-enhanced multi-scale contrast learning framework proposed by the present invention compared to the state-of-the-art benchmark models. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is an example of a heterophilic node;

[0042] Figure 2 is a flowchart of a method for anomaly detection in a literature citation network based on graph-enhanced multi-scale contrast learning according to an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the GCLAD framework provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] For ease of understanding, the following explanations are given for some of the terms that appear in the specific embodiments of the present invention:

[0045] (I) Problem description

[0046] Generally speaking, an attribute network can be represented as G=(V, E, X), where V={v1, v2,..., v n} represents the set of nodes, E={e1, e2,..., e m} represents the set of edges, and X={x1, x2,..., x n}∈R n×h represents the attributes of n nodes, and each node has h dimensions. The structural information of the attribute network can be represented by a binary adjacency matrix, denoted as A∈R n×n , where if there is a connection between nodes v i and v j , then A i,j = 1, otherwise A i,j = 0. Since both V and E are embedded in A, the attribute network can also be represented as G=(A, X). In this case, the anomaly detection task can be formulated as follows:

[0047] Problem 1 (Anomaly detection on attribute networks): The goal is to train a function using multiple unlabeled samples. Given an attribute network G=(V, E, X) composed of nodes V={v1,..., v n}. This function aims to calculate an anomaly score for each node to indicate the degree of anomaly. By sorting all nodes according to the anomaly scores, the anomaly nodes can be detected based on their respective sorting positions.

[0048] (II) Denoising Diffusion Probability Model

[0049] The Denoising Diffusion Probability Model (DDPM) belongs to a class of generative models that perform extremely well in unconditional image generation. It learns a Markov chain that gradually transforms a simple distribution (e.g., an isotropic Gaussian distribution) into the data distribution. The generation process is the inverse of the DDPM forward (diffusion) process, in which the Markov chain gradually adds noise to the data. Here, each step of the forward process is a Gaussian transformation:

[0050]

[0051] where β1,..., β T is a fixed variance schedule and not a learned parameter. Equation (1) is a process of finding z t-1 by adding small Gaussian noise to the latent variable z t .

[0052] Given the clean data z 0 , z t can be sampled in closed form as:

[0053]

[0054] where α t := 1 - β t and Thus, z t can be expressed as a linear combination of z 0 and ∈:

[0055]

[0056] where has the same dimension as the data z 0 and the latent variable z 1 ,…, z T .

[0057] Since the inverse process q(z t-1 |z t ) of the forward process is intractable, DDPM learns a parameterized Gaussian transformation p θ (z t-1 |z t ). The generation (or reverse) process has the same functional form as the forward process and is represented as a Gaussian transformation with a learned mean and a fixed variance:

[0058]

[0059] By decomposing μ θ into a linear combination of z t and the noise approximator ∈ θ , the generation process is represented as:

[0060]

[0061] where ∈ is the noise, which indicates that each generation step is stochastic. Here ∈ θ represents a neural network with the same input and output dimensions, and the noise ∈ θ predicted by the neural network at each step is used for the denoising process in Equation (5).

[0062] The present invention will be further explained and illustrated below with reference to the accompanying drawings and specific embodiments:

[0063] As Figure 2 shown, a method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrast learning includes:

[0064] Step S101: Construct an attribute network of the literature citation network. In the attribute network, nodes represent literature, edges represent citation relationships, and the attributes of nodes correspond to the attributes of the literature.

[0065] Step S102: Construct a graph augmentation multi-scale contrastive learning framework for anomaly detection in the literature citation network. The graph augmentation multi-scale contrastive learning framework includes a graph augmentation module and a multi-scale contrast module. The graph augmentation module, based on the denoising diffusion probability model, injects the features of the target node in the attribute network into some of its neighbors, thereby generating an augmented graph. The multi-scale contrast module first extracts meta-paths and subgraphs from the augmented graph, then calculates node-level representations and subgraph-level representations to construct a multi-scale contrastive loss. Finally, an anomaly score is calculated based on the similarity of the node-level and subgraph-level representations.

[0066] Step S103: Perform anomaly detection on the literature citation network based on the trained graph augmentation multi-scale contrastive learning framework.

[0067] Specifically, the framework GCLAD proposed by the present invention consists of two main modules, as Figure 3 shown. The graph augmentation module aims to operate on the neighbor nodes of the target node to generate an augmented graph. The multi-scale contrast module first extracts meta-paths and subgraphs from the augmented graph, then calculates node-level and subgraph-level representations to construct a multi-scale contrastive loss. Finally, an anomaly score is calculated based on the similarity of the node-level and subgraph-level representations.

[0068] (I) Graph Augmentation Module

[0069] Typical contrast-based anomaly detection methods attempt to enforce the consistency between node-level representations and subgraph-level representations to enhance node representation capabilities. However, for a heterogeneous node, the subgraph-level representation (extracted from its neighbors) may have different class labels from the node-level representation (extracted from node attributes). Enforcing their consistency will have a negative impact on the model's effectiveness. To address this issue, we design a diffusion-based node transformer that attempts to inject the features of the target node into its neighbor nodes through adaptive weights to generate an augmented graph. The augmented graph can make the subgraph-level representation closer to the node-level representation, alleviate the inconsistency problem between them, and enhance the robustness of the model.

[0070] The present invention introduces the denoising diffusion probability model (DDPM) as the basic framework for constructing the node transformer. For a given neighbor node of the target node, this transformer takes the prior from this neighbor (source node) as input and the target node (reference node) as a condition, and these conditions act on the reverse diffusion (denoising) process of the DDPM to generate a transformed node. The transformer first adds noise to the neighbor node z src to form a prior z T。Then, the prior z T is input into the reverse diffusion process to generate a clean embedding by gradually denoising, that is In this denoising process, the features of the reference nodes are extracted and iteratively injected into the latent variables using adaptive weights . In this way, the generated embedding z 0 retains the features of the source nodes because the model starts from the prior originating from the source nodes, and also has the features of the reference nodes because they are regarded as the conditions acting on the generation process. At the same time, the injection weights are adaptively adjusted according to the node types, that is, larger weights are assigned to heterogeneous nodes and smaller weights are assigned to homogeneous nodes. Therefore, for heterogeneous nodes, this injection can effectively alleviate the inconsistency problem. For homogeneous nodes, it hardly affects the results. So, this feature fusion strategy can improve the model performance

[0071] Adaptive weights. For a given target node, when the class labels of its neighbor nodes are different, more features from the target (reference) node should be incorporated into the neighbor nodes to alleviate the inconsistency problem. Therefore, in the injection process of denoising diffusion, we try to assign larger weights to such nodes. We introduce curvature considering attribute similarity as the weight

[0072] Curvature is used to measure the interaction and overlap strength between a pair of nodes. Node pairs with the same class label will exhibit higher curvature values due to having more shared neighbors. Therefore, curvature can reflect the structural correlation between the target node and its neighbors, that is, the curvature value between a heterogeneous node and its neighbors is lower, and vice versa. However, in addition to the structure, anomalies in the graph may also be caused by node attributes. Therefore, we further incorporate attribute information into the curvature and design attribute curvature so that it can reflect both structural information and attribute information, and use this as the weight

[0073] Specifically, we consider attribute similarity when calculating the curvature, that is, node pairs with higher attribute similarity will be assigned higher curvature values. This consideration can increase the curvature values of homogeneous node pairs and decrease the curvature values of heterogeneous node pairs, which helps to alleviate the inconsistency problem. For a pair of nodes (v i , v j ), we calculate its attribute curvature as follows

[0074]

[0075] where d(i,j) is the graph distance between v i and v j , and W(,) refers to the Wasserstein distance between two probability distributions, that is, the minimum average moving distance through any transportation plan. m i(j) is the probability distribution vector defined for node v after considering node attributes; m i relative to node v j ; m j (i) is the probability distribution vector defined for node v after considering node attributes; j relative to node v i . For the calculation of m i (j), we do not consider all neighbors of the node equally. Instead, we divide the neighbors into two groups: the common neighbors of v i and v j , and the non-common neighbors. When calculating m i (j), we treat these two groups of neighbors differently to emphasize the importance of the common neighbors. Based on their attribute similarity, the common neighbors are given a greater weight. Specifically, as follows:

[0076]

[0077] where x = 1, …, N represents the index of the distribution vector m i (j), δ is a coefficient with values in [0, 1], represents the set of common neighbors of v i and v j , while refers to the difference set of the neighbor sets of v i and v j . S i ′ j = S ij / (1 - δ) represents the normalized similarity, where S ij represents the attribute similarity between v i and v j .

[0078] Conditional diffusion. To handle neighbor nodes smoothly, we select multiple nodes as reference nodes, which are regarded as the conditions for the generation process. Here, we select reference nodes according to the node embedding similarity with the target node. To achieve this, we apply a graph convolutional network (GCN) layer to calculate the local embedding, which captures the local information of each node. Given the input node features X and the adjacency matrix A, the output embedding is represented as:

[0079]

[0080] In this process, we embed the feature vectors into the hidden representation. Based on these embeddings, we calculate the similarity between the target node embedding and other node embeddings, and select the top n nodes with the highest similarity as reference nodes (e.g., z ref1 , …, z refn ).

[0081] Then, we use the forward process q(z t |z0) to generate the prior z by adding noise to the source node z src : T

[0082]

[0083] where β1, …, β T is a fixed variance schedule. Further, we input the prior z T into the reverse diffusion process and impose the condition c ref (features of the reference node) in this reverse process to generate an embedding with the features of both z src and z ref .

[0084] To inject the features of the reference node into the generated data, we approximate the Markov transition under the condition c ref as follows:

[0085]

[0086] where is sampled from Equation (3), f l (·) is a low-pass filter, and σ is an aggregation function like mean that concatenates

[0087] Equation 10 attempts to incorporate into the generated data. Thus, the generated data will have their combined features. According to this equation, in each transition from z t to z t-1 , the features of the reference node are extracted and then injected into the latent variables. To this end, we use the forward process (Equation 3) to calculate ref1 , …, z refn separately from z

[0088]

[0089] Next, we use the reverse process (Equation 4) to calculate the latent variables t from z

[0090]

[0091] Thus, the generated embedding is optimized by matching with as follows:

[0092] ​

[0093] Among them, refers to the average attribute curvature calculated by formula 6, that is weight is used to adjust the degree of feature injection. In other words, when the target node is more different from its neighbors (resulting in a smaller attribute curvature), more features from the target node will be incorporated into its neighbors. This can alleviate the inconsistency problem.

[0094] In this way, during the generation process, the features of the reference node are injected into the latent variable, and the generated embedding can possess the features of z ref Meanwhile, the input of the generation process is the prior of z src from the neighbors themselves, so the generated embedding still retains the features of z src

[0095] After obtaining the generated embeddings, we replace a certain proportion of the neighbors of the target node with these generated embeddings to construct an enhanced Figure 1 enhanced graph, and select another proportion for constructing the enhanced Figure 2 These enhanced graphs will be used in the contrastive learning model.

[0096] (II) Multi-scale Contrast Module

[0097] Based on the enhanced graph, we further mine node-level and subgraph-level representations to extract the local features of nodes and the global structural information of subgraphs. These representations will be further used for multi-scale contrast.

[0098] Node-level representation. In graph anomaly detection methods based on contrastive learning, node-level representations are usually calculated only based on node attributes. These representations may not be sufficient because they ignore neighbor and structural information during the calculation. For a given target node, to obtain its effective node-level representation, we try to consider some of its similar neighbors to strengthen the node-level representation. For this purpose, we propose a meta-path-based representation method that uses the meta-path from the target node to its neighbors to calculate the node-level representation. Here, the meta-path is a sequence of object types that captures the semantic relationships between objects in the heterogeneous information network. In other words, the nodes in the meta-path can express important semantic relationships while being close to the target node. For the enhanced graph, when the processed nodes and the original nodes are regarded as different types, they can be regarded as a heterogeneous graph. Note that the target node is considered to be of the same type as the processed node because they have similar features. Therefore, we use the meta-path to identify a set of path-based important neighbors that are semantically related to the target node. The representation learned from the meta-path can capture more neighbor information while retaining the inherent features of the target node.​

[0099] Specifically, for a given target node v i , multiple meta-paths can be extracted. To select relevant neighbors, we screen the neighbors of v i induced by the meta-path and select the neighbors most relevant to v i . We use PathSim to measure the similarity between nodes. Specifically, given a meta-path P, the similarity between two nodes v i and v j is calculated by the following formula:

[0100]

[0101] where is the path instance between v i and v j . Based on these similarities, for each node, we select the top k neighbors with the highest similarity to it to learn the node-level representation.

[0102] Next, we input the meta-path into a shared graph convolutional network (GCN) to learn the node-level representation, i.e., a low-dimensional embedding. The GCN can be expressed as:

[0103]

[0104] where is 's degree matrix, is the subgraph adjacency matrix, I N is the identity matrix, φ is the activation function (such as ReLU), and W 0 ∈R F×d is the learnable weight matrix. To apply the node-subgraph matching pattern self-supervised task (the representations derived from nodes and their corresponding subgraphs are more consistent), we use the average pooling function as the readout module to obtain the node-level representation h i :

[0105]

[0106] where n i is the number of nodes in the meta-path.

[0107] Subgraph-level representation. To calculate the subgraph-level representation, we first perform subgraph sampling to obtain the corresponding subgraphs, and then input these subgraphs into the graph convolutional network (GCN) to calculate the representation. For a target node, an effective anomaly detection method is to measure the feature distance between it and its neighborhood. Therefore, we use the random walk with restart algorithm to sample the subgraphs around the node.

[0108] Then, we input the subgraphs into a new shared graph convolutional network to learn subgraph-level representations. Note that to make the resulting representations more discriminative, the attributes of the target nodes in the subgraphs are masked. The new graph convolutional network is defined as follows:

[0109]

[0110] where \(W\) ′0 is different from the parameter matrix in the node-level representation calculation. Then, a readout module is used to obtain the subgraph-level representation \(e\) i :

[0111]

[0112] where \(n\) i is the number of nodes in the subgraph.

[0113] Multi-scale contrastive loss. Based on the augmentations \(\hat{G}_1\) and \(\hat{G}_2\) generated by the graph augmentation module, we first calculate their node-level representations \(\mathbf{h}_1\) and \(\mathbf{h}_2\) as well as subgraph-level representations \(\mathbf{e}_1\) and \(\mathbf{e}_2\). Then, a multi-scale contrastive loss consisting of node-subgraph loss (\(L_{ns}\)), node-node loss (\(L_{nn}\)), and subgraph-subgraph loss (\(L_{ss}\)) is constructed. Figure 1 and augmentation Figure 2 and and

[0114] Figure 1 NS ), node-node loss (\(L_{nn}\)) NN ) and subgraph-subgraph loss (\(L_{ss}\)) SS ).

[0114] Node-subgraph loss. For the node-level representation of the target node, the subgraph-level representations in each augmented graph are regarded as positive sample pairs, while the subgraph-level representations corresponding to other nodes are regarded as negative sample pairs. Specifically, for augmentation \(\hat{G}_1\), the abnormality degree of the target node is related to the similarity between its node-level representation and subgraph-level representation: Figure 1

[0115]

[0116] where \(W\) s is a learnable matrix, denotes the Sigmoid activation function.

[0117] Generally speaking, for the similarity in positive sample pairs, the node-level representation and subgraph-level representation tend to be similar, i.e., On the contrary, in negative sample pairs, they may not be similar, i.e., Therefore, we use binary cross-entropy loss to train this contrastive relationship:

[0118]

[0119] Among them, in the positive sample pair, y i equals 1, and in the negative sample pair, y i equals 0. Similarly, for the augmentation Figure 2 , we can obtain the similarity and the loss Therefore, the final node-subgraph loss is:

[0120]

[0121] where λ is a trade-off parameter used to balance the importance of these two losses.

[0122] The node-node loss and the subgraph-subgraph loss. Similar to the node-subgraph loss, we calculate the similarity and between and construct the node-node loss as:

[0123]

[0124] where the meaning of y i is the same as that in formula (20). Similarly, the subgraph-subgraph loss is defined as:

[0125]

[0126] where represents the similarity and between.

[0127] Therefore, the final objective function is a combination of these losses:

[0128]

[0129] where γ1, γ2, and γ3 are the balance parameters of these three losses.

[0130] (III) Anomaly Score Calculation

[0131] In the comparison between the node and subgraph representations, for normal nodes, their positive sample pairs show similarity, while their negative sample pairs show dissimilarity. However, for abnormal nodes, it is often difficult to distinguish the degree of abnormality between their positive and negative sample pairs. Therefore, we define the anomaly score of the target node as:

[0132]

[0133] where and represent the similarities of the negative sample pair and the positive sample pair respectively.

[0134] We use multiple rounds of detection to mitigate the bias of random walk sampling. Accordingly, we calculate the final anomaly score based on the mean and standard deviation of the results of multiple rounds of detection:

[0135]

[0136] where R is the number of rounds of anomaly detection, and AS avg represents the mean value.

[0137] (IV) Complexity Analysis

[0138] We analyze the time complexity of the proposed framework by considering three main components separately. For the graph enhancement module, the complexity of the diffusion model is mainly dominated by U-Net and the Wasserstein distance. For U-Net, following the measurement method of deep neural networks, we use the number of parameters and the number of floating-point operations (FLOPs) to analyze the complexity of U-Net. The FLOPs are expressed as where D is the depth of the network, M is the length of the feature map, K is the kernel of the l-th layer, and C is the number of channels. For the Wasserstein distance, we utilize the closed-form solution of optimal transport and can compute this distance very efficiently with a time complexity of O(N logN). For the multi-scale contrast module, the time complexity is mainly generated by the graph neural network (GNN) module, with a time complexity of O(N 2 ) for each pair, and an overall time complexity of O(N 2 dR), where d represents the average node degree in the graph. For the anomaly score calculation, its time complexity is much lower than the above two stages, so we ignore this term. In summary, the complexities of the graph enhancement module and the multi-scale contrast module are comparable to the existing diffusion model and graph contrast learning model respectively. Therefore, the computational complexity is not too high, facilitating subsequent deployment on typical hardware platforms.

[0139] To verify the effectiveness of the present invention, the following experiments are conducted:

[0140] In this section, we conduct experimental evaluations to demonstrate the effectiveness of our method in terms of anomaly detection performance and the roles of the designed components and hyperparameters.

[0141] (a) Experimental Setup

[0142] Table 1 Dataset Information

[0143]

[0144]

[0145] Datasets. We use three widely used citation network datasets (Cora, Citeseer, and PubMed) for performance evaluation. Their statistics are shown in Table 1. These three citation networks are composed of scientific publications. Here, nodes and edges represent scientific publications and the citation relationships between two publications respectively. The attribute vector of each node is a bag-of-words representation, and its dimension is determined by the dictionary size. Since there is no real abnormal data in the above datasets, for easy evaluation, synthetic abnormal data needs to be injected into the clean attributed network. We inject a set of comprehensive anomalies (i.e., structural anomalies and context anomalies) for each dataset.

[0146] Baseline methods. We compare the proposed framework with two types of unsupervised baseline methods using different techniques, including reconstruction-based methods and contrast-based methods.

[0147] (1) Autoencoder is an unsupervised deep autoencoder model based on features, which introduces an anomaly regularization penalty term based on L1 or L2 norm.

[0148] (2) DOMINANT is an unsupervised anomaly detection framework that uses graph convolutional autoencoders to reconstruct the adjacency matrix and attribute matrix. Then, the anomaly score is calculated through the reconstruction error.

[0149] (3) Anomalous is a method based on residual analysis. It is a joint learning framework that takes attribute selection and anomaly detection as a whole based on CUR decomposition and residual analysis.

[0150] (4) ALARM is a multi-view framework for incorporating user preferences into anomaly detection and simultaneously processing heterogeneous attribute features through multiple graph encoders.

[0151] (5) CoLA is a method based on contrastive learning that uses subgraph sampling to augment graph data and constructs a contrastive loss for anomaly detection.

[0152] (6) ANEMONE is a multi-scale contrastive learning framework that uses a graph neural network backbone encoder to capture the pattern distribution of graph data by simultaneously learning the consistency between patch-level and context-level instances.

[0153] (7) SL-GAD is a self-supervised learning model for graph anomaly detection that constructs different context subgraphs based on target nodes and uses two modules, namely generative attribute regression and multi-view contrastive learning, for anomaly detection.

[0154] (8) Sub-CR is a self-supervised learning framework that jointly optimizes a module based on multi-view contrastive learning and a module based on attribute reconstruction to more accurately detect anomalies in attributed networks.

[0155] (9) GRADATE is a multi-scale contrastive learning framework that explores subgraphs for contrastive learning to capture abnormal features at different scales for graph anomaly detection.

[0156] Experimental settings. For the graph augmentation module, considering both efficiency and effectiveness, we use 1000-step diffusion. We obtain the best parameters for each dataset through grid search and record the results with the highest F1 score. For the meta-path, we select the top 3 neighbors with the highest similarity from 1 to 9 to learn node-level representations. For neighbor processing, we replace 20% and 40% of the neighbors of each target node with the generated nodes to generate two augmented graphs. We train the model using the Adam optimizer with a learning rate of 0.0001. The trade-off parameter λ in formula (21) is empirically set to 0.6. The accuracy is insensitive to the number of reference nodes n. Therefore, we set n to 4 because it gives slightly better results. For the baseline methods, we adopt the configurations recommended in their papers. We use the following three metrics for comprehensive evaluation: Macro-F1 is the unweighted average of the F1 scores of the two classes. AUC-ROC is the area under the Receiver Operating Characteristic curve (ROC curve), which refers to the probability that a randomly selected abnormal sample score is higher than that of a randomly selected normal sample. AUC-PR is the area under the precision-recall curve at different thresholds.

[0157] (b) Comparison of detection performance

[0158] We report the anomaly detection performance of GCLAD and each baseline method in Table 2 and have the following observations: (1) In terms of the overall detection results, on the three datasets, the performance of GCLAD is consistently better than that of all baseline methods. The performance improvement is mainly attributed to the fact that the augmented graphs effectively improve the quality of the selected positive sample pairs, and the meta-path effectively enhances the node-level representations. (2) Typical contrastive learning-based methods, such as CoLA, ANEMONE, and SL-GAD, are better than non-contrastive learning models, such as Autoencoder and ALARM. This shows that the contrastive learning-based mode can effectively detect anomalies by mining the feature and structure information in the graph. (3) Among the contrastive learning-based baseline methods, GRADATE performs better, which indicates that considering multi-scale contrast is helpful for graph anomaly detection. However, by improving the quality of the selected positive sample pairs and enhancing the node representations, GCLAD achieves the best performance.

[0159] In addition, we studied the convergence speed of our model and contrast-based baseline methods. The results show that the convergence speed of our model GCLAD is similar to that of contrastive learning models such as GRADATE. This is because both of them adopt the contrastive learning framework, and the differences between them (e.g., graph augmentation methods and meta-path-based representations) have little impact on the convergence speed.

[0160] Table 2 Performance comparison between GCLAD and baseline methods on five datasets

[0161]

[0162]

[0163] In summary, the present invention proposes a graph-augmented multi-scale contrastive learning framework GCLAD for anomaly detection in literature citation networks. In this framework, we design a graph augmentation module based on the diffusion model to adaptively process neighbor nodes to generate augmented graphs, which can alleviate the inconsistency problem between node-level and subgraph-level representations. In addition, we utilize some relevant neighbors on the meta-path to strengthen the node-level representation and construct multi-scale contrast to discover hidden anomaly patterns. Experiments conducted on five datasets show that the method we designed achieves the current optimal performance.

[0164] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An abnormal detection method for literature citation networks based on graph-enhanced multi-scale contrast learning, characterized in that, Including: Step 1: Construct an attribute network of the literature citation network. In the attribute network, nodes represent literature, edges represent citation relationships, and the attributes of nodes correspond to the attributes of literature. Step 2: Construct a graph-enhanced multi-scale contrast learning framework for anomaly detection of the literature citation network. The graph-enhanced multi-scale contrast learning framework includes a graph enhancement module and a multi-scale contrast module. The graph enhancement module, based on the denoising diffusion probability model, injects the features of the target node in the attribute network into some of the target node's neighbors to generate an enhanced graph. The multi-scale contrast module first extracts meta-paths and subgraphs from the enhanced graph, then calculates node-level representations and subgraph-level representations to construct a multi-scale contrast loss. Finally, the anomaly score is calculated based on the similarity of node-level and subgraph-level representations. Step 3: Perform anomaly detection on the literature citation network based on the trained graph-enhanced multi-scale contrast learning framework.

2. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 1, wherein In the graph enhancement module, a node transformer is constructed based on the denoising diffusion probability model. For a given neighbor node of the target node, the node transformer takes the prior from this neighbor node as input and the target node as a condition. These conditions act on the reverse diffusion process of the denoising diffusion probability model to generate a transformed node. The node transformer first adds noise to the given neighbor node using the forward process to form a prior, and then inputs the prior into the reverse diffusion process. By gradually denoising, a clean embedding is generated. A proportion of the target node's neighbors are replaced with the generated embedding to construct Enhanced Graph 1, and another proportion of the target node's neighbors are replaced with the generated embedding to construct Enhanced Graph 2.

3. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 2, wherein During the denoising process, the features of the target node are extracted, and the target node is iteratively injected into the latent variable using an adaptive weight.

4. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 3, wherein The adaptive weight is calculated based on the attribute curvature, which is a mixture of graph curvature and node attributes.

5. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 4, wherein Calculate the attribute curvature in the following way: where C(i,j) is the attribute curvature of the node pair (v i ,v j ), d(i,j) is the graph distance between v i and v j , W(,) refers to the Wasserstein distance between two probability distributions; m i (j) is the probability distribution vector defined by node v i relative to node v j after considering node attributes; m j (i) is the probability distribution vector defined by node v j relative to node v i after considering node attributes; In the formula: where x = 1, …, N represents the index of m i (j), δ is a coefficient with a value in [0, 1], represents v i and v j 's common neighbor set, refers to the difference set between the neighbor sets of v i and v j , S i ′ j = S ij / (1 - δ) represents the normalized similarity, S ij represents the attribute similarity between v i and v j .

6. The method for abnormal detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 5, characterized in that Take the average attribute curvature as the adaptive weight.

7. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 1, wherein In the multi-scale contrast module, for a given target node, multiple meta-paths are extracted, and the neighbors of the target node induced by the meta-paths are screened. PathSim is used to measure the similarity between nodes during the screening. The screened meta-paths are input into a shared graph convolutional network to learn node-level representations, and an average pooling function is used as the readout module to obtain node-level representations.

8. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 7, wherein In the multi-scale contrast module, the random walk with restart algorithm is used to sample the subgraph around the target node. The subgraph is input into a new shared graph convolutional network to learn subgraph-level representations, and the attribute information of the target node in the subgraph is masked. Then, a readout module is used to obtain subgraph-level representations.

9. The method for anomaly detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 2, wherein In the multi-scale contrast module, construct the multi-scale contrast loss in the following way: Where: Among them, γ1, γ2, and γ3 are respectively and the balance parameters of three losses, and are the node-subgraph contrast loss, the node-node contrast loss, and the subgraph-subgraph contrast loss respectively; λ is a trade-off parameter used to balance and the importance of two losses; and are the node-subgraph losses corresponding to the enhanced graph 1 and the enhanced graph 2 respectively, represents the similarity between the node-level representation and the subgraph-level representation corresponding to the enhanced graph 1, W s is a learnable matrix, represents the Sigmoid activation function, represents the subgraph-level representation of the enhanced graph 1, represents and the similarity between them, and represent the node-level representations of the enhanced graph 1 and the enhanced graph 2 respectively. In the positive sample pair, y i is equal to 1, and in the negative sample pair, y i is equal to 0, represents and the similarity between them.

10. The method for abnormal detection of a literature citation network based on graph augmentation multi-scale contrastive learning according to claim 1, wherein, In the multi-scale contrast module, calculate the anomaly score of the target node in the following way: Where: where R is the number of rounds of anomaly detection, denotes the anomaly score for the r-th round of anomaly detection, AS avg denotes the average of the anomaly scores over multiple rounds, and denote the similarities of negative sample pairs and positive sample pairs, respectively.