Attribute network anomaly detection method based on reconstruction bias learning

Through the method based on reconstruction bias learning, the loss function and classifier are dynamically adjusted, the model is forced to fit the normal mode and highlight the abnormal mode, which solves the problem of degradation of abnormal detection performance in the existing technology, and achieves more efficient attribute network abnormal detection.

CN120337079APending Publication Date: 2025-07-18HUNAN NORMAL UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510473323.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the existing attribute network anomaly detection method accounts for a high proportion of abnormal samples, the model tends to fit the normal mode and ignores the abnormal mode, resulting in a significant decline in detection performance and it is difficult to effectively identify abnormal nodes in the graph data.

Method used

Using a method based on reconstruction bias learning, the loss function and classifier are dynamically adjusted through the graph reconstruction module, the reconstruction bias dynamic adjustment module, the exception enhancement classification module and the exception score calculation module, the loss function and classifier are dynamically adjusted, the model bias fits the normal mode and highlights the abnormal mode, and the graph autoencoder and graph convolution network are used to extract node features, and the abnormal score is calculated based on structure and attribute reconstruction errors.

Benefits of technology

Without significantly increasing the computational complexity, the performance of attribute network abnormal detection is significantly improved, and the recognition accuracy and detection effect of abnormal nodes are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_11
    Figure SMS_11
  • Figure SMS_17
    Figure SMS_17
Patent Text Reader

Abstract

The invention discloses an attribute network anomaly detection method based on reconstruction bias learning. The attribute network anomaly detection method based on reconstruction bias learning is composed of a graph reconstruction module, a reconstruction bias dynamic adjustment module, an anomaly enhancement classification module and an anomaly score calculation module. And under the condition that the calculation complexity is not obviously increased, the property network anomaly detection performance is obviously improved. The method comprises the following specific conditions: firstly, a graph reconstruction module adopts a graph auto-encoder, learns a potential mode of graph data by minimizing a reconstruction error, and measures an abnormal degree by using a difference degree between node reconstruction information and original information; secondly, a reconstruction deviation dynamic adjustment module continuously interacts with the graph reconstruction module in the iterative optimization process of the graph reconstruction module, and dynamically modifies a loss function penalty coefficient to force the graph reconstruction module to deviate to fit a normal mode; thirdly, an anomaly enhancement classification module takes a pseudo normal node set and a pseudo abnormal node set which are finally screened out by the former as training samples, and an anomaly score is calculated by utilizing a classification probability, so that the anomaly performance is enhanced; and finally, the abnormal score calculation module combines the abnormal scores of the graph reconstruction module and the abnormal enhancement classification module to calculate the final abnormal score of each node, thereby achieving the purpose of abnormal detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and network security, and particularly to an attribute network anomaly detection method based on reconstruction-biased learning. Background Art

[0002] With the in-depth development of the Internet era, the data generated in various fields is becoming increasingly complex and interrelated. These data usually exist in the form of attribute networks or attribute graphs, which are represented as sets of nodes and edges, carrying rich entity and relationship information. However, the massive growth and quality differences of data have brought great challenges to the modeling and analysis of complex networks, especially in the crucial task of anomaly detection. Anomaly detection on attribute networks has become the focus of attention in academia and industry due to its wide application value, such as financial fraud detection, social spam detection, network intrusion detection, and malicious content filtering in social media.

[0003] Attribute network anomaly detection refers to the process of identifying abnormal nodes that are significantly different from most nodes in an attribute network. According to previous studies, node anomalies on attribute networks present two classic situations: attribute anomalies and topological anomalies. When a node has an attribute anomaly, its attributes are significantly different from those of adjacent nodes, but the network structure may be similar to other nodes. For example, in a computer network, a damaged device will exhibit attribute characteristics different from other devices; topological anomalies refer to those nodes that are significantly different from other nodes in terms of network structure but similar in attributes. For example, members of a fraud gang often collude to carry out malicious activities, forming a densely connected subgraph.

[0004] Although the definition of anomalies is clear, due to the sparsity of anomaly labels in practice, attribute graph anomaly detection often proceeds in an unsupervised manner. The complexity of graph data and the diversity of anomaly patterns make the task of attribute graph anomaly detection still extremely challenging. Many early traditional methods relied on specific expert knowledge and used methods such as feature engineering or statistical identification for anomaly detection, but were limited by the difficulty of effectively mining deep non-linear information in graph data, resulting in insufficient model detection performance.

[0005] The introduction of graph neural networks (GNNs) has brought significant progress to attribute graph anomaly detection. Deep learning methods are mainly divided into two categories: reconstruction learning-based and contrast learning-based. Reconstruction learning methods measure anomalies by calculating the reconstruction error of nodes, while contrast learning methods establish positive and negative instance pairs through node subgraph sampling strategies and use contrast learning to measure node anomalies. The core idea of these methods is based on the imbalance between the number of normal and abnormal nodes. The model tends to fit the majority normal pattern, but the fitting effect on abnormal patterns is limited. When the proportion of abnormal samples increases, the model will pay more attention to fitting abnormal patterns, resulting in a significant decline in anomaly detection performance.

[0006] To verify this conjecture, the present invention conducted experiments on two classic models (Dominant and CoLA) and observed the changes in model performance by adjusting the number of abnormal samples in the data. The results show that as the anomaly ratio increases, the detection performance of both models is significantly affected. Therefore, guiding the model to pay more attention to the fitting of normal patterns and neglect the fitting of abnormal patterns has become a feasible direction for improving the anomaly detection performance of the model. Based on this, the present invention proposes an attribute network anomaly detection method based on reconstruction bias learning. This method can significantly improve the performance of attribute network anomaly detection without significantly increasing the computational complexity. Summary of the invention

[0007] The present invention proposes an attribute network anomaly detection method based on reconstruction bias learning. The method includes a graph reconstruction module, a reconstruction bias dynamic adjustment module, an anomaly enhancement classification module, and an anomaly score calculation module. The performance of attribute network anomaly detection is significantly improved without significantly increasing the computational complexity.

[0008] The present invention is implemented through the following technical solutions: First, the graph reconstruction module adopts a graph autoencoder model to learn the potential pattern of the graph data by minimizing the reconstruction error, and uses the degree of difference between the node reconstruction information and the original information to measure its abnormality. Secondly, the reconstruction bias dynamic adjustment module continuously interacts with the graph reconstruction module during the iterative optimization process, and dynamically modifies the loss function penalty coefficient to force the graph reconstruction module to fit the normal mode. In addition, the anomaly enhancement classification module uses the pseudo-normal node set and pseudo-abnormal node set finally screened out by the former as training samples, and uses the classification results to highlight the anomalies. Finally, the anomaly score calculation module combines the anomaly scores of the graph reconstruction module and the anomaly enhancement classification module to calculate the final anomaly score of each node.

[0009] In the graph reconstruction module, in order to effectively reconstruct the graph node information, the present invention takes the graph autoencoder as the basic model. The proposed framework can effectively capture the complex relationships in the graph and achieve considerable anomaly detection effect.

[0010] The shared encoder uses the information aggregation capability of the common graph convolutional network (GCN) to extract the attributes and structural features of the nodes, thereby achieving a higher-order integrated representation. The formula for each graph convolution layer is as follows:

[0011] in, It is The latent representation of the layer input, Represented in the attribute network Add the adjacency matrix after self-connection, and . It is The weight matrix of the layer, is a non - linear activation function.

[0012] The present invention makes Then, the weight - sharing encoder formula based on GCN can be simplified as:

[0013] Among them, the attribute matrix is regarded as the original input feature: while , , respectively represent the node embedding matrix, the normalized adjacency matrix, and the weight matrix of the encoder.

[0014] The structure reconstruction decoder uses the high - order representation extracted by the encoder to recover the structural information of the nodes for calculating the reconstruction error. Similar to the encoder, the decoder also adopts a graph convolutional network (GCN), which is expressed as follows:

[0015] Among them, , , respectively represent the low - dimensional latent representation matrix after the decoder, the weight matrix of the structure reconstruction decoder, and the reconstructed adjacency matrix.

[0016] The attribute reconstruction decoder reconstructs the original attributes of the nodes to reflect the degree of their attribute anomalies. Similarly, the attribute reconstruction decoder can be expressed as:

[0017] Among them, , respectively represent the reconstructed attribute matrix and the weight matrix of the attribute reconstruction decoder.

[0018] So far, the present invention has obtained the reconstruction results of the structures and attributes of all nodes. The overall reconstruction error can be defined as:

[0019] Among them, is an important control parameter used to balance the influences of structure reconstruction and attribute reconstruction.

[0020] An anomaly detection method for attribute networks based on reconstruction bias learning according to claim 1, characterized in that the reconstruction bias dynamic adjustment module can interact dynamically with the graph reconstruction module, and during the training process, the model is biased to pay attention to the normal patterns with a larger proportion in the graph data and ignore the abnormal patterns with a smaller proportion, so as to achieve better anomaly pattern detection performance. The dynamic adjustment is divided into two stages: pseudo-normal node dynamic search and model adaptability fine-tuning. Previously, the graph reconstruction anomaly score was composed of a structure reconstruction score and an attribute reconstruction score, and the anomaly scoring formula for each node was expressed as:

[0021] Nodes with higher reconstruction scores have a higher probability of being abnormal nodes because their patterns are difficult to conform to the normal patterns learned by the graph autoencoder.

[0022] The first stage of the dynamic adjustment is the pseudo-normal node dynamic search. By ranking the node anomaly scores, a sufficient number of pseudo-normal nodes with high normal confidence (i.e., low anomaly scores) are selected, and then in the second stage, the model is forced to pay attention to the reconstruction of the pseudo-normal nodes during training. Since the performance of the model is weak in the early stage, only the nodes with relatively low anomaly scores have considerable confidence. Therefore, the number of selected nodes should start from a small value. As the number of model training rounds increases, the number of selected nodes also gradually increases. At the same time, considering that normal nodes account for the vast majority of all nodes, the number of selected nodes should end with a large value. Thus, in the second stage, the model is more likely to reconstruct normal patterns. The dynamic selection number function is expressed as follows:

[0023] where, , , , respectively represent the maximum proportion of all nodes finally expected to be selected, the total number of graph nodes, the current training round, and the total number of training rounds in this stage. Due to the monotonically increasing property of this function, as the model is iteratively updated, the number of pseudo-normal nodes selected at training round increases from 0; at round , the number of selected nodes is equal to . The value of should not be too large to prevent too many abnormal nodes from being selected, resulting in the opposite effect. At training round , first calculate the anomaly score

[0024]

[0025] for each node, and then rank all the scores in ascending order: is The positive order node index list of the round, the pseudo normal node set of the round is expressed as:

[0026]

[0027] Therefore, this round The reconstruction loss function will be adjusted to:

[0028] in, is the Hadamard product, and is the penalty coefficient matrix, which is defined as:

[0029] in, is the penalty coefficient. The model will be updated according to the new loss function in this round of iteration Back propagation is performed to emphasize the reconstruction of the selected pseudo-normal nodes. During the training process, the reconstruction bias dynamic adjustment module and the graph reconstruction module interact continuously. The current performance of the model determines the screening of pseudo-normal nodes. The improvement of the next round of model performance after iteration depends on the biased reconstruction of the current pseudo-normal nodes. Through the dynamic interaction between the two, the model can achieve self-optimization to achieve better results. In order to give the model basic detection capabilities, the adjustment module starts to adjust the loss function after several rounds of iterations.

[0030] The second stage of dynamic adjustment is model adaptive fine-tuning. After rounds of pseudo-normal node search training process, the number of screening reaches the predetermined maximum value, through Get the final pseudo normal node set ,in At this point, the model has good detection performance. Compared with the early stage of training, the anomaly confidence of nodes with high anomaly scores has been significantly improved. Next, the set of pseudo-anomaly nodes with high anomaly scores is screened in the same way. ,in .Notice, And it is necessary to choose a smaller value, such as 0.05, in order to ensure the quality of the pseudo-abnormal node set. The model will be run on these two sets. For each round of adaptive fine-tuning training, the penalty coefficient matrix will be fixed as:

[0031] in ,forces the model to underestimate the reconstruction of pseudo-anomalous nodes.

[0032] An anomaly detection method for attribute networks based on reconstruction-biased learning according to claim 1, characterized in that the anomaly enhancement classification module uses two node sets ([ and ) given by the reconstruction bias dynamic adjustment module as negative samples and positive samples respectively to train a binary classifier, and uses the classification probability to quantify the node anomaly confidence, so as to enhance the anomaly score of the anomaly node and be more distinguishable compared with the normal node. Specifically, considering that the high-order embeddings learned by the shared encoder of the graph reconstruction model not only fuse the structural information and attribute information of the nodes, but also have normal-anomaly distinguishability to a certain extent, the present invention freezes the encoder that has completed training, encodes the classification training set to obtain high-order embeddings, and then replaces the original features for classification training. A general multi-layer perceptron (MLP) is used, and its simplified formula is as follows:

[0033] Wherein, , , respectively represent the classification probability matrix, the logical sigmoid function, and the weight matrix of the multi-layer perceptron. \mathrm{P}[i:1] represents the probability that the node is classified as an anomaly. Similarly, \mathrm{P}[i:0] represents the probability of being classified as normal. The loss function of the classifier is given by the following formula:

[0034] Wherein, represents the labels 1 and 0 of the anomaly and normal nodes, {p}_{i}=\mathrm{P}[i:{y}_{i}] , and .

[0035] An anomaly detection method for attribute networks based on reconstruction-biased learning according to claim 1, characterized in that in the anomaly score calculation module, the reconstruction anomaly score of node is denoted as . After the anomaly enhancement classification module completes training, all nodes are classified to obtain the classification probability matrix , and the classification anomaly score of each node is defined as:

[0036] After determining the two-part anomaly scores of each node, the aggregator is used to complete the fusion operation. First, maxmin normalization is performed on each type of score to unify its size. Then, the aggregator sums the two scores according to the weights to obtain the final anomaly score of each node:

[0037] where \(\beta\in[0,1]\) is a trade-off parameter used to balance the importance between two scores.

[0038] According to the method for detecting anomalies in an attribute network based on reconstruction-biased learning described in claim 1, characterized in that the graph reconstruction module and the reconstruction-bias dynamic adjustment module, based on the designed reconstruction-bias dynamic adjustment strategy, force the model to focus on fitting the normal mode in a biased manner and reduce the influence of the abnormal mode. The anomaly enhancement classification module is used to further highlight the anomalies. The anomaly score calculation module evaluates the anomaly of each node by integrating different anomaly scores. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Overall flowchart of the solution of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the technical solutions, objectives, and advantages of the present invention clearer and more explicit, the present invention will be further described in detail below with reference to the accompanying drawings.

[0041] As Figure 1 shown, the present invention provides a method for detecting anomalies in an attribute network based on reconstruction-biased learning, which includes a graph reconstruction module, a reconstruction-bias dynamic adjustment module, an anomaly enhancement classification module, and an anomaly score calculation module.

[0042] As Figure 1 shown, the graph reconstruction module uses a graph autoencoder to minimize the reconstruction error. Among them, the weight-sharing encoder based on GCN is used to extract the attribute and structural features of the nodes. The attribute matrix is regarded as the original input feature: , while , , respectively represent the node embedding matrix, the normalized adjacency matrix, and the weight matrix of the encoder. The structure reconstruction decoder uses the high-order representation extracted by the encoder to restore the structural information of the nodes in order to calculate the reconstruction error. , , , respectively represent the low-dimensional latent representation matrix after the decoder, the weight matrix of the structure reconstruction decoder, and the reconstructed adjacency matrix. The attribute reconstruction decoder reconstructs the original attributes of the nodes, thereby reflecting the degree of their attribute anomalies. Among them, , respectively represent the reconstructed attribute matrix and the weight matrix of the attribute reconstruction decoder. Through the above node structure and the reconstruction results of the attributes, the overall reconstruction error is obtained as follows: . Among them, is an important control parameter used to balance the influence of structure reconstruction and attribute reconstruction.

[0043] As Figure 1 shown, the dynamic adjustment of the reconstruction bias dynamic adjustment module is divided into two stages: pseudo-normal node dynamic search and model adaptability fine-tuning. The graph reconstruction anomaly score is composed of the structure reconstruction score and the attribute reconstruction score. The anomaly scoring calculation formula for each node is expressed as:

[0044] Nodes with higher reconstruction scores have a higher probability of becoming abnormal nodes because their patterns are difficult to conform to the normal patterns learned by the graph autoencoder.

[0045] The first stage of the dynamic adjustment is the pseudo-normal node dynamic search. By ranking the node anomaly scores, a sufficient number of pseudo-normal nodes with high normal confidence (i.e., low anomaly scores) are selected, and then in the second stage, the model is forced to pay attention to the reconstruction of pseudo-normal nodes during training. Since the performance of the model is weak in the early stage, only nodes with relatively low anomaly scores have considerable confidence. Therefore, the screening quantity should start from a small value and gradually increase as the model is continuously trained. Considering that normal nodes account for the vast majority of all nodes, the screening quantity should end with a large value. In the second stage, the model can reconstruct the normal pattern as much as possible. Therefore, the dynamic screening quantity function is expressed as follows:

[0046] Among them, , , , respectively represent the maximum proportion of all nodes that are finally expected to be screened, the total number of graph nodes, the current training round, and the total number of training rounds in this stage. Due to the monotonically increasing property of this function, as the model is iteratively updated, the number of pseudo-normal nodes screened at training round increases from 0; at round , the screening quantity is equal to . The value of should not be too large to prevent too many abnormal nodes from being screened, resulting in the opposite effect. At training round , first calculate the anomaly score

[0047] in, yes The positive order node index list of the round, the pseudo normal node set of the round is expressed as:

[0048] Therefore, this round The reconstruction loss function will be adjusted to:

[0049] in, is the Hadamard product, and is the penalty coefficient matrix, which is defined as:

[0050] in, is the penalty coefficient. The model will be updated according to the new loss function in this round of iteration Back propagation is performed to emphasize the reconstruction of the selected pseudo-normal nodes. During the training process, the reconstruction bias dynamic adjustment module and the graph reconstruction module interact continuously. The current performance of the model determines the screening of pseudo-normal nodes. The improvement of the next round of model performance after iteration depends on the biased reconstruction of the current pseudo-normal nodes. Through the dynamic interaction between the two, the model can achieve self-optimization to achieve better results. In order to give the model basic detection capabilities, the adjustment module starts to adjust the loss function after several rounds of iterations.

[0051] The second stage of dynamic adjustment is model adaptive fine-tuning. After rounds of pseudo-normal node search training process, the number of screening reaches the predetermined maximum value, through Get the final pseudo normal node set ,in At this point, the model has good detection performance. Compared with the early stage of training, the anomaly confidence of nodes with high anomaly scores has been significantly improved. Next, the set of pseudo-anomaly nodes with high anomaly scores is screened out in the same way. ,in .Notice, And it is necessary to choose a smaller value, such as 0.05, in order to ensure the quality of the pseudo-abnormal node set. The model will be run on these two sets. For each round of adaptive fine-tuning training, the penalty coefficient matrix will be fixed as:

[0052] in ,forces the model to underestimate the reconstruction of pseudo-anomalous nodes.

[0053] As shown Figure 1 in the figure, the anomaly enhancement classification module uses the two node sets given by the reconstruction bias dynamic adjustment module ( and ) as negative samples and positive samples respectively to train a binary classifier, and uses the classification probability to quantify the node anomaly confidence, so as to enhance the anomaly score of the abnormal node, which is more distinguishable compared with the normal node. Specifically, considering that the high-order embedding learned by the shared encoder of the graph reconstruction model not only fuses the structural information and attribute information of the nodes, but also has normal-abnormal distinguishability to a certain extent, the present invention freezes the encoder that has completed training, encodes the classification training set to obtain the high-order embedding, and then replaces the original features for classification training. A general multi-layer perceptron (MLP) is used, and its simplified formula is as follows:

[0054] where , , represent the classification probability matrix, the logical sigmoid function, and the weight matrix of the multi-layer perceptron respectively. \mathrm{P}[i:1] represents the probability that the node is classified as abnormal. Similarly, \mathrm{P}[i:0] represents the probability of being classified as normal. The loss function of the classifier is given by the following formula:

[0055] where represent the labels 1 and 0 of abnormal and normal nodes, {p}_{i}=\mathrm{P}[i:{y}_{i}] , and .

[0056] As shown Figure 1 in the figure, in the anomaly score calculation module, the reconstruction anomaly score of node is denoted as . After the anomaly enhancement classification module completes training, all nodes are classified to obtain the classification probability matrix , and the classification anomaly score of each node is defined as:

[0057] After determining the two-part anomaly scores of each node, the aggregator is used to complete the fusion operation. First, maxmin normalization is performed on each type of score to unify its magnitude. Then, the aggregator sums the two scores according to the weights to obtain the final anomaly score of each node:

[0058] where \(\beta\in[0,1]\) is a trade-off parameter used to balance the importance between the two scores.

Claims

1. An attribute network anomaly detection method based on reconstruction bias learning, characterized in that, Based on the designed reconstruction bias dynamic adjustment strategy, the graph reconstruction module and the reconstruction bias dynamic adjustment module force the model to focus on fitting the normal mode in a biased manner and reduce the impact of the abnormal mode; The abnormal enhancement classification module is used to further highlight the abnormality; The abnormal score calculation module evaluates the abnormality of each node by integrating different abnormal scores; Specifically, it includes the following steps: S1: The graph reconstruction module adopts a graph autoencoder to learn the latent pattern of graph data by minimizing the reconstruction error, and measures its abnormality degree by the difference between the node reconstruction information and the original information; S2: The reconstruction bias dynamic adjustment module continuously interacts with the graph reconstruction module during the iterative optimization process, and dynamically modifies the loss function penalty coefficient to force the graph reconstruction module to bias towards fitting the normal mode; S3: The abnormal enhancement classification module uses the pseudo-normal node set and the pseudo-abnormal node set finally selected by the former as training samples, and calculates the abnormal score using the classification probability to strengthen the abnormal performance; S4: The abnormal score calculation module combines the abnormal scores of the graph reconstruction module and the abnormal enhancement classification module to calculate the final abnormal score of each node, so as to achieve the purpose of anomaly detection.

2. The attribute network anomaly detection method based on reconstruction bias learning according to claim 1, wherein In order to effectively reconstruct the graph node information, the present invention is based on a graph autoencoder as the basic model to effectively capture the complex relationships in the graph. The step S1 includes: S1-1: The shared encoder uses the information aggregation ability of the common graph convolutional network (GCN) to extract the attribute and structural features of the nodes, so as to achieve a higher-order integrated representation. The formula of each graph convolutional layer is as follows: ; Among them, is the latent representation of the -th layer input, represents the adjacency matrix after adding self-connections in the attribute network , and , is the weight matrix of the -th layer, is a non-linear activation function; The present invention makes , then the weight-sharing encoder formula based on GCN can be simplified as: ; Among them, the attribute matrix is regarded as the original input feature: , while , , represent the node embedding matrix, the normalized adjacency matrix, and the weight matrix of the encoder, respectively; S1-2: The structure reconstruction decoder uses the higher-order representation extracted by the encoder to restore the structural information of the nodes for calculating the reconstruction error; similar to the encoder, the decoder also adopts the graph convolutional network (GCN), and its representation is as follows: ; ; Among them, , , respectively represent the low-dimensional latent representation matrix after the decoder, the weight matrix of the structure reconstruction decoder, and the reconstructed adjacency matrix; S1-3: The attribute reconstruction decoder reconstructs the original attributes of the nodes to reflect the degree of their attribute abnormality; similarly, the attribute reconstruction decoder can be expressed as: ; Among them, , respectively represent the reconstructed attribute matrix and the weight matrix of the attribute reconstruction decoder; S1-4: So far, the present invention has obtained the reconstruction results of the structures and attributes of all nodes, and the overall reconstruction error can be defined as: ; Among them, is an important control parameter for balancing the effects of structural reconstruction and property reconstruction.

3. The method for detecting anomalies in an attribute network based on reconstruction bias learning according to claim 1, wherein The reconstruction bias dynamic adjustment module can dynamically interact with the graph reconstruction module, and bias the model to pay attention to the normal mode with a larger proportion in the graph data and ignore the abnormal mode with a smaller proportion during the training process to achieve better anomaly detection performance. The step S2 includes: S2-1: The dynamic adjustment is divided into two stages: pseudo-normal node dynamic search and model adaptive fine-tuning; previously, the graph reconstruction abnormal score is composed of the structure reconstruction score and the attribute reconstruction score, and the abnormal score calculation formula for each node is expressed as: ; Nodes with higher reconstruction scores have a higher probability of being abnormal nodes because their patterns are difficult to conform to the normal patterns learned by the graph autoencoder; S2-2: The first stage of dynamic adjustment is the dynamic search for pseudo-normal nodes. By ranking the node anomaly scores, a sufficient number of pseudo-normal nodes with relatively high normal confidence (i.e., low anomaly scores) are selected. Then, in the second stage, the model is forced to pay attention to the reconstruction of pseudo-normal nodes during training. Since the model's performance is weak in the early stage, only nodes with relatively low ranks in the anomaly score have considerable confidence. Therefore, the screening quantity should start from a small value and gradually increase as the number of model training rounds increases. Considering that normal nodes account for the vast majority of all nodes, the screening quantity should end with a large value. Thus, it is more likely for the model to reconstruct the normal pattern in the second stage. The dynamic screening quantity function is expressed as follows: ; in, , , , They represent the maximum proportion of all nodes that are expected to be screened, the total number of graph nodes, the current training round, and the total training rounds in this stage. Due to the monotonically increasing nature of this function, as the model is iterated, the training rounds The number of pseudo-normal nodes to be screened when Starts from 0 and increases in rounds When filtering the number equal ; The value of should not be too large to prevent screening out too many abnormal nodes, which would have a counterproductive effect. When , first calculate the abnormal score of each node , and then rank all the ratings in positive order: ; Among them, is the list of positive-order node indices of the round, and the pseudo-normal node set of this round is expressed as: ; Therefore, the reconstruction loss function at this round will be adjusted to: ; wherein, is the Hadamard product, and is the penalty coefficient matrix, which is defined as: ; Among them, is the penalty coefficient. When the model is updated in this round of iteration, it will perform backpropagation according to the new loss function to emphasize the reconstruction of the selected pseudo-normal nodes; during the training process, the reconstruction bias dynamic adjustment module and the graph reconstruction module continuously interact. The current performance of the model determines the screening of pseudo-normal nodes, and the improvement of the model performance in the next round after iteration depends on the bias reconstruction of the current pseudo-normal nodes; through the dynamic interaction of the two, the model can achieve self-optimization, so as to achieve better results; in order to enable the model to have basic detection capabilities, the adjustment module starts to adjust the loss function after several rounds of iteration; S2-3: The second stage of dynamic adjustment is the adaptive fine-tuning of the model; after rounds of the pseudo-normal node search training process, the screening quantity reaches the predetermined maximum value, and the final pseudo-normal node set is obtained through , where ; at this time, the model already has good detection performance; compared with the early stage of training, the anomaly confidence of nodes with high anomaly scores has been significantly improved; then, in the same way, a pseudo-abnormal node set with relatively high anomaly scores is screened out, where ; note that and a relatively small value, such as 0.05, needs to be selected, aiming to ensure the quality of the pseudo-abnormal node set; the model will perform rounds of adaptive fine-tuning training on these two sets, and the penalty coefficient matrix will be fixed as: ; in , forcing the model to underestimate the reconstruction of pseudo-anomalous nodes.

4. A method for attribute network anomaly detection based on reconstruction bias learning according to claim 1, characterized in that The step S3 includes: S3-1: The anomaly enhancement classification module uses the two node sets given by the reconstruction bias dynamic adjustment module ( and ) as negative and positive samples respectively to train a binary classifier, and quantifies the node anomaly confidence with classification probabilities, thereby enhancing the anomaly scores of anomaly nodes, which are more distinguishable compared with normal nodes; S3-2: Considering that the high-order embeddings learned by the shared encoder of the graph reconstruction model not only integrate the structural information and attribute information of nodes but also have a certain degree of distinguishability between normal and abnormal, the present invention freezes the encoder that has completed training, encodes the classification training set to obtain high-order embeddings, and then replaces the original features for classification training. A general multi-layer perceptron (MLP) is used, and its simplified formula is as follows: ; Among them, , , respectively represent the classification probability matrix, the logical sigmoid function, and the weight matrix of the multi-layer perceptron; represents the probability that a node is classified as abnormal; similarly, represents the probability of being classified as normal; the loss function of the classifier is given by the following formula: ; Among them, represent the labels 1 and 0 for abnormal and normal nodes, , and .

5. The attribute network anomaly detection method based on reconstruction bias learning according to claim 1, characterized in that In the abnormal score calculation module, the node 's reconstruction abnormal score is denoted as ; after the abnormal enhancement classification module is trained, all nodes are classified to obtain the classification probability matrix , and the classification abnormal score of each node is defined as: ; After determining the two-part anomaly scores of each node, the aggregator is used to complete the fusion operation. First, each type of score is normalized by maxmin to unify its magnitude. Then, the aggregator sums the two scores according to the weights to obtain the final anomaly score of each node: ; Among them, is a trade-off parameter used to balance the importance between two scores.

Citation Information

Cited By

  • Clinical test data anomaly detection method and system based on machine learning

    CN120767001A

  • A machine learning-based clinical trial data anomaly detection method and system

    CN120767001B

  • Method for detecting node anomaly of reconstruction type graph based on priori-guided self-paced mask

    CN121881218A

  • A priori guided self-walking mask reconstruction type graph node anomaly detection method

    CN121881218B