Anomaly Detection Method for Complex Network Attack Scenarios
Through anomaly detection method for complex network attack scenarios, time segmentation processing and attribute heterogeneous subgraphs are used to construct feature snapshots, combined with HASGNet module and adaptive batch normalization technology, the existing detection methods solve the false alarm and missed alarm problems in the face of complex network attacks, and achieve more efficient and accurate abnormality detection.
Patent Information
- Application Number
- CN202510140774.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-08
AI Technical Summary
Existing host anomaly detection methods are difficult to detect effectively when facing complex network attack scenarios, especially long-term hidden attacks and new attack methods, resulting in high false alarm rates, high false alarm rates, and large computing resources.
An anomaly detection method for complex network attack scenarios is adopted. By dividing the original data set into a test set and a training set, using a strategy of processing data in time segments, a property heterogeneous subgraph and feature snapshot is constructed, and a HASGNet module is used for model training and verification, combining adaptive batch normalization and 3-hop depth neighbor sampling, a multi-head attention layer is set for feature aggregation.
It significantly improves the detection ability of zero-day attacks, can more accurately identify scattered anomalies hidden in a large number of normal activities, reduces the computing burden, and improves the robustness and efficiency of detection.
Smart Images

Figure CN119628961B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of host security anomaly detection, and particularly to an anomaly detection method for complex network attack scenarios. Background Art
[0002] Currently, security protection technologies such as firewalls and anomaly detection systems are constantly evolving. However, they still face huge detection pressure when dealing with some complex attack methods. For example, due to its long attack cycle, APT attacks are more concealed and complex, and can still bypass these defense measures. They lurk in the target system for a long time, collect sensitive data, and gradually expand the control range.
[0003] In traditional host anomaly detection methods, obvious deficiencies are shown when dealing with large-scale data sets. Some new attacks are often concealed and last for a long time, resulting in extremely large-scale logs or detection data generated, posing huge challenges to existing detection systems. Existing detection systems face problems of huge storage data volume and high computing resource occupancy when analyzing these long-term data. At the same time, various new attack methods have emerged, making the anomaly detection system relying on expert knowledge gradually ineffective.
[0004] Host anomaly detection can be roughly classified into the following categories according to the detection method. Signature-based detection methods generally rely on predefined attack patterns or signature databases to identify known attack behaviors. Their defects are reflected in the inability to identify unknown attacks, such as new attack means or zero-day attacks. Secondly, the signature library maintenance cost is relatively high. With the development of the times and technological progress, attack means are constantly changing, and the signature library needs to be continuously updated. Anomaly-based detection methods construct a normal model by learning the normal behavior patterns of the host. When the system detects operations significantly different from the normal behavior, an alarm will be triggered. Anomaly-based detection methods can detect unknown attacks, but due to changes in the normal mode of the system, the current anomaly detection methods have a high false alarm rate. In the face of some highly concealed attack means such as APT attacks, there are only a small number of abnormal behaviors among a large number of benign behaviors, and the distinction between abnormal data and benign data is not obvious, further affecting the detection effect. At the same time, the anomaly-based detection model requires a large amount of benign data support for training, and the calculation scale also increases significantly. Misuse-based detection methods detect operations that do not conform to the normal system behavior by predefined a set of specific rules or policies. Since specific rules need to be predefined in advance, it is difficult for this detection method to cover all situations. Moreover, manual writing of rules requires detailed knowledge. For complex attack chains such as APT attacks, it is difficult for us to capture the complete abnormal behavior. If the attacker's attack means change or bypass the existing rules, the misuse-based system may not work properly.
[0005] Recent anomaly detection research has begun to focus on methods that detect threats in hosts by leveraging the rich context information in data sources. Combining richer context information, compared with the original audit-data-based anomaly detection, richer data sources help us better separate anomalies from benign information. However, current such methods have significantly insufficient detection performance when facing large data volumes and stealthy attacks, and there are a large number of missed detections and false positives. Summary of the Invention
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] An anomaly detection method for complex network attack scenarios, comprising the following steps:
[0008] Step 1, divide the content of the original data set into a test set and a training set. Here, the training set only contains benign data, and the test set is divided into two categories: a test set with all benign data and a test set with all abnormal data.
[0009] Step 2, use a strategy of processing data in time segments to divide the source graph into several subgraphs, and each subgraph corresponds to a specific time interval.
[0010] Step 3, construct an attribute heterogeneous subgraph for each subgraph constructed in Step 2, hereinafter simply referred to as a feature snapshot, and each node will have a feature array , which is used to store the features accumulated through the node and its adjacent edges.
[0011] Step 4, perform model training on the feature snapshots processed in Step 3. For each feature snapshot, initialize a new HASGNet module, and at the same time freeze all old HASGNet modules, establish a horizontal connection between the new and old modules, and share parameters. If there is no old module, the freezing and connection operations are not performed. Use the adaptive batch normalization method to normalize the input data, use 3-hop depth neighbor sampling, set 8-head attention layers to assign different weights to aggregate features for node neighbors, use the cross-entropy loss function to calculate the loss, and update the model parameters through backpropagation.
[0012] Step 5, use the cumulative model generated by training in Step 4 to verify each node in the current feature snapshot to be trained, and at the same time set a trust threshold. When the ratio of the first prediction probability to the second prediction probability is greater than the threshold, then remove the node, otherwise consider it a classification error and retain the node. Finally, the remaining nodes are used as a new feature snapshot to execute Step 4;
[0013] Step 6, in the test phase, test the test data set respectively and calculate the corresponding metrics.
[0014] As a preferred embodiment of the present invention, in step 1, the content of the original data set is divided into a test set and a training set, where the training set only contains benign data, and the test set is divided into two categories: a test set with all benign data and a test set with all abnormal data.
[0015] As a preferred embodiment of the present invention, in step 2, the strategy of processing data by time segments is adopted. Assuming that G is the overall source graph, after division, it is as shown in formula (1), where each sub-graph corresponds to the corresponding time period ;
[0016] (1)
[0017] where G represents the entire set of sub-graphs, represents the stage sub-graph.
[0018] As a preferred embodiment of the present invention, in step 3, the method for constructing the feature array is as follows: for node , two edge type feature vectors are defined: , , which respectively represent the out-edge feature vector and the in-edge feature vector of node v; for the feature update of a node v: for each edge connected to this node, the symbol is used to represent the type of the edge. If the edge is operation and is an out-edge and points from node to other nodes, then this operation will affect the out-edge feature vector , and vice versa for the in-edge, that is, it points from other nodes to node , then this operation will affect the in-edge feature vector . Therefore, for each node , its edge feature update is represented by the following formulas (2) and (3):
[0019] (2)
[0020] (3)
[0021] where represents the out-edge feature of this node, represents the in-edge feature of this node; represents the type of the edge, represents the out-edge node sampled from this node, represents the in-edge node sampled from this node; and for the update of the node type feature, two node type features of the two nodes are defined , statistical updates will be performed through other node types connected to this node. At the same time, considering the in-edge and out-edge connection situations separately, other nodes connected to this node are classified into out-edge nodes and in-edge nodes. The specific updates result in formulas (4) and (5):
[0022] (4)
[0023] (5)
[0024] Among them represents the node type; represents the out-edge node feature, represents the in-edge node feature; The final feature representation of the node will concatenate the above four features, as shown in formula (6): (6), where represents the total feature of the current node.
[0025] As a preferred solution of the present invention, in step 41, a new HASGNet module will be initialized for each feature snapshot. At the same time, all old HASGNet modules will be frozen, and horizontal connections will be established between the new and old modules to share parameters. If there are no old modules, the freezing and connection operations will not be performed;
[0026] In step 42, a single HASGNet module is a graph attention neural network combined with a neighbor sampling strategy. It adopts a mini-batch sampling strategy and realizes sampling and calculation by only focusing on some nodes and their local neighbors; Specifically, use to perform multi-hop sampling on the neighbor nodes of the graph; For a given node , sample a subset ; By implementing multi-hop propagation to sample neighbors, higher-order dependencies can be captured; By setting , it means performing 3-hop neighbor sampling. The feature aggregation process of feature k in each layer is expressed as formula (7): (7); where represents the updated feature of node in the k-th layer, is the weight matrix to be learned, is the activation function, is the weight calculated by the attention mechanism, is the feature of neighbor node ;
[0027] In step 43, a variant of the graph attention network is adopted to achieve feature aggregation. For each node , the feature aggregation calculation is as shown in formula (8) (8); where is the attention coefficient between node and node , is the current node update feature, W is a learnable transformation matrix, and the attention coefficient is calculated by formula (9): (9); where is an activation function, represents the attention vector, , respectively represent the embedding vectors of nodes and ;
[0028] Step 44: The custom batch normalization layer dynamically updates the mean and variance. During the normalization process, when calculating the mean and variance of each batch, an adaptive learning rate is added, and the update strategy is dynamically adjusted in combination with the characteristics of the current batch of data, described by formula (10):
[0029] (10); where is node ; , are the mean and variance of the batch respectively, and are learnable scaling and offset parameters, is a small constant to prevent the denominator from being zero; during training, by dynamically adjusting the mean and variance, an adaptive factor is introduced; as shown in formula (11), the global statistic update formula of the model is as shown in (12) and (13): (11);
[0030] (12)
[0031] (13)
[0032] where represents the activation function, represents the mean function, is the global running mean, is the global running variance;
[0033] Step 45: In the inference stage, the model no longer calculates the mean and variance of the current batch, but uses the and accumulated during training, as shown in formula (14):
[0034] (14).
[0035] As a preferred embodiment of the present invention, for the trust threshold set in step 5, when the ratio of the first prediction probability to the second prediction probability is greater than the threshold, the node is removed; otherwise, it is considered a classification error and the node is retained. Finally, the remaining nodes are used as a new snapshot to continue executing step 4. The judgment of the trust threshold is as shown in formula (15): (15), where represents the set threshold, represents the first prediction probability, and the second prediction probability.
[0036] As a preferred embodiment of the present invention, in step 6, the test work needs to be carried out separately for the benign data test set and the abnormal data test set, recording TN, TP, FN, FP, and finally calculating the relevant evaluation indicators.
[0037] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0038] When training the model, the present invention uses unsupervised learning, which can automatically learn normal behavior patterns without relying on known attack features or signatures, significantly improving the detection ability for zero-day attacks.
[0039] With the help of the heterogeneous attention mechanism, by assigning different attention weights to the key nodes and edges in the graph structure, the system can more accurately focus on potential abnormal activity areas and effectively identify scattered abnormal events hidden in a large number of normal activities.
[0040] The present invention designs an efficient segmented attribute heterogeneous graph processing and feature snapshot generation strategy, focusing on smaller time-segmented subgraphs, reducing the computational burden of global graph analysis.
[0041] 4. When detecting anomalies in different data sets, the present invention has strong robustness. By using progressive network training and reusing the features learned by the previous model, the robustness of the model is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the overall flowchart of the detection method of the embodiment of the present invention.
[0043] Figure 2 is the overall network architecture diagram of the HASGNet model of the embodiment of the present invention.
[0044] Figure 3 is the segmented source graph processing strategy diagram of the embodiment of the present invention.
[0045] Figure 4 is the single-node feature construction flowchart of the embodiment of the present invention.
[0046] Figure 5 It is a comparison chart of the F1 experimental value results between the present invention and several other benchmark models under the SC dataset.
[0047] Figure 6 It is a comparison chart of the F1 experimental value results between the present invention and several other benchmark models under the E3 dataset. Detailed implementation manners
[0048] The present invention will be further clarified below in conjunction with the accompanying drawings and specific implementation manners. It should be understood that the following specific implementation manners are only used to illustrate the present invention and not to limit the scope of the present invention. It should be noted that the terms "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to the directions in the accompanying drawings, and the terms "inner" and "outer" respectively refer to the directions towards or away from the geometric center of a specific component.
[0049] As Figure 1 shown, an anomaly detection method for complex network attack scenarios proposed in this embodiment is executed as follows in practical applications:
[0050] Step 1: Divide the content of the original dataset into a test set and a training set. Here, the training set only contains benign data, and the test set is divided into two categories: a test set with all benign data and a test set with all abnormal data.
[0051] Step 2: Use the strategy of processing data in time segments to divide the source graph into several subgraphs. Each subgraph corresponds to a specific time interval. Note that the time intervals here are not necessarily exactly equal. To control the consistent node scale of the subgraphs, the upper limit value N of the subgraph nodes will be set. When the number of collected nodes reaches the upper limit, the start timestamp will be recorded, thereby reducing the analysis scale each time. The overall process is as Figure 3 . Assuming that G is the overall source graph, after division, it is as shown in formula (1):
[0052] (1)
[0053] where each subgraph corresponds to a specific time period .
[0054] Construct an attribute heterogeneous subgraph for each subgraph constructed in Step 2, which will be simply referred to as a feature snapshot below. The simple construction process of the feature array of a single node is as Figure 4 , and each node will have a feature array , used to store the features accumulated through the node and its adjacent edges. The method for constructing the feature array is for a node , define two feature vectors: , , respectively representing the out-edge feature vector and the in-edge feature vector of node v; for the feature update of a node v: for each edge connected to this node , use the symbol to represent the type of the edge. If the edge is operation, and it is an out-edge, and it points from node to other nodes, then this operation will affect the out-edge feature vector , and vice versa for the in-edge, that is, it points from other nodes to node , then this operation will affect the in-edge feature vector . Therefore, for each node , its edge feature update is represented by the following formulas (2) and (3):
[0055] (2)
[0056] (3)
[0057] where represents the out-edge feature of this node, represents the in-edge feature of this node; represents the type of the edge, represents sampling the out-edge nodes of this node, represents sampling the in-edge nodes of this node; for the update of the node type feature, define the node type features of two nodes, which will be statistically updated through the other node types connected to this node. At the same time, considering the in-out edge connection situations separately, classify the other nodes connected to this node into out-edge nodes and in-edge nodes. The specific updates result in formulas (4) and (5):
[0058] (4)
[0059] (5)
[0060] where represents the node type; represents the out-edge node feature, represents the in-edge node feature; the final feature representation of the node will concatenate the above four features, as shown in formula (6): (6), where represents the total feature of the current node.
[0061] Step 4 performs model training on the feature snapshots processed in Step 3. For each feature snapshot, a new HASGNet module will be initialized. At the same time, all old HASGNet modules will be frozen, and a horizontal connection will be established between the new and old modules to share parameters. If there are no old modules, the freezing and connection operations will not be performed. The adaptive batch normalization method is used to normalize the input data. 3-hop depth neighbor sampling is used, and an 8-head attention layer is set to assign different weights to node neighbors to aggregate features. The cross-entropy loss function is used to calculate the loss, and the model parameters are updated by backpropagation. The model construction is as shown in Figure 2 .
[0062] In Step 41, for each feature snapshot, a new HASGNet module will be initialized. At the same time, all old HASGNet modules will be frozen, and a horizontal connection will be established between the new and old modules to share parameters. If there are no old modules, the freezing and connection operations will not be performed;
[0063] In Step 42, a single HASGNet module is a graph attention neural network combined with a neighbor sampling strategy. The mini-batch sampling strategy is adopted, and sampling and calculation are achieved by only focusing on some nodes and their local neighbors. Specifically, is used to perform multi-hop sampling on the neighbor nodes of the graph; for a given node , a subset is sampled; by implementing multi-hop propagation to sample neighbors, higher-order dependencies can be captured; by setting , it means performing 3-hop neighbor sampling. The feature aggregation process of feature k in each layer is expressed by formula (7): (7) where represents the updated feature of node in the k-th layer, is the weight matrix to be learned, is the activation function, is the weight calculated by the attention mechanism, is the feature of neighbor node ;
[0064] Step 43 adopts a variant of the graph attention network to achieve feature aggregation. For each node , the feature aggregation calculation is shown in formula (8) (8); where is the attention coefficient between node and node , is the current node update feature, and W is the learnable transformation matrix. The attention coefficient is calculated by formula (9): (9); where is an activation function, represents the attention vector, 、 respectively represent the embedding vectors of nodes and ;
[0065] In step 44, the custom batch normalization layer dynamically updates the mean and variance. During the normalization process, when calculating the mean and variance of each batch, an adaptive learning rate is added, and when updating the mean and variance, the update strategy is dynamically adjusted in combination with the characteristics of the current batch of data, as described by formula (10):
[0066] (10); where is the node ; 、 are the mean and variance of the batch respectively, and are the learnable scaling and offset parameters, is a small constant to prevent the denominator from being zero; during training, by dynamically adjusting the mean and variance, an adaptive factor is introduced; as shown in formula (11), the global statistic update formula of the model is as shown in (12) and (13):
[0067] (11);
[0068] (12)
[0069] (13)
[0070] where represents the activation function, represents the mean function, is the global running mean, is the global running variance;
[0071] In step 45, during the inference phase, the model no longer calculates the mean and variance of the current batch, but uses the and accumulated during training, as shown in formula (14):
[0072] (14).
[0073] Step 5: Use the cumulative model generated in Step 4 to verify each node in the current feature snapshot to be trained, and set a confidence threshold. The confidence threshold judgment formula is as shown in (15). When the ratio of the first prediction probability to the second prediction probability is greater than the threshold, remove the node; otherwise, consider it a classification error and retain the node. Finally, the remaining nodes are used as the new feature snapshot to continue Step 4.
[0074] The confidence threshold judgment is as shown in formula (15): (15), where represents the set threshold, represents the first prediction probability, the second prediction probability.
[0075] In Step 6, the testing work needs to be carried out separately for the benign data test set and the abnormal data test set, record TN, TP, FN, FP, and finally calculate the relevant evaluation metrics.
[0076] The datasets used in this example are the Unicorn dataset (two data subsets) and the DARPA E3 (three data subsets) dataset, which are generated in a controlled laboratory environment and simulate the typical network attack chain process. By capturing the detailed activity logs at the system level, it provides researchers with complex attack and system activity data in a real environment. The dataset contains the background noise of benign activities, ensuring the environmental consistency of benign and attack graphs, which is convenient for a comprehensive evaluation of the detection effect. Select the F1 value of anomaly detection as the performance indicator of the algorithm. Compared with the applicable methods in the traditional anomaly detection field, the F1 value of the present invention has been improved. The comparison graph with the traditional methods is as shown in Figure 5 、 Figure 6 shown. Figure 5 shows the bar chart of the F1 value comparison results of this example for the SC-1 subset and the SC-2 subset of the Unicorn dataset, corresponding to Dataset 1 and Dataset 2 respectively, compared with Comparative Method 1 (Unicorn), Comparative Method 2 (ProvDetector), Comparative Method 3 (Threatrace), Comparative Method 4 (Prov2Vec), and Comparative Method 5 (WLSubtree). Figure 6It shows the bar chart of the comparison results of the F1 values of the Trace subset, Cadets subset, and Fivedirections subset of this example under the DARPA E3 dataset, corresponding to Dataset 1, Dataset 2, and Dataset 3 respectively. The comparison is made with the F1 values of Comparative Method 1 (Threatrace), Comparative Method 2 (Log2Vec), and Comparative Method 3 (LogGAN). The F1 value is a harmonic mean that measures precision and recall, providing a single score that takes both into account. A high F1 score indicates a good balance between precision and recall, meaning the system is both accurate and comprehensive in anomaly detection. A low F1 score indicates an imbalance between precision and recall, meaning the system may miss many anomaly nodes (low recall) or mis-detect too many benign nodes (low precision).
[0077] In summary, the anomaly detection method designed in the present invention demonstrates significant performance advantages on multiple datasets. Whether it is the precision or recall of anomaly detection, it performs excellently, further proving its application potential and stability in anomaly detection in complex network attack scenarios.
[0078] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features.
Claims
1. An anomaly detection method for complex network attack scenarios, characterized in that: The steps include: Step 1: Divide the test data set into a test set and a training set; Step 2: Use the strategy of processing data by time segment to divide the source graph into several sub-graphs, each of which corresponds to a specific time interval; Step 3: construct attribute heterogeneous subgraphs for each subgraph constructed in step 2, which will be referred to as feature snapshots below. Each node will have a feature array , used to store the features accumulated through the nodes and their adjacent edges; The characteristic array construction method in step 3 is for nodes , define two edge type feature vectors: , , respectively represent the outgoing edge feature vector and incoming edge feature vector of node v; for the feature update of a node v: for each edge connected to the node , using the symbol To indicate the type of edge, if the edge yes Operation, and it is an outgoing edge, and from the node Pointing to other nodes, this operation will affect the outgoing edge feature vector , the opposite is an inbound edge, that is, it points from other nodes to the node , then this operation will affect the input edge feature vector , so for each node , its edge feature update is expressed as the following formulas (2) and (3): (2) (3) in Represents the outgoing edge feature of the node, Represents the incoming edge features of the node; Represents the type of edge, Represents sampling of the node's outgoing edge nodes. Represents sampling of the node's incoming edge node; and the update of the node type feature defines the node type features of the two nodes , the types of other nodes connected to the node are statistically updated, and the inbound and outbound edge connections are considered separately, and the other nodes connected to the node are classified into outbound nodes and inbound nodes. The specific update obtains formulas (4) and (5): (4) (5) in Represents the node type; represents the outgoing node feature, represents the feature of the incoming edge node; the final feature representation of the node will concatenate four features, as shown in formula (6): (6), where Indicates the total characteristics of the current node; Step 4: Perform model training on the feature snapshot processed in step 3. During the training process, a new HASGNet module is initialized for each feature snapshot, and all old HASGNet modules are frozen. A horizontal connection is established between the new and old modules to share parameters. If the old module does not exist, the freezing and connection operations are not performed. The input data is normalized using the adaptive batch normalization method, 3-hop deep neighbor sampling is used, and an 8-head attention layer is set to assign different weights to node neighbors to aggregate features. The cross entropy loss function is used to calculate the loss, and the model parameters are updated by back propagation. Step 5: Use the cumulative model generated by step 4 training to verify each node in the current feature snapshot to be trained, and set a credible threshold. When the ratio of the first predicted probability to the second predicted probability is greater than the threshold, the node is removed. Otherwise, it is considered to be misclassified and the node is retained. Finally, the remaining nodes are used as new feature snapshots to execute step 4. Step 6: Testing phase: test the test sets separately and calculate the corresponding indicators.
2. According to claim 1, the anomaly detection method for complex network attack scenarios is characterized in that: In step 1, the test data set content is divided into a test set and a training set, wherein the training set contains only benign data, and the test set is divided into two categories, a test set consisting entirely of benign data and a test set consisting entirely of abnormal data.
3. The anomaly detection method for complex network attack scenarios according to claim 1 is characterized in that: The strategy of processing data by time segment in step 2 is as follows: assuming that G is the entire subgraph set, the division is as shown in formula (1), where each subgraph Corresponding time period ; (1) Where G represents the entire subgraph set, Represents a subgraph.
4. The anomaly detection method for complex network attack scenarios according to claim 1 is characterized in that: The step 4 includes the following: Step 41 initializes a new HASGNet module for each feature snapshot, freezes all old HASGNet modules, establishes lateral connections between new and old modules, and shares parameters. If there is no old module, the freezing and connection operations are not performed. The single HASGNet module in step 42 is a graph attention neural network combined with a neighbor sampling strategy, which adopts a mini-batch sampling strategy to achieve sampling and calculation by focusing only on some nodes and their local neighbors; specifically, using Multi-hop sampling of neighbor nodes in the graph; for a given node , sample a subset ; is the set of all neighbor nodes of node v; yes A subset of the set; by implementing multi-hop propagation sampling neighbors to capture higher-order dependencies; by setting , which means performing 3-hop neighbor sampling. The feature aggregation process of feature k at each layer is expressed as formula (7): (7), where Represents the nodes in the kth layer The updated features of is the weight matrix to be learned, is the activation function, is the weight calculated by the attention mechanism, Neighbor node Features; Step 43 uses a variant of the graph attention network to achieve feature aggregation. For each node , the feature aggregation calculation is shown in formula (8) (8); among which For Node and nodes The attention coefficient between Update the features for the current node, and W is the learnable transformation matrix, the attention coefficient By calculating formula (9), we know: (9); among which is an activation function, represents the attention vector, , Represents nodes and The embedding vector of Step 44 The custom batch normalization layer dynamically updates the mean and variance. When calculating the mean and variance of each batch, the normalization process adds an adaptive learning rate. When updating the mean and variance, the update strategy is dynamically adjusted based on the characteristics of the current batch data, which is described as formula (10): (10), , are the mean and variance of the batch, respectively. and are learnable scaling and offset parameters, To prevent a small constant with a denominator of 0; During the training process, the adaptive factor is introduced by dynamically adjusting the mean and variance. ; As in formula (11), the global statistics update formulas of the model are as follows (12) and (13): (11); (12) (13), in represents the activation function, represents the mean function, is the global runtime mean, is the global runtime variance; Step 45 During the inference phase, the model no longer calculates the mean and variance of the current batch, but uses the accumulated mean and variance during training. and , as shown in formula (14): (14)。 5. The anomaly detection method for complex network attack scenarios according to claim 1 is characterized in that: The trust threshold set in step 5 is that when the ratio of the first predicted probability to the second predicted probability is greater than the threshold, the node is removed; otherwise, it is considered that the classification is wrong and the node is retained. Finally, the remaining node is used as a new snapshot to continue to execute step 4; The credible threshold is determined as follows: (15), where Represents the set threshold value, represents the first predicted probability, The second predicted probability.
6. The anomaly detection method for complex network attack scenarios according to claim 1 is characterized in that: In step 6, tests should be carried out on the benign data test set and the abnormal data test set respectively, and TN, TP, FN, FP should be recorded, and finally the relevant evaluation indicators should be calculated.
Citation Information
Patent Citations
Attack detection method based on knowledge graph recommendation model
CN117596044A
APT attack detection method based on traceability graph behavior information
CN119232465A