Anomaly detection method and device based on multi-period balanced attribute subgraph and heterogeneous graph attention
By adopting the anomaly detection method of multi-period balanced attribute subgraphs and diffusion-enhanced heterogeneous graph attention network layer in IDS, the data imbalance and classification uncertainty problems in traditional IDS are solved, and the accuracy and adaptability of network traffic anomaly detection are improved.
Patent Information
- Application Number
- CN202411244042.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-09-05
AI Technical Summary
Traditional IDS has data imbalance and classification uncertainty problems during construction, resulting in insufficient identification and in timely processing when detecting network traffic anomalies.
Anomaly detection method based on multi-time balanced attribute subgraphs and heterogeneous graph attention is adopted, and the multi-time balanced attribute subgraphs and diffusion-enhanced heterogeneous graph attention network layer is constructed to process unbalanced data and characterize the spatiotemporal relationship in attack traffic.
The network anomaly detection model's ability to identify attack traffic is improved, the classification accuracy is improved, data imbalance and classification uncertainty are solved, and higher classification accuracy and stronger model adaptability are achieved.
Smart Images

Figure CN119182576B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to an anomaly detection method and device based on multi-period balanced attribute subgraphs and heterogeneous graph attention. Background Art
[0002] Network security has become a crucial topic in the field of network technology. Network traffic anomaly detection is an important technical means of network security situation awareness. Therefore, a corresponding intrusion detection system (IDS) for network traffic anomaly detection has emerged.
[0003] However, traditional IDSs are based on predefined rules or signatures to identify known attack patterns. Faced with the emerging and increasingly complex network attack methods, IDSs have problems with data imbalance and classification uncertainty when they are built, resulting in inaccurate identification and untimely processing when performing network traffic anomaly detection. Its limitations are gradually becoming apparent, which undoubtedly lays hidden dangers for the network security system and increases the risk of data leakage. It is therefore necessary to propose a new IDS that can solve the problems of data imbalance and classification uncertainty. Summary of the invention
[0004] In view of this, the present invention proposes an anomaly detection method and device based on multi-period balanced attribute subgraphs and heterogeneous graph attention to solve the problems of data imbalance and classification uncertainty in the construction of current IDS.
[0005] The technical solution of the present invention is achieved in this way:
[0006] According to a first aspect, an embodiment of the present invention provides an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention, the method comprising:
[0007] Acquire the network data to be detected, and input the network data to be detected into a trained network anomaly detection model to obtain the anomaly data output by the network anomaly detection model and the category of the anomaly data; the network anomaly detection model includes a diffusion enhanced heterogeneous graph attention network layer and a classification network layer, and the network anomaly detection model is obtained based on sample data, similarity features of the sample data and event embedding training;
[0008] The network anomaly detection model is trained by the following steps:
[0009] Acquire the sample data, and perform time-discrete processing on the sample data according to the time series correlation degree to obtain a number of time-discrete data groups within a similar time period; the time-discrete data groups contain a number of the sample data;
[0010] Determine the characteristic information of each sample data in each of the time discrete data groups, and construct a data attribute subgraph according to the characteristic information; the data attribute subgraph is composed of a plurality of meta-paths, and the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths;
[0011] According to the ratio of abnormal data to normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, the abnormal data and the normal data are proportionally balanced to obtain a balanced attribute subgraph corresponding to each data attribute subgraph;
[0012] Inputting the balanced attribute subgraph into the diffusion-enhanced heterogeneous graph attention network layer, and iteratively updating the network parameters of the diffusion-enhanced heterogeneous graph attention network layer based on the total loss function, to obtain the similarity features and the event embedding of the balanced attribute subgraph output by the diffusion-enhanced heterogeneous graph attention network layer; the total loss function is composed of a class distance loss function, a normalized loss function, and an attribute-level structure-aware loss function;
[0013] The similarity feature is used as input data and the event is embedded as label information used for training, which is input into the classification network layer and the network parameters of the classification network layer are iteratively updated based on a supervised machine learning method to obtain a network anomaly detection model.
[0014] In combination with the first aspect, in a first implementation of the first aspect, the proportion of abnormal data to normal data in the sample data of each data attribute subgraph, the sample data, and the neighbor data of the sample data is balanced to obtain a balanced attribute subgraph corresponding to each data attribute subgraph, specifically including:
[0015] Determine the number of abnormal data and normal data in the sample data of each data attribute subgraph respectively, and obtain a first number of abnormal data and a second number of normal data in each data attribute subgraph;
[0016] When it is determined that the ratio of the first number to the second number exceeds the preset ratio and the first number is greater than the second number, generating first incremental data based on the normal data and the neighboring nodes of the normal data, and adding the first incremental data to the normal data until the ratio of the first number to the second number does not exceed the preset ratio;
[0017] When it is determined that the ratio of the first number to the second number exceeds the preset ratio and the first number is less than the second number, second incremental data is generated based on the abnormal data and the neighboring nodes of the abnormal data, and the second incremental data is added to the abnormal data until the ratio of the first number to the second number does not exceed the preset ratio.
[0018] In combination with the first implementation of the first aspect, in the second implementation of the first aspect, the first incremental data is generated by the following steps:
[0019] Use the nearest neighbor algorithm to find a preset number of neighbor nodes for normal data;
[0020] One of the neighbor nodes is used as constructed sample data, and linear interpolation is performed according to the normal data and the constructed sample data to generate the first incremental data.
[0021] In combination with the first aspect, in a third implementation of the first aspect, the inputting the balanced attribute subgraph into the diffusion enhanced heterogeneous graph attention network layer, and iteratively updating the network parameters of the diffusion enhanced heterogeneous graph attention network layer based on the total loss function, to obtain the similarity feature of the balanced attribute subgraph output by the diffusion enhanced heterogeneous graph attention network layer and the event embedding, specifically includes:
[0022] Obtaining the attention weights between the attribute nodes in each meta-path in each data attribute subgraph;
[0023] According to the attention weight, obtaining the node feature of the attribute node;
[0024] Performing a preset number of node-level attention processing on the node features to obtain a path node embedding of the attribute node, and concatenating the path node embeddings of all attribute nodes in a meta-path to obtain a meta-path embedding of the meta-path;
[0025] weighting the meta-path embedding according to semantic-level attention to obtain the importance of the meta-path;
[0026] determining the meta-path weight of the meta-path according to the importance of the meta-path;
[0027] Obtaining an attribute node embedding of the attribute node according to all meta-paths through which the attribute node passes in the balanced attribute subgraph and the meta-path weights;
[0028] Obtaining the event embedding of the balanced attribute subgraph according to the attribute node embeddings of all the attribute nodes in each balanced attribute subgraph;
[0029] According to the correlation degree between the attribute node and the attribute node embeddings of the neighboring nodes of the attribute node, the similarity between the attribute node embeddings is obtained, and the similarity feature of the balanced attribute subgraph is determined according to the similarity.
[0030] In combination with the first aspect, in a fourth implementation of the first aspect, the method further includes the following steps:
[0031] The original data is obtained and preprocessed to obtain sample data.
[0032] In combination with the fourth implementation of the first aspect, in the fifth implementation of the first aspect, obtaining the original data and preprocessing the original data to obtain the sample data specifically includes:
[0033] Obtain the original data and clean up the duplicate and incomplete data in the original data;
[0034] Discretize the cleaned raw data;
[0035] The discretized original data is converted into binary form to obtain sample data.
[0036] In combination with the first aspect, in a sixth implementation of the first aspect, the expression of the total loss function is:
[0037] L=L1+L2+L3
[0038] Among them, L represents the total loss function; L1 represents the class distance loss function, L2 represents the normalization loss function, and L3 represents the attribute-level structure-aware loss function.
[0039] In combination with the sixth implementation of the first aspect, in the seventh implementation of the first aspect, the expression of the class distance loss function is:
[0040] L1=-||c1-c2||
[0041] Where c1 represents the geometric center of abnormal type events, and c2 represents the geometric center of normal type events;
[0042] The expression of the normalized loss function is:
[0043]
[0044] Among them, ε j represents the loss parameter, n represents the total number of attribute nodes in the balanced attribute subgraph;
[0045] The expression of the attribute-level structure-aware loss function is:
[0046]
[0047] Among them, V represents the set of all attribute values v in the data attribute subgraph, Y v represents the label corresponding to the node v, and C represents the parameters corresponding to the classifier.
[0048] In combination with the seventh implementation manner of the first aspect, in the eighth implementation manner of the first aspect, the expression of the geometric center is:
[0049]
[0050] Among them, c represents the geometric center of the event corresponding to the abnormal / normal type, N represents the total number of events corresponding to the abnormal / normal type, and B represents all events corresponding to the abnormal / normal type.
[0051] According to a second aspect, an embodiment of the present invention provides an anomaly detection device based on multi-period balanced attribute subgraphs and heterogeneous graph attention, the device comprising:
[0052] An anomaly detection module is used to obtain network data to be detected, and input the network data to be detected into a trained network anomaly detection model to obtain anomaly data output by the network anomaly detection model and the category of the anomaly data; the network anomaly detection model includes a diffusion enhanced heterogeneous graph attention network layer and a classification network layer, and the network anomaly detection model is obtained based on sample data, similarity features of the sample data and event embedding training;
[0053] The network anomaly detection model includes:
[0054] A time discrete module is used to obtain the sample data and perform time discrete processing on the sample data according to the time series correlation degree to obtain a number of time discrete data groups within a similar time period; the time discrete data group contains a number of the sample data;
[0055] A graph construction module, used to determine the characteristic information of each sample data in each of the time discrete data groups, and to construct a data attribute subgraph according to the characteristic information; the data attribute subgraph is composed of a plurality of meta-paths, and the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths;
[0056] A data balancing module, used to balance the abnormal data and the normal data according to the ratio of the abnormal data to the normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, so as to obtain a balanced attribute subgraph corresponding to each data attribute subgraph;
[0057] A feature learning module, used for inputting the balanced attribute subgraph into the diffusion enhanced heterogeneous graph attention network layer, and iteratively updating the network parameters of the diffusion enhanced heterogeneous graph attention network layer based on a total loss function, to obtain the similarity feature of the balanced attribute subgraph output by the diffusion enhanced heterogeneous graph attention network layer and the event embedding; the total loss function is composed of a class distance loss function, a normalization loss function and an attribute-level structure perception loss function;
[0058] The classification detection module is used to input the similarity feature as input data and the event embedding as label information used for training into the classification network layer and iteratively update the network parameters of the classification network layer based on a supervised machine learning method to obtain a network anomaly detection model.
[0059] Compared with the prior art, the anomaly detection method and device based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention have the following beneficial effects:
[0060] By inputting the network data to be detected into the trained network anomaly detection model, the anomaly data and the category of the anomaly data output by the network anomaly detection model are obtained. The network anomaly detection model processes unbalanced data by constructing balanced attribute subgraphs in multiple time periods, thereby avoiding overfitting and improving the network anomaly detection model's ability to identify attack traffic. At the same time, the network anomaly detection model is based on the diffusion enhanced heterogeneous graph attention network to characterize and reason about the spatiotemporal relationship of different fields in the attack traffic. It not only considers the relationship between different types of data, but also mines the relationship between data fields, making the difference between the features of abnormal data and normal data more obvious, thereby improving the classification accuracy. The network anomaly detection model solves the problems of data imbalance and classification uncertainty, achieves higher classification accuracy and stronger model adaptability, and can process large-scale network traffic data, improving the accuracy and robustness of the model, so that the trained network anomaly detection model can more effectively identify and defend against network attacks, and provide more powerful protection measures for network security. The network anomaly detection model has important application value in the field of network security. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0062] Figure 1 It is a flow chart of the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0063] Figure 2 It is a schematic diagram of the training process of the network anomaly detection model in the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0064] Figure 3 A schematic diagram of a balanced attribute subgraph in an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0065] Figure 4 It is a schematic diagram of a balanced attribute subgraph node meta-path in the anomaly detection method based on multi-period balanced attribute subgraph and heterogeneous graph attention of the present invention;
[0066] Figure 5 It is another flow chart of the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0067] Figure 6 It is another flow chart of the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0068] Figure 7 It is another flow chart of the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0069] Figure 8 It is another flow chart of the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0070] Fig. 9 It is a schematic diagram of the Diffusion-HAN layer optimization process in the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention;
[0071] Fig.10 The schematic diagram of the structure of the anomaly detection device based on multi-period balanced attribute subgraphs and heterogeneous graph attention provided by the present invention is shown; DETAILED DESCRIPTION
[0072] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] Network security has become a crucial topic in the field of network technology. Network traffic anomaly detection is an important technical means of network security situation awareness. Therefore, the corresponding intrusion detection system (IDS) has emerged. IDS is a network security device that monitors network transmission in real time and issues alarms or takes active response measures when suspicious transmission is found. It is a proactive security protection technology.
[0074] However, traditional intrusion detection systems (IDS) identify known attack patterns based on predefined rules or signatures. Faced with the ever-emerging and increasingly complex network attack methods, traditional IDS has the defects of inaccurate network anomaly identification and untimely processing. Its limitations are gradually emerging, which undoubtedly lays hidden dangers for the network security system and increases the risk of data leakage.
[0075] This is because there are problems of data imbalance and classification uncertainty when building IDS. Specifically, the data imbalance problem is mainly reflected in the uneven distribution of the number of samples of each category in the training data, which leads to the traditional IDS tending to pay too much attention to the category data with a large number and ignore the category data with a small number. This deviation directly affects the accuracy and efficiency of detecting abnormal behaviors. At the same time, the traditional IDS cannot effectively model the time and space relationship between different fields in the attack traffic, which makes it difficult to solve the data imbalance problem. That is, when there are fewer samples of certain attack types in the data set for building IDS, the IDS performs poorly in detecting these uncommon attack types, which seriously affects the accuracy of detection; the classification uncertainty problem stems from the diversity and complexity of network security threats. There are fuzzy boundaries between the spatial distribution of different categories or potential manifolds. It is difficult for traditional IDS to accurately characterize the differentiated characteristics between categories. This classification uncertainty problem makes it difficult for IDS to distinguish and determine the categories of network traffic or events, increasing the risk of false positives and false negatives, thereby reducing the overall performance of traditional IDS.
[0076] The anomaly detection method and device based on multi-period balanced attribute subgraphs and heterogeneous graph attention provided in this specification are intended to establish multi-period balanced attribute subgraphs and characterize and reason about the spatiotemporal relationships between different fields in attack traffic to solve the data imbalance and classification uncertainty problems existing in the prior art, thereby improving the overall performance and reliability of related models and providing a more solid technical guarantee for network security.
[0077] The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention provided in this specification can be applied to electronic devices with network security processing capabilities. The electronic device may include a notebook, a desktop computer, a smart phone, a smart wearable device (virtual reality glasses, smart watches, etc.), a tablet computer, etc. Of course, the anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention provided in this specification can also be applied to applications running in the above-mentioned electronic devices. For example, the message display method can be applied to a browser with network security processing capabilities, and can also be applied to immediately corresponding network security processing software.
[0078] See also Figure 1, Figure 1 A flowchart of an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to an embodiment of the present invention is shown. The method may include the following steps:
[0079] S101, obtain the network data to be detected, and input the network data to be detected into the trained network anomaly detection model, and finally obtain the anomaly data and the category of the anomaly data output by the network anomaly detection model. The anomaly data and the category of the anomaly data are the anomaly detection classification results output by the network anomaly detection model. When the network anomaly detection model detects that there is an intrusion in the network data, it can respond in time, protect the log, cut off the connection or report to the network administrator in time, etc., which can resist the intrusion in real time and prevent the loss from continuing to expand.
[0080] Of course, it is understandable that if there is no abnormal traffic data that affects network security in the network data, the network anomaly detection model will not output abnormal data and the category of abnormal data.
[0081] In this embodiment, the network anomaly detection model includes a diffusion-enhanced heterogeneous graph attention network (Diffusion-HAN) layer and a classification network layer, wherein the Diffusion-HAN layer is used to obtain similarity features and time embedding of events containing attribute nodes based on feature information of attribute nodes and feature information of neighbor nodes of attribute nodes, and the classification network layer is used to determine abnormal data and the category of abnormal data based on the similarity features and embedding of meta-paths.
[0082] During training, the training data of the network anomaly detection model includes: sample data, similarity features of sample data, and event embedding.
[0083] In this embodiment, all known historical intrusion offline data are collected as sample data to provide massive data support for subsequent training. The historical intrusion data can be stored in the electronic device in advance or obtained by the electronic device from the outside. There is no restriction on the specific acquisition form of historical intrusions, as long as the electronic device can obtain the sample data.
[0084] See also Figure 2 , where the network anomaly detection model is trained through the following steps:
[0085] S102, acquiring sample data, and performing time-discrete processing on the sample data according to the degree of temporal correlation, to obtain a number of time-discrete data groups within a similar time period.
[0086] The time (segment) discrete processing performed according to the time series correlation degree can divide the sample data in different time periods into different groups, and obtain several of the above-mentioned time discrete data groups. It can be understood that the time discrete data group is composed of several sample data, that is, the time discrete data group contains several sample data, and these sample data include abnormal data and normal data.
[0087] S103, determining the characteristic information of each sample data in each time discrete data group, and constructing a data attribute subgraph according to the characteristic information. The data attribute subgraph is composed of a number of meta-paths, and the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths.
[0088] In this embodiment, the characteristic information includes but is not limited to: service type, duration, number of data packet bits, and flags, etc. By taking these characteristic information as the attributes corresponding to the data, and then dividing each sample data in each discrete time data group into a number of attribute nodes according to the attribute value of the sample data, connecting the attribute nodes of corresponding types to construct a number of meta-paths, and then constructing a data attribute subgraph according to these meta-paths.
[0089] See also Figure 3 and Figure 4 In the data attribute subgraph, the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths. Through the data attribute subgraph, the connection between nodes in the same meta-path can be established, and the connection between different data with the same attribute value can also be established. The data attribute subgraph not only considers the relationship between different data, but also mines the relationship between attribute values.
[0090] The data attribute subgraph can be expressed as:
[0091] G=(V,E,R)(1)
[0092] In formula (1), G represents the data attribute subgraph; V represents the set of all attribute values v in the data attribute subgraph G; E represents the set of edges between attribute nodes in the data attribute subgraph G, and each element e in E represents the relationship between the attribute values of the edge consisting of two attribute nodes; R represents the set of relationships between different types of attribute nodes in the data attribute subgraph G. The set R in the balanced attribute subgraph includes all meta-path types, and each element r in R can represent attribute relationships between different types, such as duration-protocol type, duration-flag value, etc.
[0093] S104, according to the ratio of abnormal data to normal data in the sample data of each data attribute subgraph, and the sample data and the neighbor data of the sample data, balance the ratio of abnormal data to normal data to obtain a balanced attribute subgraph corresponding to each data attribute subgraph.
[0094] After time discretization processing, construction of data attribute subgraphs based on data feature information, and data balancing processing, balanced attribute subgraphs for multiple time periods can be constructed. Then, Diffusion-HAN multi-attention representation learning with spatiotemporal attribute subgraph diffusion enhancement is performed based on these multi-time period balanced attribute subgraphs.
[0095] S105, input the balanced attribute subgraph into the Diffusion-HAN layer, and iteratively update the network parameters of the Diffusion-HAN layer based on the total loss function L, to obtain the similarity features and event embedding of the balanced attribute subgraph output by the Diffusion-HAN layer. The total loss function of the Diffusion-HAN layer is composed of the class distance loss function L1, the normalization loss function L2, and the attribute level structure perception loss function L3.
[0096] The class distance loss function L1 is used to maximize the distance between categories, the regularization loss function L2 is used to minimize the regularization loss, and the attribute-level structure-aware loss function L3 is used to perform attribute-level structure-aware loss.
[0097] The Diffusion-HAN layer is used to perform Diffusion-HAN multi-attention representation learning for spatiotemporal attribute subgraph diffusion enhancement. Since there are different types of attribute nodes in the balanced attribute subgraph, and the relationship between attribute nodes also has multiple types, the balanced attribute subgraph is essentially a heterogeneous graph. In this embodiment, the purpose of Diffusion-HAN multi-attention representation learning for spatiotemporal attribute subgraph diffusion enhancement is to find the relationship between different types of attribute nodes. The diffusion process refers to the diffusion enhancement of the balanced attribute subgraph representation learning ability. Diffusion is a physical process of balancing concentration differences without generating or destroying quality. After obtaining the node representation after attention processing, a first-order diffusion process is performed to capture the feature information of neighboring nodes. The embedding of Diffusion-HAN mainly focuses on the preservation of structural information based on meta-paths. The network anomaly detection model based on Diffusion-HAN attention representation learning uses a hierarchical attention structure. After generating the initial embedding, for each meta-path, node-level attention is also considered, that is, the attribute node will have neighbor nodes based on the meta-path and learn new feature representations through the self-attention mechanism.
[0098] Different from feature-based classification methods (such as support vector machines (SVMs) and decision trees) and deep learning methods (such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), which often find it difficult to accurately characterize the differentiated features between different attack categories when processing network security data, resulting in blurred boundaries between data types, the network anomaly handling model in this embodiment uses Diffusion-HAN multi-attention representation learning enhanced by spatiotemporal attribute subgraph diffusion, which not only considers the relationship between different types of data volumes, but also mines the relationship between data fields, thereby improving the classification accuracy.
[0099] S106, using the similarity feature as input data and the event embedding as label information for training, inputting into the classification network layer and iteratively updating the network parameters of the classification network layer based on a supervised machine learning method to obtain a network anomaly detection model. The network anomaly detection model is used to output the anomaly data and the category of the anomaly data in the network data to be detected.
[0100] The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention obtains anomaly data and the category of anomaly data output by the network anomaly detection model by inputting the network data to be detected into the trained network anomaly detection model. The network anomaly detection model processes unbalanced data by constructing a balanced attribute subgraph of multiple periods, thereby avoiding overfitting and improving the recognition ability of the network anomaly detection model for attack traffic. At the same time, the network anomaly detection model is based on a diffusion enhanced heterogeneous graph attention network to characterize and reason about the spatiotemporal relationship of different fields in the attack traffic, not only considering the relationship between different types of data, but also mining the relationship between data fields, so that the difference between the features of abnormal data and normal data is more obvious, and the classification accuracy is improved. The network anomaly detection model solves the problems of data imbalance and classification uncertainty, achieves higher classification accuracy and stronger model adaptability, and can process large-scale network traffic data, improves the accuracy and robustness of the model, so that the trained network anomaly detection model can more effectively identify and defend against network attacks, and provide more powerful protection measures for network security. The network anomaly detection model has important application value in the field of network security.
[0101] See also Figure 5 , Figure 5 Another flow chart of an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to an embodiment of the present invention is shown. In this method, a network anomaly detection model is trained by the following steps:
[0102] S201, obtaining original data, and preprocessing the original data to obtain sample data.
[0103] Before the sample data is time discretized, the present embodiment also pre-processes the acquired raw data to obtain sample data for training. The purpose of pre-processing is to clean the data and discretize the continuous values. The sample data obtained after the above pre-processing is convenient for the neural network to process, so as to determine whether there is intrusion in the network in the next step.
[0104] S202: Perform time-discrete processing on the sample data according to the time series correlation degree to obtain a number of time-discrete data groups within a similar time period. For details, refer to step S102.
[0105] S203, determining the characteristic information of each sample data in each time discrete data group, and constructing a data attribute subgraph according to the characteristic information. For details, refer to step S103.
[0106] S204: According to the ratio of abnormal data to normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, the abnormal data and the normal data are proportionally balanced to obtain a balanced attribute subgraph corresponding to each data attribute subgraph. For details, refer to step S104.
[0107] S205, input the balanced attribute subgraph into the Diffusion-HAN layer, and iteratively update the network parameters of the Diffusion-HAN layer based on the total loss function L, to obtain the similarity features and event embedding of the balanced attribute subgraph output by the Diffusion-HAN layer. The total loss function of the Diffusion-HAN layer is composed of the class distance loss function L1, the normalization loss function L2, and the attribute level structure perception loss function L3. For details, refer to step S105.
[0108] S206: Use the similarity feature as input data and the event embedding as label information for training, input it into the classification network layer, and iteratively update the network parameters of the classification network layer based on supervised machine learning to obtain a network anomaly detection model. For details, refer to step S106.
[0109] See also Figure 6 , Figure 6 Another flow chart of an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to an embodiment of the present invention is shown. The method may include the following steps:
[0110] S3011. Obtain original data, and clean up duplicate data and incomplete data in the original data, so that the original data after cleaning are all continuous data.
[0111] S3012. Perform data discretization processing on the cleaned original data.
[0112] Preferably, in the embodiment of the present invention, the data discretization processing mentioned above can be performed by using the equal width method, the equal frequency method and the like. No limitation is imposed on the specific means of the discretization processing, and it is only necessary to ensure that the discretized data can be obtained.
[0113] S3013. Perform binary conversion on the discretized original data to obtain sample data that can be applied to machine learning.
[0114] Preferably, the above-mentioned binary conversion processing is performed through one-hot encoding, which is used to convert discrete classification labels into binary vectors, so that enumeration fields such as protocols in sample data can be converted into binary vectors to facilitate the construction of attribute subgraphs.
[0115] S302: Perform time-discrete processing on the sample data according to the time series correlation degree to obtain a number of time-discrete data groups within a similar time period. For details, refer to step S102.
[0116] S303: Determine the characteristic information of each sample data in each time discrete data group, and construct a data attribute subgraph according to the characteristic information. For details, refer to step S103.
[0117] S304: According to the ratio of abnormal data to normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, the abnormal data and the normal data are proportionally balanced to obtain a balanced attribute subgraph corresponding to each data attribute subgraph. For details, refer to step S104.
[0118] S305, input the balanced attribute subgraph into the Diffusion-HAN layer, and iteratively update the network parameters of the Diffusion-HAN layer based on the total loss function L, to obtain the similarity features and event embedding of the balanced attribute subgraph output by the Diffusion-HAN layer. The total loss function of the Diffusion-HAN layer is composed of the class distance loss function L1, the normalization loss function L2, and the attribute level structure perception loss function L3. For details, refer to step S105.
[0119] S306: Use the similarity feature as input data and the event embedding as label information for training, input it into the classification network layer, and iteratively update the network parameters of the classification network layer based on supervised machine learning to obtain a network anomaly detection model. For details, refer to step S106.
[0120] and Figure 1Compared with the embodiment shown, in this embodiment, the sample data obtained through the related processing of data cleaning, continuous value discretization and binary conversion can effectively eliminate noise data. Therefore, the attribute subgraph constructed based on the sample data can be better processed by the neural network to determine whether there is intrusion behavior in the network at this time in the next step.
[0121] See also Figure 7 , Figure 7 Another flow chart of an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to an embodiment of the present invention is shown. In this method, a network anomaly detection model is trained by the following steps:
[0122] S401, obtain original data, and pre-process the original data to obtain sample data. For details, refer to step S201.
[0123] S402: Perform time-discrete processing on the sample data according to the time series correlation degree to obtain a number of time-discrete data groups within a similar time period. For details, refer to step S102.
[0124] S403, determining the characteristic information of each sample data in each time discrete data group, and constructing a data attribute subgraph according to the characteristic information. For details, refer to step S103.
[0125] S4041. Determine the number of abnormal data and the number of normal data in the sample data of each data attribute subgraph respectively, and obtain a first number of abnormal data and a second number of normal data in each data attribute subgraph.
[0126] S4042: When it is determined that the ratio of the first number to the second number exceeds the preset ratio and the first number is greater than the second number, generate first incremental data based on the normal data and the neighboring nodes of the normal data, and add the first incremental data to the normal data until the ratio of the first number to the second number does not exceed the preset ratio. It is understandable that the first incremental data is normal data.
[0127] S4043: When it is determined that the ratio of the first number to the second number exceeds the preset ratio and the first number is less than the second number, generate second incremental data based on the abnormal data and the neighboring nodes of the abnormal data, and add the second incremental data to the abnormal data until the ratio of the first number to the second number does not exceed the preset ratio. It is understandable that the second incremental data is abnormal data.
[0128] Not exceeding the preset ratio indicates that the first quantity and the second quantity are very close, that is, the normal data and the abnormal data are in a balanced state. When the ratio of the first quantity to the second quantity does not exceed the preset ratio, it means that the ratio of normal data to abnormal data has reached a balanced ratio, thereby solving the problem of sample data imbalance that may exist in the model building process.
[0129] In order to solve the problem of unbalanced sample data, in this embodiment, by generating abnormal data or normal data of the minority class in the same attribute subgraph, the balance of the ratio of normal data to abnormal data in the data attribute subgraph is achieved, and then a balanced attribute subgraph is obtained.
[0130] At the same time, in this embodiment, constructing a balanced attribute subgraph for multiple time periods does not rely on the modification of sample data, but rather achieves data between different types of data by generating a smaller proportion of abnormal / normal data in a data balanced subgraph, thereby avoiding overfitting and improving the ability of the network anomaly detection model to identify attack traffic.
[0131] More specifically, the first / second incremental data is generated in the following manner:
[0132] First, the K nearest neighbor algorithm is used to find the K nearest neighbor nodes of a normal data / abnormal data, where K is a preset number, which is a pre-set positive integer, such as 5;
[0133] A neighbor node is randomly selected from the K neighbor nodes as the constructed sample data, and the original normal data / abnormal data and the constructed sample data are linearly interpolated to generate the first / second incremental data. The position of the new first / second incremental data is based on the original data and the feature value of the selected neighbor node and is obtained by random weight combination. This process is repeated multiple times to generate multiple first / second incremental data until a predetermined number of minority class sample data is expanded, so that the ratio of the first number to the second number does not exceed a preset ratio.
[0134] All the generated first / second incremental data are added to the normal data / abnormal data, thereby increasing the proportion of minority class sample data and making the total sample data set categories more balanced.
[0135] S405, input the balanced attribute subgraph into the Diffusion-HAN layer, and iteratively update the network parameters of the Diffusion-HAN layer based on the total loss function L, to obtain the similarity features and event embedding of the balanced attribute subgraph output by the Diffusion-HAN layer. The total loss function of the Diffusion-HAN layer is composed of the class distance loss function L1, the normalization loss function L2, and the attribute level structure perception loss function L3. For details, refer to step S105.
[0136] S406: Use the similarity feature as input data and the event embedding as label information for training, input it into the classification network layer, and iteratively update the network parameters of the classification network layer based on supervised machine learning to obtain a network anomaly detection model. For details, refer to step S106.
[0137] See also Figure 8 , Figure 8 Another flow chart of an anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to an embodiment of the present invention is shown. In this method, a network anomaly detection model is trained by the following steps:
[0138] S501, obtain original data, and pre-process the original data to obtain sample data. For details, refer to step S201.
[0139] S502: Perform time-discrete processing on the sample data according to the time series correlation degree to obtain a number of time-discrete data groups within a similar time period. For details, refer to step S102.
[0140] S503: Determine the characteristic information of each sample data in each time discrete data group, and construct a data attribute subgraph according to the characteristic information. For details, refer to step S103.
[0141] S504: According to the ratio of abnormal data to normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, the abnormal data and the normal data are proportionally balanced to obtain a balanced attribute subgraph corresponding to each data attribute subgraph. For details, refer to step S104.
[0142] S5051. First, obtain the attention weights between the attribute nodes in each meta-path in each data attribute subgraph, that is:
[0143]
[0144] In formula (2), Represents a meta path The attention weight of attribute node i to attribute node j; σ represents the activation function; h i represents the feature information of attribute node i, i.e., the initial embedding; h i Represents the feature information of attribute node j, i.e., the initial embedding; Represents the meta-path to be learned The transpose of the node-level attention vector.
[0145] S5052. Then, according to the attention weight, the node features of the attribute node after attention processing are obtained, that is:
[0146]
[0147] In formula (3), h' i represents the node features of attribute node i after attention processing; Represents attribute node i about the meta path All nodes of Contains the attribute node i itself.
[0148] S5053. Perform node-level attention processing on the node features of the attribute nodes for a preset number of times to obtain the path node embedding of the attribute nodes, and concatenate the path node embeddings of all attribute nodes in the meta-path together to obtain the meta-path embedding of the meta-path.
[0149] This embodiment introduces the diffusion process, which is a physical process that balances concentration differences without creating or destroying mass. After obtaining the node representation after attention processing, a first-order diffusion process is performed to capture the characteristic information of the neighbors, namely:
[0150]
[0151] In formula (4), Z i Represents a meta path β represents the path node embedding of attribute node i in the balanced data subgraph; β represents the pre-set hyperparameter. It can be understood that there are P meta-paths (P ≥ 1) in a balanced data subgraph, and the path node embedding combination of all attribute nodes in the pth meta-path in a balanced data subgraph can obtain the corresponding meta-path Metapath embedding (group) Among them
[0152] S5054. Perform weighted processing on the meta-path embedding according to the semantic-level attention to obtain the importance of the meta-path.
[0153] The purpose of node-level attention is to learn the importance between a node and its adjacent nodes based on meta-paths, and the purpose of semantic-level attention is to learn the importance of different meta-paths. Considering that node-level attention can only reflect nodes from one aspect, in an embodiment of the present invention, semantic-level attention is used to perform weighted processing on meta-path embedding to learn the importance of different meta-paths, namely:
[0154]
[0155] In formula (5), Represents the pth meta-path The importance of; W represents the weight matrix; f represents the bias vector; q Trepresents the preset semantic-level attention vector; U represents the meta-path A collection of attribute nodes in .
[0156] It is particularly noted that in the embodiment of the present invention, all meta-paths share the parameters W, f, and q for training and intermediate data calculation.
[0157] S5055. Determine the meta-path weight of the meta-path according to the importance of the meta-path.
[0158] Get the pth meta-path Importance After that, based on importance The meta-path weight of each meta-path can be further obtained Right now:
[0159]
[0160] Preferably, in the embodiment of the present invention, the softmax function is used to After normalization, we get the meta-path weight
[0161] S5056. Obtain an attribute node embedding of the attribute node according to all meta-paths and meta-path weights that the attribute node passes through in the balanced attribute subgraph.
[0162] Since each attribute node in the balanced attribute subgraph may pass through several meta-paths, it is also necessary to obtain the meta-path weights Then determine the attribute node embedding of each attribute node, that is:
[0163]
[0164] In formula (7), Z represents the attribute node embedding of the attribute node.
[0165] S5057. Obtain event embedding of the balanced attribute subgraph according to the attribute node embedding of all attribute nodes in each balanced attribute subgraph.
[0166] From the obtained attribute node embedding Z of the attribute node, we can further obtain the event embedding d of the balanced data subgraph composed of attribute nodes, that is, the event, that is:
[0167]
[0168] In formula (8), ε j represents the loss parameter; Z j represents the attribute node embedding of attribute node j; n represents the total number of attribute nodes in the balanced attribute subgraph.
[0169] At the same time, in order to maximize the distance between abnormal nodes and normal nodes, that is, for both abnormal and normal events, the Euclidean distance between them is expected to be as large as possible, in this embodiment, the inter-category distance and the inter-category distance are set, and the corresponding class distance loss function is L1.
[0170] More specifically, in the embodiment of the present invention, for each type of event, the geometric center corresponding to the event is first obtained, that is:
[0171]
[0172] In formula (9), c represents the geometric center of the event corresponding to the abnormal / normal type; N represents the total number of events corresponding to the abnormal / normal type; and B represents all events corresponding to the abnormal / normal type.
[0173] Since there are only data of two types of events, abnormal and normal, the class distance loss function L1 can be expressed as:
[0174] L1=-||c1-c2||(10)
[0175] In formula (10), c1 represents the geometric center of abnormal type events; c2 represents the geometric center of normal type events.
[0176] In the process of continuous training and optimization of the loss function L1, the parameter ε j The value of will tend to become very large, so it is also necessary to use the normalized loss function l2, that is:
[0177]
[0178] In the embodiment of the present invention, the attribute-level structure-aware loss function l3 of Diffusion-HAN is set to:
[0179]
[0180] In formula (12), Y v Represents the label corresponding to the attribute node v, and C represents the parameters corresponding to the classifier.
[0181] See also Fig. 9 , after determining the three corresponding loss functions, the total loss function L can be obtained as:
[0182] L=L1+L2+L3(13)
[0183] After determining the total loss function L, the total loss function L is used to train and optimize the network parameters of the Diffusion-HAN layer, and finally the optimal parameters are obtained.
[0184] In the process of training Diffusion-HAN, all balanced attribute subgraphs are taken as inputs in turn and all balanced attribute subgraphs are trained using the same Diffusion-HAN network parameters. According to the structure of the attribute graph, the attribute fields of the same data are adjacent. The diffusion process diffuses and enhances the learning ability of the balanced attribute subgraph representation. Diffusion is a physical process that balances concentration differences without generating or destroying mass. After obtaining the node representation after attention processing, a first-order diffusion process is performed to capture the characteristic information of neighboring nodes. Since the training process adopts the diffusion attention mechanism, the same node will be processed by the attention mechanism multiple times, and the results will be connected in series, and then the neighboring node information will be aggregated in the current node. In this way, the diffusion attention mechanism enhanced by the diffusion of spatiotemporal attribute subgraphs makes the Diffusion-HAN training process more stable, avoiding the problem of too many connection types caused by long meta-paths, and also avoiding the problem of excessive computational burden.
[0185] S5058. Based on the attribute node embedding of each attribute node j in a certain event b, and the degree of correlation between the attribute node embeddings of the attribute node j and its adjacent attribute node k (that is, the neighbor node k of the attribute node j), the similarity sim(b) between the attribute node embeddings of any two attribute nodes in a certain event can be obtained. j ,b k ), and then after corresponding processing, the similarity characteristics of the event can be obtained, specifically:
[0186]
[0187] In formula (14), sim(b j ,b k ) represents the similarity between the attribute node embedding of attribute node j and the attribute node embedding of its adjacent attribute node k; b j b represents the attribute node embedding of attribute node j; k Represents the attribute node embedding of attribute node k. By traversing all attribute nodes in event b, the similarity group of event b can be obtained.
[0188] Then, based on the similarity group of event b, the similarity average and variance of event b are obtained, namely:
[0189]
[0190] In formulas (15) to (16), sim_avg(b) represents the average similarity of event b; sim_var(b) represents the variance of the similarity of event b. j ,b k )=(sim(b j ,bk )-sim_avg(b)) 2 Then, according to the similarity average value and similarity variance of event b, the similarity vector s of the event can be obtained, and the similarity vector s is the similarity feature.
[0191] The trained Diffusion-HAN layer is used to learn the representation of network data, which can obtain the embedded representation of network data fields, calculate the similarity features between field nodes, and then determine the embedding of network data.
[0192] S506: Use the similarity features of the meta-path as input data and the embedding of the meta-path as label information for training, input them into the classification network layer, and iteratively update the network parameters of the classification network layer based on a supervised machine learning method to obtain a network anomaly detection model. For details, refer to step S106.
[0193] Abnormal data generally has numerically abnormal attributes. The similarity between attributes in an event can be used to predict whether an event is abnormal. The softmax classifier uses supervised machine learning to perform iterative training. The classifier parameters are optimized according to the loss function of the softmax classifier. Finally, the trained softmax classifier is used for classification and abnormal data in the network data is detected for network anomaly detection.
[0194] The binary classification function used by the softmax classifier is the logistic function, specifically:
[0195]
[0196] In formula (17), if g(s)≤0.5, it is determined as a normal event, that is, there is no network abnormality, and if g(s)>0.5, it is determined as an abnormal event, that is, there is a network abnormality. The above 0.5 can be specifically configured according to the actual application scenario.
[0197] The following describes an apparatus provided by an embodiment of the present invention. The apparatus described below and the method described above can refer to each other.
[0198] See also Fig.10 , Fig.10 A schematic diagram of the structure of an anomaly detection device based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to an embodiment of the present invention is shown. The device may include:
[0199] The anomaly detection module 10 is used to obtain the network data to be detected, and input the network data to be detected into the trained network anomaly detection model, and finally obtain the anomaly data and the category of the anomaly data output by the network anomaly detection model. The anomaly data and the category of the anomaly data are the anomaly detection classification results output by the network anomaly detection module. When the network anomaly detection model detects that there is an intrusion in the network data, it can respond in time, protect the log, cut off the connection, or report to the network administrator in time, etc., which can resist the intrusion in real time and prevent the loss from continuing to expand.
[0200] Of course, it is understandable that if there is no abnormal traffic data that affects network security in the network data, the network anomaly detection model will not output abnormal data and the category of abnormal data.
[0201] In this embodiment, the network anomaly detection model includes a Diffusion-HAN layer and a classification network layer, wherein the Diffusion-HAN layer is used to obtain the similarity features and embedding of the meta-path containing the attribute nodes based on the feature information of the attribute nodes and the feature information of the neighbor nodes of the attribute nodes, and the classification network layer is used to determine the abnormal data and the category of the abnormal data based on the similarity features and embedding of the meta-path.
[0202] During training, the training data of the network anomaly detection model includes: sample data, similarity features of sample data, and event embedding.
[0203] In this embodiment, all known historical intrusion offline data are collected as sample data to provide massive data support for subsequent training. The historical intrusion data can be stored in the electronic device in advance or obtained by the electronic device from the outside. There is no restriction on the specific acquisition form of historical intrusions, as long as the electronic device can obtain the sample data.
[0204] The network anomaly detection models include:
[0205] The time discrete module 20 is used to obtain sample data and perform time discrete processing on the sample data according to the degree of temporal correlation to obtain a number of time discrete data groups within a similar time period.
[0206] The time (segment) discrete processing performed according to the degree of temporal correlation can divide the sample data in different time periods into different groups, and obtain several time discrete data groups in similar time periods. It can be understood that the time discrete data group is composed of several sample data, that is, the time discrete data group contains several sample data, and these sample data include abnormal data and normal data.
[0207] The graph construction module 30 is used to determine the characteristic information of each sample data in each time discrete data group, and construct a data attribute subgraph according to the characteristic information. The data attribute subgraph is composed of a plurality of meta-paths, and the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths.
[0208] The data balancing module 40 is used to balance the proportion of abnormal data and normal data according to the proportion of abnormal data and normal data in the sample data of each data attribute subgraph, as well as the sample data and the neighbor data of the sample data, to obtain a balanced attribute subgraph corresponding to each data attribute subgraph.
[0209] After time discretization processing, construction of data attribute subgraphs based on data feature information, and data balancing processing, balanced attribute subgraphs for multiple time periods can be constructed. Then, Diffusion-HAN multi-attention representation learning with spatiotemporal attribute subgraph diffusion enhancement is performed based on these multi-time period balanced attribute subgraphs.
[0210] The feature learning module 50 is used to input the balanced attribute subgraph into the Diffusion-HAN layer, and iteratively update the network parameters of the Diffusion-HAN layer based on the total loss function L, to obtain the similarity features and event embedding of the balanced attribute subgraph output by the Diffusion-HAN layer. The total loss function of the Diffusion-HAN layer is composed of the class distance loss function L1, the normalization loss function L2 and the attribute level structure perception loss function L3.
[0211] The classification detection module 60 is used to input the similarity feature as input data and the event embedding as label information used for training into the classification network layer and iteratively update the network parameters of the classification network layer based on a supervised machine learning method to obtain a network anomaly detection model. The network anomaly detection model is used to output the anomaly data and the category of the anomaly data in the network data to be detected.
[0212] The anomaly detection device based on multi-period balanced attribute subgraphs and heterogeneous graph attention of the present invention obtains anomaly data and the category of anomaly data output by the network anomaly detection model by inputting the network data to be detected into the trained network anomaly detection model. The network anomaly detection model processes unbalanced data by constructing a multi-period balanced attribute subgraph, thereby avoiding overfitting and improving the recognition ability of the network anomaly detection model for attack traffic. At the same time, the network anomaly detection model is based on a diffusion enhanced heterogeneous graph attention network to characterize and reason about the spatiotemporal relationship of different fields in the attack traffic, not only considering the relationship between different types of data, but also mining the relationship between data fields, so that the difference between the features of abnormal data and normal data is more obvious, and the classification accuracy is improved. The network anomaly detection model solves the problems of data imbalance and classification uncertainty, achieves higher classification accuracy and stronger model adaptability, and can process large-scale network traffic data, improves the accuracy and robustness of the model, so that the trained network anomaly detection model can more effectively identify and defend against network attacks, and provide more powerful protection measures for network security. The network anomaly detection model has important application value in the field of network security.
[0213] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention, characterized by: The method comprises: Acquire the network data to be detected, and input the network data to be detected into a trained network anomaly detection model to obtain the anomaly data output by the network anomaly detection model and the category of the anomaly data; the network anomaly detection model includes a diffusion enhanced heterogeneous graph attention network layer and a classification network layer, and the network anomaly detection model is obtained based on sample data, similarity features of the sample data and event embedding training; The network anomaly detection model is trained by the following steps: Acquire the sample data, and perform time-discrete processing on the sample data according to the time series correlation degree to obtain a number of time-discrete data groups within a similar time period; the time-discrete data groups contain a number of the sample data; Determine the characteristic information of each sample data in each of the time discrete data groups, and construct a data attribute subgraph according to the characteristic information; the data attribute subgraph is composed of a plurality of meta-paths, and the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths; According to the ratio of abnormal data to normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, the abnormal data and the normal data are proportionally balanced to obtain a balanced attribute subgraph corresponding to each data attribute subgraph; Inputting the balanced attribute subgraph into the diffusion-enhanced heterogeneous graph attention network layer, and iteratively updating the network parameters of the diffusion-enhanced heterogeneous graph attention network layer based on the total loss function, to obtain the similarity features and the event embedding of the balanced attribute subgraph output by the diffusion-enhanced heterogeneous graph attention network layer; the total loss function is composed of a class distance loss function, a normalized loss function, and an attribute-level structure-aware loss function; The similarity feature is used as input data and the event embedding is used as label information for training, which is input into the classification network layer and the network parameters of the classification network layer are iteratively updated based on a supervised machine learning method to obtain a network anomaly detection model; The expression of the total loss function is: ; in, represents the total loss function, represents the class distance loss function, represents the normalized loss function, represents the attribute-level structure-aware loss function; The expression of the class distance loss function is: ; in, Indicates the geometric center of the abnormal type event, Represents the geometric center of normal type events; The expression of the normalized loss function is: ; in, represents the loss parameter, Represents the total number of attribute nodes in the balanced attribute subgraph; The expression of the attribute-level structure-aware loss function is: ; in, Represents all attribute values in the data attribute subgraph A collection of Representation Node The corresponding label, C Represents the parameters corresponding to the classifier, For attribute nodes The corresponding attribute node is embedded.
2. The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to claim 1 is characterized in that: The method of performing a proportional balance on the abnormal data and the normal data according to the ratio of the abnormal data to the normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, and obtaining a balanced attribute subgraph corresponding to each data attribute subgraph specifically includes: Determine the number of abnormal data and normal data in the sample data of each data attribute subgraph respectively, and obtain a first number of abnormal data and a second number of normal data in each data attribute subgraph; When it is determined that the ratio of the first number to the second number exceeds the preset ratio and the first number is greater than the second number, generating first incremental data based on the normal data and the neighboring nodes of the normal data, and adding the first incremental data to the normal data until the ratio of the first number to the second number does not exceed the preset ratio; When it is determined that the ratio of the first number to the second number exceeds the preset ratio and the first number is less than the second number, second incremental data is generated based on the abnormal data and the neighboring nodes of the abnormal data, and the second incremental data is added to the abnormal data until the ratio of the first number to the second number does not exceed the preset ratio.
3. The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to claim 2 is characterized in that: The first incremental data is generated by the following steps: Use the nearest neighbor algorithm to find a preset number of neighbor nodes for normal data; One of the neighbor nodes is used as constructed sample data, and linear interpolation is performed according to the normal data and the constructed sample data to generate first incremental data.
4. The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to claim 1 is characterized in that: The step of inputting the balanced attribute subgraph into the diffusion enhanced heterogeneous graph attention network layer, and iteratively updating the network parameters of the diffusion enhanced heterogeneous graph attention network layer based on the total loss function, to obtain the similarity feature of the balanced attribute subgraph output by the diffusion enhanced heterogeneous graph attention network layer and the event embedding, specifically includes: Obtaining the attention weights between the attribute nodes in each meta-path in each data attribute subgraph; According to the attention weight, obtaining the node feature of the attribute node; Performing a preset number of node-level attention processing on the node features to obtain a path node embedding of the attribute node, and concatenating the path node embeddings of all attribute nodes in a meta-path to obtain a meta-path embedding of the meta-path; weighting the meta-path embedding according to semantic-level attention to obtain the importance of the meta-path; determining the meta-path weight of the meta-path according to the importance of the meta-path; Obtaining an attribute node embedding of the attribute node according to all meta-paths through which the attribute node passes in the balanced attribute subgraph and the meta-path weights; Obtaining the event embedding of the balanced attribute subgraph according to the attribute node embeddings of all the attribute nodes in each balanced attribute subgraph; According to the correlation degree between the attribute node and the attribute node embeddings of the neighboring nodes of the attribute node, the similarity between the attribute node embeddings is obtained, and the similarity feature of the balanced attribute subgraph is determined according to the similarity.
5. The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention as claimed in claim 1, characterized in that: The method further comprises the following steps: The original data is obtained and preprocessed to obtain sample data.
6. The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to claim 5 is characterized in that: The obtaining of raw data and preprocessing of the raw data to obtain sample data specifically includes: Obtain the original data and clean up the duplicate and incomplete data in the original data; Discretize the cleaned raw data; The discretized original data is converted into binary form to obtain sample data.
7. The anomaly detection method based on multi-period balanced attribute subgraphs and heterogeneous graph attention according to claim 1 is characterized in that: The expression of the geometric center is: ; in, Indicates the geometric center of the event corresponding to the abnormal / normal type, Indicates the total number of events corresponding to abnormal / normal types, Indicates all events corresponding to abnormal / normal types.
8. An anomaly detection device based on multi-period balanced attribute subgraphs and heterogeneous graph attention, characterized in that: The device comprises: An anomaly detection module is used to obtain network data to be detected, and input the network data to be detected into a trained network anomaly detection model to obtain anomaly data output by the network anomaly detection model and the category of the anomaly data; the network anomaly detection model includes a diffusion enhanced heterogeneous graph attention network layer and a classification network layer, and the network anomaly detection model is obtained based on sample data, similarity features of the sample data and event embedding training; The network anomaly detection model includes: A time discrete module is used to obtain the sample data and perform time discrete processing on the sample data according to the time series correlation degree to obtain a number of time discrete data groups within a similar time period; the time discrete data group contains a number of the sample data; A graph construction module, used to determine the characteristic information of each sample data in each of the time discrete data groups, and to construct a data attribute subgraph according to the characteristic information; the data attribute subgraph is composed of a plurality of meta-paths, and the attribute nodes in the same meta-path are fully connected to each other, and the attribute nodes with the same attribute value are shared between different meta-paths; A data balancing module, used to balance the abnormal data and the normal data according to the ratio of the abnormal data to the normal data in the sample data of each data attribute subgraph, the sample data and the neighbor data of the sample data, so as to obtain a balanced attribute subgraph corresponding to each data attribute subgraph; A feature learning module, used for inputting the balanced attribute subgraph into the diffusion enhanced heterogeneous graph attention network layer, and iteratively updating the network parameters of the diffusion enhanced heterogeneous graph attention network layer based on a total loss function, to obtain the similarity feature of the balanced attribute subgraph output by the diffusion enhanced heterogeneous graph attention network layer and the event embedding; the total loss function is composed of a class distance loss function, a normalization loss function and an attribute-level structure perception loss function; A classification detection module, used to input the similarity feature as input data and the event embedding as label information used for training into the classification network layer and iteratively update the network parameters of the classification network layer based on a supervised machine learning method to obtain a network anomaly detection model; The expression of the total loss function is: ; in, represents the total loss function, represents the class distance loss function, represents the normalized loss function, represents the attribute-level structure-aware loss function; The expression of the class distance loss function is: ; in, Indicates the geometric center of the abnormal type event, Represents the geometric center of normal type events; The expression of the normalized loss function is: ; in, represents the loss parameter, Represents the total number of attribute nodes in the balanced attribute subgraph; The expression of the attribute-level structure-aware loss function is: ; in, Represents all attribute values in the data attribute subgraph A collection of Representation Node The corresponding label, C Represents the parameters corresponding to the classifier, For attribute nodes The corresponding attribute node is embedded.
Citation Information
Patent Citations
Deep learning-oriented network anomaly detection method and device, storage medium and equipment
CN118070107A