Network abnormal traffic identification model training method, device and system and medium
By processing and updating the traffic graph samples of the network anomaly traffic identification model, and optimizing the model parameters using activation layers, aggregation networks, and classification networks, the problem of insufficient self-adjustment ability of the model is solved, and the identification accuracy and robustness are improved.
Patent Information
- Application Number
- CN202511015126.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-24
AI Technical Summary
The network abnormal traffic identification model based on machine learning lacks self-regulation ability, resulting in limited model generalization ability and low recognition accuracy.
By acquiring traffic graph samples from the network anomaly traffic identification model, the node features are processed and updated using activation layers, aggregation networks, and classification networks. The model parameters are then optimized using a loss function to achieve self-adjustment and learning.
It improves the accuracy and robustness of the network anomaly traffic identification model in complex network environments, and can dynamically adjust internal parameters to optimize the identification effect.
Smart Images

Figure CN120834951A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network traffic identification, and in particular to a network abnormal traffic identification model training method, device, system and medium. BACKGROUND
[0002] As a core component of the network security protection system, network abnormal traffic detection is committed to quickly identifying abnormal network behavior, preventing potential attack events from interfering with the system, and thus ensuring the reliability of the network environment and the security of data transmission.
[0003] Traditional machine learning-based methods are widely used due to their good interpretability, computational efficiency and small sample adaptability. A large number of normal and abnormal traffic data are trained by machine learning algorithms to obtain models such as support vector machine (SVM) and K-nearest neighbor algorithm (KNN), and subsequently, the trained model is used to classify and identify new network traffic. However, the model trained based on machine learning algorithms lacks self-regulation ability, and thus cannot automatically correct model parameters, resulting in limited generalization ability of the model and low accuracy of model identification. SUMMARY
[0004] The present application provides a network abnormal traffic identification model training method, device, system and medium, which can solve at least one of the above technical problems.
[0005] In a first aspect, the present application provides a network abnormal traffic identification model training method, comprising:
[0006] Obtaining a traffic graph sample of a network abnormal traffic identification model, wherein the traffic graph sample includes node features of each node, an abnormal traffic label, and edge features of edges between each node and its neighbor nodes;
[0007] For each node, the following node update operation is performed:
[0008] Based on the activation layer in the abnormal traffic identification model, the node features of the node and the node features of each neighbor node of the node are processed to obtain the weight factor of the edge between the node and each neighbor node;
[0009] Based on the aggregation network in the abnormal traffic identification model, the node features of each neighbor node, the weight factor of the edge between the node and each neighbor node, and the edge features are aggregated to obtain the aggregated flow features from the node to each neighbor node;
[0010] splicing the node feature of the node and the aggregated flow feature from the node to each of the neighbor nodes based on a splicing network in the abnormal traffic identification model, and updating the node feature of the node based on a splicing result; and
[0011] performing classification prediction on the updated node feature of the node based on a classification network in the abnormal traffic identification model to obtain an abnormal traffic prediction result of the node.
[0012] In a case that the node updating operation has been performed on each of the nodes to obtain the abnormal traffic prediction information of each of the nodes, updating an activation layer, an aggregation network, a splicing network and a classification network in the network abnormal traffic identification model based on the abnormal traffic label and the abnormal traffic prediction result of each of the nodes.
[0013] In a second aspect, an embodiment of the present application provides a training device of a network abnormal traffic identification model, comprising:
[0014] a traffic graph sample acquisition module configured to acquire a traffic graph sample of the network abnormal traffic identification model, wherein the traffic graph sample comprises a node feature of each node and an abnormal traffic label, and an edge feature of an edge between each of the nodes and a neighbor node thereof;
[0015] a node updating module configured to perform the following node updating operation on each of the nodes:
[0016] a weight factor calculation unit configured to process the node feature of the node and the node feature of each of the neighbor nodes of the node based on an activation layer in the abnormal traffic identification model to obtain a weight factor of an edge between the node and each of the neighbor nodes;
[0017] an aggregation processing unit configured to perform aggregation processing on the node feature of each of the neighbor nodes, and the weight factor and the edge feature of the edge between the node and each of the neighbor nodes based on an aggregation network in the abnormal traffic identification model to obtain an aggregated flow feature from the node to each of the neighbor nodes;
[0018] a splicing unit configured to splice the node feature of the node and the aggregated flow feature from the node to each of the neighbor nodes based on a splicing network in the abnormal traffic identification model, and update the node feature of the node based on a splicing result; and
[0019] a classification prediction unit configured to perform classification prediction on the updated node feature of the node based on a classification network in the abnormal traffic identification model to obtain an abnormal traffic prediction result of the node.
[0020] The updating module is configured to, in a case where the node updating operation has been performed on each of the nodes to obtain the abnormal traffic prediction information of each of the nodes, update the activation layer, the aggregation network, the concatenation network and the classification network in the network abnormal traffic identification model based on the abnormal traffic label and the abnormal traffic prediction result of each of the nodes.
[0021] In a third aspect, an embodiment of the present application further provides a training system of a network abnormal traffic identification model, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in any one of the embodiments of the present application.
[0022] In a fourth aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to perform the method in any one of the embodiments of the present application.
[0023] According to the technical scheme, the traffic graph sample is used as the training sample of the network abnormal traffic identification model, and the traffic graph sample includes the node features of each node, the abnormal traffic label of each node, and the edge features of the edges between each node and its neighbor nodes. Based on the node features and the edge features of each node in the traffic graph sample, the updating operation of the node features is performed on each node to obtain the abnormal traffic prediction information of each node. Then, the loss function is calculated by comparing the abnormal traffic prediction result of each node with the corresponding abnormal traffic label, and the parameters of the activation layer, the aggregation network, the concatenation network and the classification network in the network abnormal traffic identification model are updated based on the loss function. Thus, the network abnormal traffic identification model can adaptively adjust the internal parameters according to the actual prediction effect in the continuous training process, and has the self-adjusting and learning ability. In this way, the network abnormal traffic identification model can efficiently identify the network abnormal traffic, and continuously optimize the identification through self-adjusting, thereby improving the accuracy and robustness of the model in identifying the abnormal traffic in the complex network environment. Subsequently, the network abnormal traffic can be identified based on the network abnormal traffic identification model.
[0024] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings are used to better understand the present application, and do not limit the present application. Among them:
[0026] Figure 1 is a flow chart of a training method of a network abnormal traffic identification model according to an embodiment of the present application;
[0027] Figure 2 is a structural block diagram of a training device of a network abnormal traffic identification model according to an embodiment of the present application;
[0028] Figure 3 is a structural block diagram of a weight factor calculation unit in a training device of a network abnormal traffic identification model according to an embodiment of the present application;
[0029] Figure 4 is a structural block diagram of an aggregation processing unit in a training device of a network abnormal traffic identification model according to an embodiment of the present application;
[0030] Figure 5 is a structural block diagram of a splicing unit in a training device of a network abnormal traffic identification model according to an embodiment of the present application;
[0031] Figure 6 is a block diagram of an electronic device for implementing the method according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] Exemplary embodiments of the present application will be described hereinafter with reference to the accompanying drawings, in which various specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. It will be apparent to one of ordinary skill in the art, however, that the embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring the embodiments of the present application. The following exemplary embodiments are described herein with reference to the figures, in which like reference numerals refer to like elements throughout the various figures.
[0033] Figure 1 is a flow chart of a training method of a network abnormal traffic identification model according to an embodiment of the present application.
[0034] As shown in Figure 1 , the training method of the network abnormal traffic identification model can include:
[0035] S110, obtaining a traffic pattern sample of a network abnormal traffic identification model, wherein the traffic pattern sample includes node features of each node and an abnormal traffic label, and edge features of edges between each node and its neighbor nodes;
[0036] S120, performing the following node update operation for each node:
[0037] S121, based on an activation layer in the abnormal traffic identification model, processing the node features of the node and the node features of each neighbor node of the node to obtain weight factors of edges between the node and each neighbor node;
[0038] S122, based on the aggregation network in the abnormal traffic identification model, aggregate the node features of each neighbor node and the weight factor and edge features of the edges between the node and each neighbor node to obtain aggregated flow features from the node to each neighbor node;
[0039] S123, based on the splicing network in the abnormal traffic identification model, splice the node features of the node and the aggregated flow features from the node to each neighbor node, and update the node features of the node based on the splicing result; and
[0040] S124, based on the classification network in the abnormal traffic identification model, classify and predict the updated node features of the node to obtain an abnormal traffic prediction result of the node.
[0041] S130, in a case where the node update operation has been performed on each node to obtain the abnormal traffic prediction information of each node, update the activation layer, the aggregation network, the splicing network and the classification network in the network abnormal traffic identification model based on the abnormal traffic label and the abnormal traffic prediction result of each node.
[0042] In the embodiment of the present application, the traffic pattern sample is used as the training sample of the network abnormal traffic identification model, which includes the node features of each node, the abnormal traffic label of the node and the edge features of the edges between each node and its neighbor nodes. The node update operation is performed on each node based on the node features of each node and the edge features of the edges between each node and its neighbor nodes to obtain the abnormal traffic prediction information of each node. Then, the activation layer, the aggregation network, the splicing network and the classification network in the network abnormal traffic identification model are updated based on the abnormal traffic label and the abnormal traffic prediction result of each node. In this way, the network abnormal traffic identification model in the present application can update the model parameters based on the abnormal traffic label and the abnormal traffic prediction result, so that it has the self-adjusting ability. It can be seen that the network abnormal traffic identification model in the present application can improve the accuracy of identifying the network abnormal traffic.
[0043] Specifically, the network abnormal traffic identification model comprises an activation layer, an aggregation network, a splicing network and a classification network, wherein the node update operation performed by each node comprises: processing, by the activation layer, the node feature of the node and the node features of neighbor nodes of the node to obtain a weight factor of an edge between the node and each neighbor node; performing, by the aggregation network, aggregation processing on the node features of each neighbor node, the weight factor of the edge between the node and each neighbor node and the edge feature to obtain an aggregated flow feature from the node to each neighbor node; performing, by the splicing network, splicing on the node feature of the node and the aggregated flow feature from the node to each neighbor node to obtain a splicing result, and updating the node feature of the node based on the splicing result to obtain an updated node feature of the node; and performing, by the classification network, classification prediction on the updated node feature of the node to obtain an abnormal traffic prediction result of the node.
[0044] Exemplarily, the traffic pattern sample is obtained by constructing a graph structure based on five-tuple data in original traffic data. A data packet is the smallest unit constituting the original traffic data (or the original traffic data is composed of at least one data packet), and the five-tuple data is stored in the data packet. The five-tuple data comprises a source Internet Protocol (IP) address (abbreviated as source IP address), a destination IP address, a source port, a destination port and a protocol number.
[0045] Exemplarily, before constructing the traffic pattern sample, the original traffic data is pre-processed (for example, missing value processing, repeated value processing, and outlier processing, etc.), to ensure data format consistency. The original traffic data includes normal network traffic data and abnormal network traffic data. The normal network traffic data can be collected by a switch, a router, or other traffic collection devices within a preset time period. The abnormal network traffic data can be obtained by using the publicly available CIC-IDS-2017 dataset (Canadian Institute for Cybersecurity-Intrusion Detection System 2017), the CICDDOS2019 dataset (Canadian Institute for Cybersecurity-Distributed Denial of Service 2019), the MQTT-IoT-IDS 2020 dataset (Message Queuing Telemetry Transport-Internet of Things-Intrusion Detection System 2020), and the CICIoT 2023 dataset (Canadian Institute for Cybersecurity-Internet of Things 2023).
[0046] Specifically, the missing value processing: for the missing values in the data packets of the original traffic data, for example, some protocol fields are missing, and the mean, median or mode is used for interpolation or deletion according to the specific situation. For another example, if the missing proportion of a certain feature is high and does not affect the overall model, the record can be deleted; if the missing proportion is low, the mean or median can be used to fill in according to the distribution of the feature in the same type of data. The repeated value processing: detecting and deleting the completely same repeated records, to avoid bias to the subsequent analysis. The outlier processing: identifying and processing outliers, for example, the data packet size exceeds the reasonable range, and the abnormal data packet can be deleted, truncated or replaced (such as replaced by the median) according to the specific business scenario.
[0047] Exemplarily, the preprocessed data packets are divided into different data flows using the quintuple data in the data packets. One data flow represents one complete Transmission Control Protocol / Internet Protocol (TCP / IP) connection, which contains a plurality of associated data packets. For example, all Synchronize Sequence Number (SYN), Synchronize-Acknowledge Sequence Number (SYN-ACK), Acknowledge Character (ACK) and subsequent data packets of one TCP connection belong to the same data flow.
[0048] The.pcap files of the data flows are taken as inputs, and the headers of the data packets in the data flows are parsed by the libpcap library to extract basic data information. The basic data information includes quintuple data and other basic data (for example, a timestamp). Each identified data flow is managed, and data flow state information is recorded. The data flow state information includes, for example, the start time of the flow, the duration, the total number of bytes, the total number of packets, and the like. When a data flow ends (for example, a TCP connection is closed), feature calculation is triggered. Feature calculation is to calculate various traffic features for each data flow based on all data packets and byte numbers contained in the data flow, and the time sequence of the data packets. These traffic features mainly include statistical features (for example, the average packet size of the data flow, the maximum / minimum packet size of the data flow, the packet rate of the data flow, the byte rate of the data flow, and the like); time-related features (for example, the duration of the data flow, the average / standard deviation of the packet arrival time interval, the time difference between the first packet and the last packet); protocol behavior features (for example, the frequency of connection reset, and the like); and the like.
[0049] Then, the extracted traffic features are output in a labeled Comma-Separated Values (CSV) format to form an original feature matrix. Each row represents an independent data flow, and each column represents an extracted traffic feature. The label is used to distinguish whether the data flow is "normal" traffic or "abnormal" traffic.
[0050] Exemplarily, the original feature matrix is shown in Table 1.
[0051] Table 1
[0052]
[0053]
[0054]
[0055] Exemplarily, since too many features in the original feature matrix will affect the model effect and the calculation efficiency, the original feature matrix is reduced in dimension through a feature selection algorithm to obtain 30 key feature values that can be used for anomaly detection, thereby forming a key feature matrix. The feature selection algorithm can be a random forest, a Spearman correlation coefficient, an Extreme Gradient Boosting (XGBoost) algorithm, or the like.
[0056] In this example, the key feature matrix is shown in Table 2.
[0057] Table 2
[0058]
[0059]
[0060] Exemplarily, the node (symbolized as V) is defined as follows: the source IP address and the destination IP address in the data stream and the corresponding port numbers are taken as the node set of the graph. The edge (symbolized as E) is defined as follows: the communication relationship between the source IP / port and the destination IP / port in the same data stream is defined as an edge. Through the above definition, the five-tuple data and the extracted features in all data streams are mapped to a graph structure G=(V, E). As can be seen, the traffic graph sample is essentially a graph structure composed of multiple nodes and edges connecting two nodes. Based on the above, the discrete network traffic data is converted into structured graph data, which provides the required sample data for the training of the subsequent network anomaly traffic identification model. Since the traffic graph sample contains rich topological structure information and multi-level feature information, it helps the model to more deeply understand and identify complex network anomaly traffic.
[0061] Exemplarily, the edge attribute (edge feature of the edge): the edge feature is obtained through feature vector conversion of the 30-dimensional key feature values in Table 2, and the edge feature of the edge can reflect the overall communication attribute of the data stream.
[0062] Exemplarily, for each node in the graph, the node feature of each node is obtained by aggregating the features of the data stream between all nodes connected to the node, wherein all nodes connected to the node are the neighbor nodes in the application.
[0063] For example, a structure diagram composed of node a and nodes b, c, d connected to node a, where nodes b, c, d are the neighbor nodes of node a. If the aggregated flow (aggregated flow is also referred to as data flow) between node a and node b is a one-row 30-column feature representation, if the aggregated flow between a and node c is a one-row 30-column feature representation, and if the aggregated flow between a and node d is a one-row 30-column feature representation, then the feature representation of node a is a 3-row 30-column feature representation (the feature representation has the same meaning as the node feature), that is, the node feature of node a aggregates the features of the data flow between each node connected thereto.
[0064] For example, the abnormal flow label of a node. If there is at least one data flow marked as abnormal in the communication in which a certain IP address or IP address plus port combination participates, the node can be marked as abnormal flow (that is, the node is marked with an abnormal flow label). Otherwise, it is marked as normal flow.
[0065] According to the above embodiment, the training sample of the network abnormal flow identification model uses a flow graph sample. Based on the node features of each node and the edge features of the edges, a node update operation is performed on each node, and then the abnormal flow prediction information containing the abnormal flow prediction result of each node is obtained. Then, according to the abnormal flow label and the abnormal flow prediction result in the abnormal flow prediction information of each node, the activation layer, the aggregation network, the splicing network and the classification network in the model are updated. Based on the flow graph sample containing the node features, the edge features and the abnormal flow label, the correlation characteristics of the network flow can be more comprehensively captured, and the prediction of the model on the node anomaly is more in line with the actual network situation. In addition, the prediction information is generated by the node update operation, and then the label and the prediction result are combined to update the multiple network levels (activation layer, aggregation network, splicing network and classification network) of the model, so that the feature processing, information aggregation and classification ability of the model can be optimized, the identification accuracy and reliability of the model on the network abnormal flow are improved, and the model can more accurately discover the abnormal flow in actual application.
[0066] In one embodiment, based on the activation layer in the abnormal traffic identification model, the node features of the node and the node features of each of the node's neighboring nodes are processed to obtain the weight factors of the edges between the node and each of the neighboring nodes, including: based on the activation function of the activation layer, processing the product between the node features of the node and the transposed feature representations of the node features of each of its neighboring nodes to obtain a first calculation score of the edge between the node and each of the neighboring nodes; exponentially calculating the first calculation score of the edge between the node and each of the neighboring nodes to obtain a second calculation score of the edge between the node and each of the neighboring nodes; summing the second calculation scores of the edges between the node and each of the neighboring nodes to obtain a third calculation score; and determining the weight factors of the edges between the node and each of the neighboring nodes based on the ratio of the second calculation score of the edge between the node and each of the neighboring nodes to the third calculation score.
[0067] For example, node x v With neighbor node x u The calculation expression of the edge weight factor is as follows:
[0068] Among them, λ uv Represents node x v With neighbor node x u The weight factor of the edge; x v , Represents the node characteristics of each node in the flow graph sample V; Represents the transposed features of the node features of each neighboring node of the node; LeakyReLU() represents the activation function of the activation layer; exp() represents exponentialization.
[0069] For example, the activation function LeakyReLU not only introduces nonlinear characteristics, but also solves the "dead neuron" problem to increase the nonlinear expression ability of the network abnormal traffic identification model. Through the activation function of the activation layer, the product between the node feature of the node and the transposed feature representation of the node feature of each of its neighboring nodes is processed to obtain the first calculation score of the edge between the node and each neighboring node. The first calculation score of the edge between the node and each neighboring node is exponentially obtained to obtain the second calculation score of the edge between the node and each neighboring node. Then, the third calculation score is obtained by summing the second calculation scores of the edges between multiple nodes and each neighboring node. Finally, the ratio of the second calculation score of the edge between the node and each neighboring node to the third calculation score is used as the weight factor of the edge between the node and each neighboring node.
[0070] According to the above-mentioned embodiments, the product of the node feature of the node and the transposed feature of the neighbor node is processed by using the activation function, which can effectively capture the feature correlation between nodes and perform nonlinear conversion, so that the first calculation score is more consistent with the actual feature interaction. The third calculation score obtained by exponentiation and summation is combined with the ratio of the second calculation score to determine the weight factor, which not only considers the relative importance of each edge, but also makes the weight factor more comparable and reasonable through normalization processing, which helps to more accurately aggregate neighbor node information in subsequent node updating operations, improves the ability of the model to describe the relationship between nodes, and further optimizes the accuracy of network abnormal traffic identification. In summary, through the edge score weighting mechanism based on the activation function, the strength of the relationship between nodes is measured, so that the network abnormal traffic identification model can dynamically adjust the information to aggregate the weight, highlight the difference of the abnormal traffic node in the neighbor propagation process, and improve the identification accuracy and generalization ability of the model to abnormal behavior.
[0071] In an embodiment, based on the aggregation network in the abnormal traffic identification model, the node features of each neighbor node, and the weight factors and edge features of the edges between the node and each neighbor node are aggregated to obtain the aggregated flow features from the node to each neighbor node, including: multiplying the node features of the neighbor nodes and the weight factors of the edges between the node and the neighbor nodes to obtain first features; adding the first features and the edge features of the edges between the node and the neighbor nodes to obtain second features; processing the second features based on the aggregation function of the aggregation network to obtain the aggregated flow features from the node to the neighbor nodes.
[0072] For example, assume that there is a node a, which has three neighbor nodes, namely neighbor node b, neighbor node c and neighbor node d, and the edges between the node and each neighbor node are a-b edge, a-c edge and a-d edge. For node a and neighbor node d, the weight factor of a-d edge is multiplied by the node feature of neighbor node d to obtain first features; the first features are added to the edge features of a-d edge to obtain second features. The aggregation function is used to process the second features to obtain the aggregated flow features from node a to neighbor node b.
[0073] For example, as can be seen from the above-mentioned example “for each node in the graph, the node feature of each node is obtained by aggregating the features of the data flow between all nodes connected to the node”, the node features of each node in the traffic structure graph are calculated in the same way.
[0074] Exemplarily, for each neighbor node, a product of a node feature of the neighbor node and a weight factor of an edge between the node and the neighbor node is determined as a first feature, and a sum of the first feature and an edge feature of the edge between the node and the neighbor node is determined as a second feature. The calculation of the first feature and the second feature is repeated for each neighbor node to obtain a plurality of second features. The plurality of second features are processed by an aggregation function of the aggregation network to obtain an aggregated flow feature from the node to the neighbor nodes.
[0075] According to the above embodiment, the weight factor of the edge is combined when calculating the first feature, which reflects the importance of the neighbor node to the node, and can enhance the information propagation of valuable neighbor nodes and suppress noise. Some abnormal connections are difficult to identify by the node feature alone (such as port scanning, short connection burst, etc.), and the introduction of the edge feature (i.e., summing the edge feature and the first feature) can improve the accuracy of anomaly detection. In summary, the present embodiment realizes the joint modeling of the semantic information, structural features and behavior patterns of the neighbor nodes by multiplying the node feature of the neighbor node with the edge weight factor, fusing the edge feature and processing by the aggregation function, effectively improving the identification accuracy, feature expression ability and robustness of the network anomaly traffic identification model for network anomaly traffic, and is particularly suitable for anomaly detection tasks in complex attack scenarios.
[0076] In an embodiment, based on the splicing network in the anomaly traffic identification model, the node feature of a node and the aggregated flow features from the node to each neighbor node are spliced, and based on the splicing result, the node feature of the node is updated, including: splicing the node feature of the node and the aggregated flow features from the node to each neighbor node to obtain a first splicing feature; processing the first splicing feature based on a weight parameter matrix of the splicing network to obtain a second splicing feature; processing the second splicing feature based on a nonlinear activation function of the splicing network to obtain a splicing result; and updating the node feature of the node based on the splicing result.
[0077] Exemplarily, assuming that there is now a node q, the node q has three neighbor nodes, namely a neighbor node w, a neighbor node e and a neighbor node r, and the aggregated flow features of the node q and each neighbor node are calculated based on the foregoing steps, which are aggregated flow feature q-w, aggregated flow feature q-e and aggregated flow feature q-r respectively. For the node q, the node feature of the node q, the aggregated flow feature q-w, the aggregated flow feature q-e and the aggregated flow feature q-r are spliced to obtain a first splicing feature. The product of the weight parameter matrix and the first splicing feature is taken as a second splicing feature. Then, the second splicing feature is processed by the nonlinear activation function of the splicing network to obtain a splicing result. Finally, the node feature of q is updated with the splicing result.
[0078] According to the above-mentioned embodiments, the node feature of the node itself is spliced with the aggregated flow features from the node to each neighbor node to form first spliced features; the first spliced features are processed by using a weight parameter matrix of the splicing network to obtain second spliced features; the second spliced features are operated by using a nonlinear activation function of the splicing network to obtain a splicing result; and finally, the node feature of the node is updated according to the splicing result. By splicing the node feature of the node itself and the neighbor aggregated flow features, the node self information and the neighbor association information can be retained at the same time, and the effective fusion of multi-dimensional features is realized. By processing the spliced features by using the weight matrix and the nonlinear activation function of the splicing network, the expression ability of the features can be enhanced, and the deep association between the features can be mined. Finally, the node feature is updated based on the processed splicing result, so that the node feature can more comprehensively and accurately reflect the comprehensive information of the node and its surrounding relationship, and provide a more reliable feature basis for subsequent model calculation and prediction.
[0079] In an embodiment, based on the classification network in the abnormal flow identification model, the updated node feature of the node is classified and predicted to obtain an abnormal flow prediction result of the node, including: calculating the product of the classification parameter matrix of the classification network and the node feature of the node to obtain a third feature; summing the bias matrix of the classification network and the third feature to obtain an abnormal flow prediction score of the node; and determining the abnormal flow prediction result of the node based on the abnormal flow prediction score of the node.
[0080] For example, the classification network can be expressed by a function expression as follows: score v =W,·x v +b;
[0081] Wherein, score v represents the abnormal flow prediction score of the node; W, represents the classification parameter matrix of the classification network; x v represents the node feature matrix of the node; and b represents the bias matrix value of the classification network.
[0082] For example, each node can obtain a node feature matrix x v composed of the node features of each node by continuously iteratively updating the node according to the foregoing node updating operation. It should be noted that the node features in the x v are the node features after the node is updated, which are different from the node features of the node mentioned above, which are obtained based on the features of the data flow aggregated between all nodes connected to the node. Therefore, it can be understood that the node features in the foregoing embodiments are the node features before the node is updated, and the node features in the present embodiment are the node features after the node is updated.
[0083] According to the above embodiment, the node features of the node are linearly transformed with a classification parameter matrix in the classification network to obtain third features of the node; the third features are subjected to an addition operation with a bias matrix of the classification network (specifically, a bias matrix value of the classification network) to obtain an abnormal traffic prediction score corresponding to the node; wherein the classification parameter matrix and the bias term can be trained by back propagation and have automatic learning ability. Finally, the abnormal traffic prediction result of the node is determined based on the abnormal traffic prediction score of the node (specifically, by comparing the prediction score with a preset score threshold, it can be determined whether the node is an abnormal traffic node).
[0084] In an embodiment, based on the abnormal traffic prediction score of the node, the abnormal traffic prediction result of the node is determined, including: in the case that the abnormal traffic prediction score of the node is less than a preset score threshold, the node is determined to be normal traffic; in the case that the abnormal traffic prediction score of the node is greater than or equal to the preset score threshold, the node is determined to be abnormal traffic.
[0085] Exemplarily, assuming that the preset score threshold is 0.5, the abnormal traffic prediction score of the existing node a is 0.3, the abnormal traffic prediction score of the node b is 0.15, and the abnormal traffic prediction score of the node c is 0.8. Comparing the abnormal traffic prediction scores of the nodes with the preset score threshold, it can be determined that the nodes a and b are normal traffic, while the node c is determined to be abnormal traffic because the abnormal traffic prediction score of the node c is greater than the preset score threshold.
[0086] According to the above embodiment, by setting the prediction score threshold and classifying whether the node is abnormal according to the prediction score threshold, an abnormal node binary classification mechanism based on the graph neural network is realized. The mechanism has the advantages of simple judgment, efficient deployment, easy interpretation, etc., and is suitable for real-time monitoring and early warning of potential threat nodes in a complex network environment.
[0087] In an embodiment, based on the abnormal traffic label and the abnormal traffic prediction result of each node, the activation layer, the aggregation network, the splicing network and the classification network in the network abnormal traffic identification model are updated, including: based on the abnormal traffic label and the abnormal traffic prediction result of each node, a loss function is determined; based on the gradient of the loss function, the activation layer, the aggregation network, the splicing network and the classification network in the network abnormal traffic identification model are updated.
[0088] Exemplarily, the expression of the loss function is as follows:
[0089] Wherein, loss(Y, T) represents the loss between the abnormal traffic label of the node and the abnormal traffic prediction result; C represents the number of classes of the classifier; T represents the true classification label, i.e. the abnormal traffic label of the node; T iThe i-th anomaly traffic label of the node; Y represents the anomaly traffic prediction result; Y j The j-th anomaly traffic prediction result of the node.
[0090] Exemplarily, the update formula of the gradient descent algorithm is:
[0091]
[0092] Wherein, θ new The updated network anomaly traffic is activated in the model, aggregated network, splicing network and classification network; θ old The activated layer, aggregation network, splicing network and classification network in the network anomaly traffic identification model; η is the learning rate.
[0093] According to the above embodiment, according to the anomaly traffic true label of each node (that is, the anomaly traffic label of each node) and the anomaly traffic prediction result output by the model, the loss function is calculated to measure the prediction error; Based on the loss function, the model is back propagated, and the gradient value of each network module is calculated; and the parameters of the activated layer, aggregation network, splicing network and classification network in the anomaly traffic identification model are updated accordingly, thereby improving the overall identification performance of the model.
[0094] In an embodiment, the network traffic data collected within a period of time is constructed into a network traffic structure diagram. In the feature engineering stage, the node features of each node in the network traffic structure diagram are extracted. The extracted node features are input into the network anomaly traffic identification model, which iteratively generates node features based on the importance of node neighborhood. Finally, the anomaly traffic prediction score of each node in the network traffic structure diagram is calculated, and the anomaly traffic prediction score of each node is determined based on the anomaly traffic prediction score of each node. If the node is abnormal traffic, the attack path is traced by analyzing the neighborhood association and topology, and the source and propagation path of the abnormal traffic are further analyzed, so as to take corresponding protection measures. In addition, the result can also trigger a linkage response, for example, blocking the attack source network connection, limiting the upper limit of the traffic transmission bandwidth, intelligent alarm, etc., combined with the threat intelligence library to perform precise access control, forming a closed-loop security system of "detection-tracing-protection".
[0095] Figure 2 The structure block diagram of the training device of the network anomaly traffic identification model of an embodiment of the present application.
[0096] As Figure 2 shown, the training device of the network anomaly traffic identification model can include:
[0097] The traffic graph sample acquisition module 510 is configured to acquire traffic graph samples of the network anomaly traffic identification model, wherein the traffic graph samples include node features of each node, and an abnormal traffic label, and edge features of edges between each node and its neighbor nodes;
[0098] The node updating module 520 is configured to perform the following node updating operation on each of the nodes:
[0099] The weight factor calculation unit 521 is configured to process the node features of the node and the node features of each neighbor node of the node based on an activation layer in the abnormal traffic identification model, to obtain weight factors of edges between the node and each of the neighbor nodes;
[0100] The aggregation processing unit 522 is configured to perform aggregation processing on the node features of each of the neighbor nodes, and the weight factors and the edge features of the edges between the node and each of the neighbor nodes based on an aggregation network in the abnormal traffic identification model, to obtain aggregated flow features from the node to each of the neighbor nodes;
[0101] The splicing unit 523 is configured to splice the node features of the node and the aggregated flow features from the node to each of the neighbor nodes based on a splicing network in the abnormal traffic identification model, and update the node features of the node based on a splicing result.
[0102] The classification prediction unit 524 is configured to perform classification prediction on the updated node features of the node based on a classification network in the abnormal traffic identification model, to obtain an abnormal traffic prediction result of the node.
[0103] The updating module 530 is configured to update the activation layer, the aggregation network, the splicing network and the classification network in the network anomaly traffic identification model based on the abnormal traffic label and the abnormal traffic prediction result of each of the nodes, in a case where the node updating operation has been performed on each of the nodes to obtain the abnormal traffic prediction information of each of the nodes.
[0104] In an embodiment, as shown in Figure 3 The weight factor calculation unit 521 includes:
[0105] The activation function sub-unit 5211 is configured to process a product between the node features of the node and a transposed feature representation of the node features of each of the neighbor nodes based on an activation function of the activation layer, to obtain a first calculation score of the edge between the node and each of the neighbor nodes.
[0106] an exponentiation subunit 5212, configured to exponentiate the first calculated score of the edge between the node and each of the neighbor nodes to obtain a second calculated score of the edge between the node and each of the neighbor nodes;
[0107] a summation subunit 5213, configured to sum the second calculated scores of the edges between the node and each of the neighbor nodes to obtain a third calculated score;
[0108] a determination subunit 5214, configured to determine a weight factor of the edge between the node and each of the neighbor nodes based on a ratio of the second calculated score of the edge between the node and each of the neighbor nodes to the third calculated score.
[0109] In an embodiment, as shown in Figure 4 the aggregation processing unit 522 includes:
[0110] a first feature calculation subunit 5221, configured to calculate a node feature of the neighbor node, and multiply the node feature by a weight factor of the edge between the node and the neighbor node to obtain a first feature;
[0111] a second feature calculation subunit 5222, configured to add the first feature to an edge feature of the edge between the node and the neighbor node to obtain a second feature;
[0112] an aggregation function processing subunit 5223, configured to process the second feature based on an aggregation function of the aggregation network to obtain an aggregated flow feature from the node to the neighbor node.
[0113] In an embodiment, as shown in Figure 5 the concatenation unit 523 includes:
[0114] a first concatenation feature subunit 5231, configured to concatenate a node feature of the node to aggregated flow features from the node to each of the neighbor nodes to obtain a first concatenation feature;
[0115] a second concatenation feature subunit 5232, configured to process the first concatenation feature based on a weight parameter matrix of the concatenation network to obtain a second concatenation feature;
[0116] a concatenation result subunit 5233, configured to process the second concatenation feature based on a nonlinear activation function of the concatenation network to obtain the concatenation result;
[0117] a node feature updating subunit 5234, configured to update the node feature of the node based on the concatenation result.
[0118] In an embodiment, the classification prediction unit includes:
[0119] a third feature calculation subunit configured to calculate a product of a classification parameter matrix of the classification network and the node feature of the node to obtain a third feature;
[0120] an abnormal traffic prediction score subunit configured to sum a bias matrix of the classification network and the third feature to obtain an abnormal traffic prediction score of the node;
[0121] an abnormal traffic prediction result subunit configured to determine an abnormal traffic prediction result of the node based on the abnormal traffic prediction score of the node.
[0122] In an implementation, the abnormal traffic prediction result subunit is specifically configured to:
[0123] determine that the node is normal traffic when the abnormal traffic prediction score of the node is less than a preset score threshold;
[0124] determine that the node is abnormal traffic when the abnormal traffic prediction score of the node is greater than or equal to the preset score threshold.
[0125] In an implementation, the update module is specifically configured to:
[0126] determine a loss function based on the abnormal traffic label and the abnormal traffic prediction result of each node;
[0127] update the activation layer, the aggregation network, the concatenation network and the classification network in the network abnormal traffic identification model based on a gradient of the loss function.
[0128] The specific functions and examples of the modules and submodules of the system of the embodiments of the present application are described in the related description of the corresponding steps in the above method embodiments, which will not be described here.
[0129] In the technical solution of the present application, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0130] The embodiments of the present application also provide a network abnormal traffic identification model training system, comprising:
[0131] at least one processor; and a memory connected in communication with the at least one processor;
[0132] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of the embodiments of the present application.
[0133] The training system of the network abnormal traffic identification model has the same beneficial effects as the training method of the network abnormal traffic identification model, which will not be repeated here.
[0134] The application also provides a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are used to make a computer execute the method in any one of the embodiments of the application.
[0135] The storage medium has the same beneficial effects as the training method of the network abnormal traffic identification model, which will not be repeated here.
[0136] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the application is shown. The electronic device 800 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device 800 can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the application described and / or claimed in this document.
[0137] As shown in Figure 6 The electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0138] Various components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc., an output unit 807, such as various types of displays, a speaker, etc., the storage unit 808, such as a magnetic disk, an optical disk, etc., and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0139] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as the training method of the network anomaly traffic identification model. For example, in some embodiments, the training method of the network anomaly traffic identification model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the training method of the network anomaly traffic identification model described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the training method of the network anomaly traffic identification model by any other suitable means, such as by means of firmware.
[0140] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0141] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0142] In the context of the present application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0143] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0144] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0145] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0146] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the flow. For example, the steps described in the present application can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present application can be achieved, which is not limited herein.
[0147] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for training a network anomaly traffic identification model, characterized in that, The method comprises: obtaining a traffic pattern sample of a network anomaly traffic identification model, wherein the traffic pattern sample comprises node features of each node and an anomaly traffic label, and edge features of edges between each node and its neighbor nodes; performing the following node update operation for each node: processing the node features of the node and the node features of each neighbor node of the node based on an activation layer in the anomaly traffic identification model to obtain a weight factor of the edge between the node and each neighbor node; performing aggregation processing on the node features of each neighbor node, the weight factor of the edge between the node and each neighbor node, and the edge features of the edge between the node and each neighbor node based on an aggregation network in the anomaly traffic identification model to obtain aggregated flow features from the node to each neighbor node; splicing the node features of the node and the aggregated flow features from the node to each neighbor node based on a splicing network in the anomaly traffic identification model, and updating the node features of the node based on a splicing result; and performing classification prediction on the updated node features of the node based on a classification network in the anomaly traffic identification model to obtain an anomaly traffic prediction result of the node; in a case where the node update operation has been performed for each node to obtain anomaly traffic prediction information of each node, updating the activation layer, the aggregation network, the splicing network, and the classification network in the network anomaly traffic identification model based on the anomaly traffic label and the anomaly traffic prediction result of each node.
2. The method of claim 1, wherein, The processing of the node features of the node and the node features of each neighbor node of the node based on the activation layer in the anomaly traffic identification model to obtain a weight factor of the edge between the node and each neighbor node comprises: processing a product between a transpose feature representation of the node features of the node and the node features of each neighbor node of the node based on an activation function of the activation layer to obtain a first calculation score of the edge between the node and each neighbor node; exponentiating the first calculation score of the edge between the node and each neighbor node to obtain a second calculation score of the edge between the node and each neighbor node; summing the second calculation score of the edge between the node and each neighbor node to obtain a third calculation score; determining the weight factor of the edge between the node and each neighbor node based on a ratio of the second calculation score of the edge between the node and each neighbor node to the third calculation score.
3. The method of claim 1, wherein, The aggregation processing on the node features of each neighbor node, the weight factor of the edge between the node and each neighbor node, and the edge features of the edge between the node and each neighbor node based on the aggregation network in the anomaly traffic identification model to obtain aggregated flow features from the node to each neighbor node comprises: multiplying the node features of the neighbor node and the weight factor of the edge between the node and the neighbor node to obtain a first feature; adding the first feature and the edge features of the edge between the node and the neighbor node to obtain a second feature; Based on an aggregation function of the aggregation network, the second feature is processed to obtain an aggregated flow feature from the node to the neighbor node.
4. The method of claim 1, wherein, Based on a concatenation network in the abnormal flow identification model, the node feature of the node and the aggregated flow feature from the node to each of the neighbor nodes are concatenated, and based on a concatenation result, the node feature of the node is updated, including: The node feature of the node and the aggregated flow feature from the node to each of the neighbor nodes are concatenated to obtain a first concatenation feature; Based on a weight parameter matrix of the concatenation network, the first concatenation feature is processed to obtain a second concatenation feature; Based on a nonlinear activation function of the concatenation network, the second concatenation feature is processed to obtain the concatenation result; Based on the concatenation result, the node feature of the node is updated.
5. The method of claim 1, wherein, Based on a classification network in the abnormal flow identification model, the updated node feature of the node is classified and predicted to obtain an abnormal flow prediction result of the node, including: A product of a classification parameter matrix of the classification network and the node feature of the node is calculated to obtain a third feature; A bias matrix of the classification network and the third feature are summed to obtain an abnormal flow prediction score of the node; Based on the abnormal flow prediction score of the node, the abnormal flow prediction result of the node is determined.
6. The method of claim 5, wherein, Based on the abnormal flow prediction score of the node, the abnormal flow prediction result of the node is determined, including: In a case where the abnormal flow prediction score of the node is less than a preset score threshold, the node is determined to be normal flow; In a case where the abnormal flow prediction score of the node is greater than or equal to the preset score threshold, the node is determined to be abnormal flow.
7. The method of claim 1, wherein, Based on the abnormal flow label and the abnormal flow prediction result of each of the nodes, an activation layer, an aggregation network, a concatenation network and a classification network in the network abnormal flow identification model are updated, including: Based on the abnormal flow label and the abnormal flow prediction result of each of the nodes, a loss function is determined; Based on a gradient of the loss function, the activation layer, the aggregation network, the concatenation network and the classification network in the network abnormal flow identification model are updated. 8.A device for training a network anomaly traffic identification model, characterized in that, Including: A traffic graph sample acquisition module is configured to acquire a traffic graph sample of a network abnormal flow identification model, wherein the traffic graph sample includes a node feature and an abnormal flow label of each node, and an edge feature of an edge between each of the nodes and a neighbor node of the node; A node updating module is configured to perform the following node updating operation on each of the nodes: A weight factor calculation unit is configured to process, based on an activation layer in the abnormal flow identification model, the node feature of the node and the node feature of each of the neighbor nodes of the node to obtain a weight factor of an edge between the node and each of the neighbor nodes. an aggregation processing unit, configured to perform aggregation processing on node features of each of the neighbor nodes and weight factors and edge features of edges between the node and each of the neighbor nodes based on an aggregation network in the abnormal traffic identification model, to obtain aggregated flow features from the node to each of the neighbor nodes; a concatenation unit, configured to perform concatenation on the node features of the node and the aggregated flow features from the node to each of the neighbor nodes based on a concatenation network in the abnormal traffic identification model, and update the node features of the node based on a concatenation result; and a classification prediction unit, configured to perform classification prediction on the updated node features of the node based on a classification network in the abnormal traffic identification model, to obtain an abnormal traffic prediction result of the node. an updating module, configured to update the activation layer, the aggregation network, the concatenation network and the classification network in the network abnormal traffic identification model based on the abnormal traffic labels and the abnormal traffic prediction results of each of the nodes, in a case where the node updating operation has been performed on each of the nodes to obtain the abnormal traffic prediction information of each of the nodes. 9.A system for training a network anomaly traffic identification model, the system comprising: comprise: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-7.