Transverse movement flow detection method
By extracting clear flow characteristics and topological structures in a containerized environment, combining multi-stage detection and pre-training models, the problems of unclear lateral movement detection characteristics and difficult detection are solved, and lateral movement flow detection with high accuracy and low false alarm rate are achieved.
Patent Information
- Application Number
- CN202510106905.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In a containerized environment, the lateral movement detection characteristics are unclear, and the classification of the detection tasks is difficult and the samples are unbalanced, resulting in low detection rates and high false alarm rates.
A lateral movement flow detection method is proposed. The target features and topological structure are extracted through the pre-processing stage, the first detection stage is used to perform maximum value detection and topological detection, and the second detection stage is used to perform traffic detection using a pre-trained lateral movement detection model, and finally deduplicate the detection result.
By extracting clear packet-level and session-level traffic characteristics, combined with topology detection and pre-training models, the accuracy and efficiency of lateral mobile traffic detection are improved, the false alarm rate is reduced and the detection rate is improved.
Smart Images

Figure CN120017340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and in particular to a lateral movement traffic detection technology in the field of network security, and more particularly to a lateral movement traffic detection method. Background Art
[0002] With the rapid development of containerization technology, containerized clusters are increasingly used in various fields. However, the widespread use of this technology has also attracted the attention of attackers, making it particularly important to detect attack behaviors in a timely manner. Attacks faced by containerized clusters are usually divided into five stages: reconnaissance, foothold, lateral movement, attack, maintaining access, and cleanup. Among these stages, lateral movement is the key stage for attack behavior detection and the most feasible stage.
[0003] Since containerized environments contain multiple components and have a wide range of attack surfaces, defects in any component may become a security risk in the containerized environment and bring security challenges. Therefore, in a containerized environment, it is necessary to obtain network traffic to perform lateral movement detection and thus implement attack behavior detection.
[0004] At present, there are still two problems in lateral movement detection in a containerized environment. On the one hand, the features that can be used for lateral movement detection are unclear, making it difficult to implement lateral movement detection. On the other hand, the classification of lateral movement detection tasks is difficult and the samples are unbalanced, making it difficult to achieve the requirements of high detection rate and low false alarm rate.
[0005] It should be noted that this background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily the prior art. If there is no evidence that the relevant information has been disclosed before the application date of the present invention, the relevant information shall not be regarded as the prior art. Summary of the invention
[0006] Therefore, the purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for detecting lateral movement flow.
[0007] The objectives of the present invention are achieved through the following technical solutions.
[0008] According to a first aspect of the present invention, a lateral movement traffic detection method is provided, which is used to detect whether there is lateral movement traffic in a traffic sequence to be detected generated in a target containerized cluster, wherein the traffic sequence to be detected includes multiple network flows that are continuous in time sequence, and the lateral movement traffic is the traffic data generated when the target containerized cluster is attacked. The method comprises: a preprocessing stage: obtaining a benign traffic sequence containing multiple benign flows generated by the target containerized cluster, performing feature extraction on the benign traffic sequence to obtain a target feature set containing multiple target features, and traversing each target feature in each benign flow in the benign traffic sequence to obtain a value range corresponding to each target feature, and performing topology extraction based on the benign traffic sequence to obtain a topological structure of the target containerized cluster; a first detection stage: maximum value detection: analyzing whether each network flow in the traffic sequence to be detected is lateral movement traffic; wherein each network flow includes multiple target features. The target feature is a network flow whose value is not within the value range of the target feature. The network flow is a lateral movement flow. All lateral movement flows obtained by the maximum value detection of the flow sequence to be detected constitute the first detection result. Topology detection: analyze whether the transmission process of each network flow in the flow sequence to be detected satisfies the topology structure of the target containerized cluster. If not, the network flow is a lateral movement flow. Among them, all lateral movement flows obtained by the topology detection of the flow sequence to be detected constitute the second detection result. The second detection stage: use the pre-trained lateral movement detection model to perform flow detection on the flow sequence to be detected to analyze the lateral movement flow in the flow sequence to be detected to obtain the third detection result. The detection result output stage: after removing the network flow that is repeatedly judged as lateral movement flow in the first detection result, the second detection result and the third detection result, all the remaining lateral movement flows are used as the final detection result of the flow sequence to be detected.
[0009] In some embodiments of the present invention, the method includes extracting features from a benign traffic sequence in the following manner to obtain a target feature set containing multiple target features: each benign traffic in the benign traffic sequence includes multiple data packets, and based on all data packets of any benign traffic, packet-level feature extraction is performed on the benign traffic to obtain multiple packet-level traffic features corresponding to the benign traffic; and based on a packet connection protocol, all data packets in the benign traffic are aggregated into multiple sessions to extract multiple session-level traffic features corresponding to the benign traffic; wherein all packet-level traffic features and all session-level traffic features constitute an initial feature set; a preset evaluation method is used to perform importance evaluation on each traffic feature in the initial feature set to obtain an importance evaluation result for each traffic feature; the importance evaluation results are sorted in descending order, and a preset number of traffic features ranked first are selected as target features.
[0010] In some embodiments of the present invention, the preset evaluation method is: using gradient boosting decision tree, random forest and mutual information to evaluate the importance of each traffic feature in the initial feature set, and calculating the mean of the importance of each traffic feature to obtain the importance evaluation result of each traffic feature.
[0011] In some embodiments of the present invention, the preset number is 57.
[0012] In some embodiments of the present invention, the target features include: standard deviation of continuous idle time, average number of bytes transmitted by forward bulk data packets, minimum forward data packet length, average forward bulk data packet rate, forward RST flag, number of forward active data packets, minimum backward inter-frame arrival time, backward PSH flag, minimum forward segment size, initial value of backward window, total TCP flow time, PSH flag count, ACK flag count, average forward inter-frame arrival time, FIN flag count, backward header length, minimum inter-flow arrival time, minimum forward inter-frame arrival time, number of forward data packets, average time interval between data packet arrivals, total length of backward data packets, SYN flag count, maximum data packet length, average backward inter-frame arrival time, ratio of backward to forward data packets, number of backward data packets, forward header length, flow duration in milliseconds, number of forward packets, Sum of packet arrival time intervals, Sum of backward packet arrival time intervals, Mean value of forward packet length, Mean value of forward segment size, Forward PSH flag, Standard deviation of packet arrival time intervals, Standard deviation of forward packet arrival time intervals, Number of bytes of backward subflow, Forward window initial value, Standard deviation of backward packet length, Mean value of packet length, Mean value of backward packet length, Packet length variation, Standard deviation of backward packet arrival time intervals, Maximum value of packet arrival time intervals, Mean value of backward segment size, Maximum value of backward packet arrival time intervals, Maximum value of forward packet arrival time intervals, Standard deviation of packet length, Mean value of packet length, Forward flow rate, Total length of forward packets, Standard deviation of forward packet length, Number of bytes of forward subflow, Total length of backward packets, Packet flow rate, Byte flow rate, Backward packet flow rate and Maximum value of forward packet length.
[0013] In some embodiments of the present invention, the method includes performing topology extraction based on a benign traffic sequence to obtain a topological structure of a target containerized cluster in the following manner: each benign traffic in the benign traffic sequence corresponds to a source node and a target node; wherein the source node represents a node that sends the benign traffic in the target containerized cluster, and the target node represents a node that receives the benign traffic in the target containerized cluster; based on the node types of the source node and the target node corresponding to each benign traffic in the benign traffic sequence, a load node set, an API server set, a load node set that communicates with the API server, and a load node set that communicates with external nodes are obtained; wherein the load node set, the API server set, the load node set that communicates with the API server, and the load node set that communicates with external nodes constitute the topological structure of the target containerized cluster.
[0014] In some embodiments of the present invention, the method includes analyzing whether the transmission process of each network flow in the flow sequence to be detected satisfies the topological structure of the target containerized cluster in the following manner: if the source node or the target node corresponding to the network flow is a new load node and is not in the load node set, then the network flow is lateral movement flow; if the source node corresponding to the network flow is not in the load node set communicating with the API server, and the target node corresponding to the network flow is the API server, then the network flow is lateral movement flow; if the target node corresponding to the network flow is not in the load node set communicating with the API server, and the source node corresponding to the network flow is the API server, then the network flow is lateral movement flow; if the source node corresponding to the network flow is not in the load node set communicating with the external node, and the target node corresponding to the network flow is an external node, then the network flow is lateral movement flow; if the target node corresponding to the network flow is not in the load node set communicating with the external node, and the source node corresponding to the network flow is an external node, then the network flow is lateral movement flow.
[0015] In some embodiments of the present invention, a pre-trained lateral movement detection model is configured to perform traffic detection on a traffic sequence to be detected in the following manner: taking the traffic sequence to be detected as input, generating a target prediction sequence of the traffic sequence to be detected in a recursive prediction manner; wherein the target prediction sequence includes multiple target predicted network flows, and each target predicted network flow corresponds to a network flow in the traffic to be detected; calculating the error value between each target predicted network flow and its corresponding actual network flow, and if the error value is greater than or equal to a threshold, determining that the network flow is lateral movement traffic.
[0016] In some embodiments of the present invention, the pre-trained lateral movement detection model is a model obtained by training in the following manner: obtaining a training benign traffic sequence, wherein the training benign traffic sequence includes a plurality of benign flows that are continuous in time series, and each benign flow includes a plurality of traffic features; performing feature extraction on the training benign traffic sequence to obtain a target feature set corresponding to each benign flow in the training benign traffic sequence; taking the benign traffic sequence after feature extraction processing as input and the target prediction sequence corresponding to the benign traffic sequence after feature extraction processing as output, performing multiple rounds of iterative training until the lateral movement detection model converges.
[0017] In some embodiments of the present invention, the pre-trained lateral movement detection model includes a Transformer model and a judgment module, wherein: the Transformer model is used to predict a target prediction sequence corresponding to a traffic sequence to be detected; the judgment module is used to calculate the error value between each target predicted network traffic and its corresponding actual network traffic, and when the error value is greater than or equal to a threshold, the network traffic is judged to be lateral movement traffic.
[0018] In some embodiments of the present invention, the pre-trained lateral movement detection model includes a spatial feature embedding module, a temporal feature embedding module, a coding and decoding module, a sequence output module and a judgment module; wherein: the spatial feature embedding module is used to extract the spatial features of the traffic sequence to be detected to obtain the spatial feature embedding vector of each network traffic in the traffic sequence to be detected, and connect the spatial feature embedding vector of each network traffic with the traffic features of the network traffic to obtain the initial features of the traffic sequence to be detected; the temporal feature embedding module is used to extract the temporal features based on the initial features of the traffic sequence to be detected to obtain the temporal feature embedding vector of the traffic sequence to be detected; the coding and decoding module includes an encoder, a first decoder and a second decoder, wherein: the encoder is used to extract the dependency relationship between each network traffic in the traffic sequence to be detected based on the initial features of the traffic sequence to be detected to obtain the first latent vector of the traffic sequence to be detected; and each initial predicted network traffic obtained based on the initial features of the traffic sequence to be detected and the sequence output module and the actual network traffic in the traffic sequence to be detected corresponding to itself The first decoder is used to perform decoding processing based on the first latent vector of the traffic sequence to be detected and the time feature embedding vector to obtain the first decoding feature vector of the traffic sequence to be detected; the second decoder is used to perform decoding processing based on the second latent vector of the traffic sequence to be detected and the time feature embedding vector to obtain the second decoding feature vector of the traffic sequence to be detected; the sequence output module is used to generate an initial prediction sequence of the traffic sequence to be detected based on the first decoding feature vector of the traffic sequence to be detected, and to generate a target prediction sequence of the traffic sequence to be detected based on the second decoding feature vector of the traffic sequence to be detected; wherein the initial prediction sequence includes multiple initial predicted network flows, and each initial predicted network flow corresponds to a network flow in the traffic sequence to be detected; the judgment module is used to calculate the error value between each target predicted network flow and its corresponding actual network flow, and when the error value is greater than or equal to a threshold value, the network flow is judged to be a lateral movement flow.
[0019] Compared with the prior art, the advantages of the present invention are: (1) extracting packet-level traffic features and session-level traffic features, and eliminating traffic features with lower importance based on the importance ranking results of the traffic features, thereby solving the problem of unclear lateral movement detection features in containerized clusters; (2) setting up two-stage detection to screen lateral movement traffic, thereby improving the detection rate and reducing the false alarm rate; (3) the lateral movement detection model uses two-step prediction to minimize the prediction error of benign traffic and maximize the prediction error of lateral movement traffic, thereby achieving more accurate lateral movement traffic detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:
[0021] Figure 1 A schematic diagram of a flow chart of a method for detecting lateral movement traffic according to an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of ranking the importance of traffic features according to an embodiment of the present invention;
[0023] Figure 3 is a box plot of lateral movement traffic under active and idle class features according to an embodiment of the present invention;
[0024] Figure 4 A data transmission density diagram of network traffic under length characteristics according to an embodiment of the present invention;
[0025] Figure 5 is a lateral movement detection model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0027] In order to better understand the present invention, the existing lateral movement detection method is briefly introduced below.
[0028] Existing technologies usually detect lateral movement from four aspects: network traffic, endpoint behavior, user behavior, and threat intelligence. Among them, detection based on network traffic: lateral movement detection can be performed through abnormal port scanning, large amounts of data transmission, and communication using non-standard protocols; however, this detection method only focuses on the network traffic generated by user login activities and is not suitable for containerized clusters. Detection based on endpoint behavior: detection can be performed through abnormal process creation, abnormal file access, abnormal registry modification, etc.; for example, a network connection diagram can be constructed using Windows systems and security events. When remote file execution occurs in the network (conventional detection scenario) or an attacked node is detected (forensic analysis scenario), the constructed network connection diagram can be used to find possible lateral movement paths; however, this detection method is deeply bound to the Windows system and cannot be used for lateral movement detection in containerized clusters. Detection based on user behavior: Detection can be performed by logging into abnormal accounts, accessing abnormal resources, performing abnormal operations, etc. For example, by modeling user operation logs (such as logging in, accessing web pages, sending emails, opening files, etc.) into graphs and generating embedding vectors for each node, smaller clusters are determined by clustering and the determined smaller clusters are identified as abnormal operations; however, in containerized clusters, there is no user behavior information available, and there is no relevant user behavior information in TCP / IP traffic, so this detection method is not suitable for containerized clusters. Detection based on threat intelligence: Attackers usually use known tools and techniques to move laterally, so by analyzing threat intelligence, tools and techniques that attackers may use can be identified and defenses can be carried out in advance; however, this detection method requires known existing threats, and the timeliness of lateral movement detection is weak.
[0029] Although lateral movement detection can be achieved in different ways, existing lateral movement detection methods mainly implement lateral movement detection through network traffic. These methods only focus on the network traffic generated by login activities during detection, and these methods are based on a key assumption when performing detection: attackers usually perform lateral movement to access machines that the initial victim cannot access. At the same time, the detection models used by these methods are different, but they all adopt the method of modeling network traffic logs into login graphs, where the nodes of the login graph represent computers and the edges represent the traffic generated between computers. Since these methods use different data sets and evaluation indicators, and all perform lateral movement detection in non-containerized environments such as enterprise networks, these methods are not suitable for containerized clusters.
[0030] In order to better understand the present invention, the following briefly introduces the existing lateral movement detection method based on network traffic.
[0031] Some researchers have proposed to extract the features of the login graph and then use the Logistic regression model to detect lateral movement. The features extracted from the login graph include the number of nodes, the number of edges, the graph density, etc. Although this detection method can detect lateral movement, the feature extraction method used in this method seems to be very old in today's popular graph neural network. It is cumbersome to implement and the effect is not good.
[0032] Some researchers have proposed to construct a triple ⟨𝑈,𝑆,𝐷> by statistically mining several most frequently communicating users, source computers, and target computers through the login graph. For a login record ⟨𝑢,𝑠,d>, if the login record meets one of the triples, namely 𝑢∈ 𝑈,𝑠∈ 𝑆,d∈ 𝐷, then the login is considered normal; otherwise, it is considered to be lateral movement. The advantage of this detection method is that it is highly interpretable; the disadvantage is that its triple mining method is inefficient and requires enumerating all possible ⟨𝑈,𝑆,𝐷>, which takes a long time.
[0033] Some researchers have also proposed a detection method based on machine learning. This method uses the Node2vec model to learn the login graph and generate an embedding vector for each node. For each login record, the element-by-element product of the two nodes in the login record is calculated and the probability of the login record being a lateral movement is predicted through the Logistic model. The advantage of this detection method is that it uses a machine learning method to extract graph features more effectively than previous methods; the disadvantage is that it regards the login graph as static and fails to reflect the dynamic nature of the network. For computers that were originally normal but later infected, it is difficult to detect their lateral movement behavior.
[0034] Other researchers have conducted in-depth research on the causal relationship between paths in the login graph and defined a method for aggregating the causal relationship between login records, that is, for two login records 𝐿1=⟨𝑠1,d1> and 𝐿2=⟨𝑠2,d2>, if 𝐿1 occurs within 24 hours before 𝐿2, and the destination computer of 𝐿1 is equal to the source computer of 𝐿2, that is, d1=𝑠2, then 𝐿1 and 𝐿2 are "causally related", and 𝐿1 is the cause of 𝐿2. Multiple causal login records are aggregated into login paths through the above rules, and then lateral movement detection is achieved by analyzing the login paths. In this detection method, lateral movement detection is performed by checking the user information on the login path. If the user information has changed, it is considered lateral movement; if the user information has not changed, it is not lateral movement. However, since the aggregation of causal relationships is not precise enough, for example, the cause of a login record may be one of multiple login records and cannot be accurately determined. For such unclear login paths where user information may change, this detection method uses historical login records and historical alarm records to infer the probability of lateral movement. The advantage of this detection method is that it takes into account the temporal relationship of the login path. The disadvantage is that its login path rule aggregation is more complicated and needs to exclude multiple whitelist situations (such as jump servers, new machines, etc.), resulting in the need for improved detection accuracy.
[0035] Some researchers have also proposed that the graph deep learning network and the recurrent neural network can be stacked in a similar way to the GCRN model to detect lateral movement. In this detection method, the time slices are first divided according to the set time interval, and a login graph is generated for each time slice; then, the graph deep learning network is used to generate an embedding vector for each node in the login graph; these vectors are then input into the recurrent neural network to obtain the prediction of the node embedding vector at the next moment; the dot product of the embedding vectors of the two nodes is then performed to obtain the probability that there are login records for the two nodes, and the one with the lower probability is judged as lateral movement. The advantage of this detection method is that it takes into account the dynamics of the network and combines deep learning methods to greatly improve its accuracy; the disadvantage is that it needs to divide the time slices, and the timing relationship within each time slice will be lost after the time slices are divided.
[0036] Some researchers have also proposed that the TGN model can be used to implement lateral movement detection to avoid the division of time slices. In this detection method, the detection model is mainly composed of a graph feature extraction module, a memory module, and an embedding module. The graph feature extraction module is used to calculate the daily features of each node, such as the number of users logging in to the node on the same day, the number of computers logging in to the node on the same day, etc.; the memory module is used to process each login record in chronological order, update the memory vectors of the source node and the destination node of the login record, and the node memory vector is updated using a recurrent neural network; the embedding module is used to enumerate each historical neighbor node of the node, and the node's memory vector at the time, the node's features at the time, the historical neighbor node's memory vector at the time, and the historical neighbor node's features at the time are used to obtain the node's current embedding vector through the attention function; finally, the embedding vectors of the two nodes are passed through the decoder to obtain the probability that the two nodes have an edge, and the one with a lower probability is judged as lateral movement. The advantage of this detection method is that there is no need to divide the time slice, so the time data will not be lost, which greatly improves its accuracy.
[0037] The common disadvantage of the above-mentioned existing lateral movement detection methods based on network traffic is that they all use the graph link prediction method to detect lateral movement, that is, to determine the probability of the existence of login records between two nodes. However, in the containerized cluster, there is no corresponding login record information available. This is because when the attacker invades the cluster, he does not log in to the cluster based on anyone's identity information, but exploits the vulnerability of the load, etc., to invade the cluster through remote code execution. In addition, although the API server of the containerized cluster is responsible for identity authentication and authorization, when it comes to internal operations of the cluster, the API server mainly identifies the identity of the load and the host, not the identity of the person. Moreover, in the containerized cluster, when using network traffic information to detect lateral movement, the graph link prediction method cannot be simply adopted, because in this prediction method, only the existence of the edge is judged, and the characteristics of the edge are not considered; but in the containerized cluster, the characteristics of the edge, that is, the characteristics of the network traffic, play a very important role.
[0038] As mentioned in the background technology section, there are still two problems with lateral movement detection in a containerized environment. On the one hand, the features that can be used for lateral movement detection are unclear, making it difficult to implement lateral movement detection. On the other hand, the classification of lateral movement detection tasks is difficult and the samples are unbalanced, making it difficult to achieve the requirements of high detection rate and low false alarm rate.
[0039] Based on the analysis of existing detection methods and the problems mentioned in the background technology, the inventor proposes a lateral movement detection method suitable for containerized clusters. In this method, it includes a preprocessing stage, a first detection stage, a second detection stage and a detection result output stage, wherein the preprocessing stage is used to extract multiple target features and determine the value range of each target feature, as well as extract the topological structure of the target containerized cluster; the first detection stage is used to perform maximum value detection and topology detection; the second detection stage is used to use a pre-trained lateral movement detection model to perform traffic detection on the traffic sequence to be detected; the detection result output stage is used to perform deduplication processing to obtain the final detection result of the traffic sequence to be detected. Among them, target feature extraction in the preprocessing stage can solve the problem of unclear lateral movement detection features; at the same time, the two-stage detection design can fully detect lateral movement traffic, thereby improving the detection rate and reducing the false alarm rate.
[0040] In summary, if Figure 1 As shown, the present invention provides a lateral movement traffic detection method, which is used to detect whether there is lateral movement traffic in a traffic sequence to be detected generated in a target containerized cluster, wherein the traffic sequence to be detected includes multiple network flows that are continuous in time sequence, and the lateral movement traffic is the traffic data generated when the target containerized cluster is attacked. The method comprises: a preprocessing stage: obtaining a benign traffic sequence containing multiple benign flows generated by the target containerized cluster, performing feature extraction on the benign traffic sequence to obtain a target feature set containing multiple target features, and traversing each target feature in each benign flow in the benign traffic sequence to obtain a value range corresponding to each target feature, and performing topology extraction based on the benign traffic sequence to obtain a topological structure of the target containerized cluster; a first detection stage: maximum value detection: analyzing whether each network flow in the traffic sequence to be detected is lateral movement traffic; wherein each network flow includes multiple target features. Characteristic: any network traffic whose target feature value is not within the value range of the target feature is lateral movement traffic; all lateral movement traffic obtained by the maximum value detection of the traffic sequence to be detected constitutes the first detection result; topology detection: analyze whether the transmission process of each network traffic in the traffic sequence to be detected satisfies the topological structure of the target containerized cluster. If not, the network traffic is lateral movement traffic; wherein, all lateral movement traffic obtained by the topology detection of the traffic sequence to be detected constitutes the second detection result; second detection stage: use the pre-trained lateral movement detection model to perform traffic detection on the traffic sequence to be detected to analyze the lateral movement traffic in the traffic sequence to be detected to obtain the third detection result; detection result output stage: after removing the network traffic that is repeatedly judged as lateral movement traffic in the first detection result, the second detection result and the third detection result, all the remaining lateral movement traffic is used as the final detection result of the traffic sequence to be detected.
[0041] In order to better understand the present invention, each stage is described in detail below in conjunction with specific embodiments.
[0042] 1. Preprocessing stage
[0043] In the preprocessing stage, a benign traffic sequence containing multiple benign traffics generated by the target containerized cluster is obtained, feature extraction is performed on the benign traffic sequence to obtain a target feature set containing multiple target features, and each target feature in each benign traffic in the benign traffic sequence is traversed to obtain the value range corresponding to each target feature, and topology extraction is performed based on the benign traffic sequence to obtain the topological structure of the target containerized cluster.
[0044] 1.1 Target feature extraction and value range extraction
[0045] According to one embodiment of the present invention, the method includes extracting features from a benign traffic sequence in the following manner to obtain a target feature set containing multiple target features: each benign traffic in the benign traffic sequence includes multiple data packets, and based on all data packets of any benign traffic, packet-level feature extraction is performed on the benign traffic to obtain multiple packet-level traffic features corresponding to the benign traffic; and based on a packet connection protocol, all data packets in the benign traffic are aggregated into multiple sessions to extract multiple session-level traffic features corresponding to the benign traffic; wherein all packet-level traffic features and all session-level traffic features constitute an initial feature set; a preset evaluation method is used to perform importance evaluation on each traffic feature in the initial feature set to obtain an importance evaluation result for each traffic feature; the importance evaluation results are sorted in descending order, and a preset number of traffic features ranked first are selected as target features.
[0046] The packet-level traffic features are features extracted at the traffic packet level, such as the time difference between packets in a session, the size of packets in a session, and TCP tag value statistics in a session. In order to better understand the feature extraction at the packet level, the extraction of packet protocol field information is taken as an example for explanation, wherein the protocol field information can be determined by reading specific bytes in the packet header. For example, in an Ethernet frame, a byte starting from the 9th byte of the IP packet header is used to mark the protocol type. When the value of this byte is 6, it indicates the TCP protocol, and when it is 17, it indicates the UDP protocol.
[0047] The session-level traffic features are features extracted at the session level after aggregating all packets in the historical network traffic into multiple sessions based on the packet connection protocol, such as session duration, total number of bytes transmitted in the session, total number of packets transmitted in the session, etc. In order to better understand the extraction of session-level features, the byte traffic of the session is extracted as an example for explanation, wherein the byte traffic of the session is determined by calculating the total number of bytes sent and received during the session. Specifically, for each HTTP request and response packet in the session, the packet length field (usually in the packet header) is read, and the lengths of all packets are added. It should be noted that when aggregating packets into sessions, the connection protocol of the packets needs to be followed. For example, for a TCP connection, the establishment of a session starts with one party sending the first handshake packet to the other party, and ends with one party sending the last wave packet or resetting the connection packet or timeout to the other party; for packets of connectionless protocols such as UDP, a certain aggregation method can be adopted for specific application layer protocols. For example, for the DNS protocol, a query request and its corresponding result can be returned as a session. It should also be noted that in the process of aggregating data packets into sessions, there is a special set of data packets called bulk (when a TCP stream transmits a large amount of data, it needs to be divided into multiple data packets for transmission, and the set of these data packets is called bulk). When extracting features from this special set of data packets in bulk, we must first determine whether the data packets can be aggregated into bulk, and then count the number of data packets. When the number of data packets reaches 4, we can start extracting bulk features. Among them, the features related to bulk are usually 0, and are greater than 0 only when the amount of transmitted data is large, indicating that packet transmission has occurred in this flow.
[0048] Through feature extraction at the packet level and feature extraction at the session level, we can extract seven categories of features, including flow duration, number of packets, packet length, flow rate, packet interval, TCP flag statistics, packet header batch transmission, and TCP window active and idle, totaling 79 traffic features. The traffic characteristics include: backward URG flag (Bwd URG Flags), backward RST flag (Bwd RST Flags), CWR flag count (CWR Flag Count), ECE flag count (ECE Flag Count), forward subflow packet number (Subflow Fwd Packets), backward subflow packet number (Subflow Bwd Packets), reverse bulk packet transmission byte average (Bwd Bytes / Bulk AVg), reverse bulk packet transmission packet average (Bwd Packet / Bulk Avg), forward URG flag (Fwd URGFlags), reverse bulk transmission rate average (Bwd Bulk Rate Avg), URG flag count (URG Flag Count), continuous active time maximum (Active Max), continuous active time minimum (Active Min), packet length minimum (Packet Length Min), reverse packet length minimum (Bwd Packet Length Min), continuous idle time mean (ldle Mean), continuous active time standard deviation (Active Std), Fwd Packet / Bulk Avg, Ldle Min, Ldle Max, RST Flag Count, IdleStd, Fwd Bytes / Bulk Avg, Fwd Packet Length Min, Fwd Bulk Rate Avg, Fwd RST Flags, Fwd Act Data Pkts, Bwd IAT Min, Bwd PSH Flags, Fwd Seg SizeMin, Bwd Init Win Bytes, Total TCP Flow Time, PSH Flag CountFlag Count), ACK Flag Count, Fwd IAT Mean, FIN Flag Count, Bwd Header Length, Flow IAT Min, Fwd IAT Min, Total Fwd Packet, Flow IAT Mean, Bwd Packet Length Max, SYN Flag Count, Packet Length Max, Bwd IAT Mean, Down / Up Ratio, Total Bwd packets, Fwd Header Length, Flow Duration, Fwd IAT Total, Bwd IAT Total, Fwd Packet Length Max Mean), Fwd Segment Size Avg, Fwd PSH Flags, Flow IAT Std, Fwd IAT Std, Subflow Bwd Bytes, FWD Init Win Bytes, Bwd Packet Length Std, Average Packet Size, Bwd Packet Length Mean, Packet Length Variance, Bwd IAT Std, Flow IAT Max, Bwd Segment Size Avg, Bwd IAT Max, Fwd IAT Max, Packet Length Std, Packet Length AverageLength Mean), forward flow rate (FwdPackets / s), total length of forward packets (Total Length of Fwd Packet), standard deviation of forward packet length (Fwd Packet Length Std), number of forward subflow bytes (Subflow Fwd Bytes), total length of backward packets (Total Length of Bwd Packet), flow rate of packets (Flow Packets / s), flow rate of bytes (Flow Bytes / s), flow rate of backward packets (Bwd Packets / s), and maximum length of forward packets (Fwd Packet LengthMax).
[0049] According to one embodiment of the present invention, the preset evaluation method is: using gradient boosting decision tree, random forest and mutual information to evaluate the importance of each traffic feature in the initial feature set, and calculating the mean of the importance of each traffic feature to obtain the importance evaluation result of each traffic feature. Specifically, the gradient boosting decision tree is used to calculate the importance of each traffic feature in the initial feature set, and the calculated importance of each traffic feature is scaled to the interval [0,1]; the random forest is used to calculate the importance of each traffic feature in the initial feature set, and the calculated importance of each traffic feature is scaled to the interval [0,1]; the mutual information is used to calculate the importance of each traffic feature in the initial feature set, and the calculated importance of each traffic feature is scaled to the interval [0,1]; the mean of the importance of each traffic feature is calculated, and the traffic features are sorted in order from high to low importance. In order to better understand the importance of different traffic features, Figure 2 The traffic characteristics ranking results are shown to illustrate, among which, Figure 2 The traffic features in the above table are sorted in order from low to high importance. Figure 2 It can be seen that the importance of the six traffic features, namely, backward URG flag (Bwd URGFlags), backward RST flag (Bwd RST Flags), CWR flag count (CWR Flag Count), ECE flag count (ECE Flag Count), forward subflow packet count (Subflow Fwd Packets) and backward subflow packet count (Subflow Bwd Packets), is 0, and these six traffic features can be directly removed.
[0050] According to an embodiment of the present invention, the preset number is 57. It should be noted that, according to the order of importance from high to low, the top 57 traffic features can be selected to be retained, or more or fewer traffic features can be retained according to actual needs, such as retaining the top 60 features, or retaining the top 50 features, and the present invention does not impose any special restrictions.
[0051] According to one embodiment of the present invention, the target features include: standard deviation of continuous idle time, average value of forward bulk data packet transmission bytes, minimum value of forward data packet length, average value of forward bulk data packet rate, forward RST flag, number of forward active data packets, minimum value of backward inter-frame arrival time, backward PSH flag, minimum value of forward message segment size, initial value of backward window, total TCP flow time, PSH flag count, ACK flag count, average value of forward inter-frame arrival time, FIN flag count, backward header length, minimum value of inter-flow arrival time, minimum value of forward inter-frame arrival time, number of forward data packets, average value of data packet arrival time interval, total length of backward data packets, SYN flag count, maximum value of data packet length, average value of backward inter-frame arrival time, ratio of backward to forward data packets, number of backward data packets, forward header length, flow duration in milliseconds, forward number Sum of packet arrival time intervals, Sum of backward packet arrival time intervals, Mean value of forward packet length, Mean value of forward segment size, Forward PSH flag, Standard deviation of packet arrival time intervals, Standard deviation of forward packet arrival time intervals, Number of bytes of backward subflow, Forward window initial value, Standard deviation of backward packet length, Mean value of packet length, Mean value of backward packet length, Packet length variation, Standard deviation of backward packet arrival time intervals, Maximum value of packet arrival time intervals, Mean value of backward segment size, Maximum value of backward packet arrival time intervals, Maximum value of forward packet arrival time intervals, Standard deviation of packet length, Mean value of packet length, Forward flow rate, Total length of forward packets, Standard deviation of forward packet length, Number of bytes of forward subflow, Total length of backward packets, Packet flow rate, Byte flow rate, Backward packet flow rate and Maximum value of forward packet length.
[0052] After extracting the target feature set based on the above embodiment, each target feature in each benign flow in the benign flow sequence is traversed to obtain the value range corresponding to each target feature. The code shown in Table 1 can be used to obtain the value range corresponding to each target feature.
[0053] Table 1
[0054]
[0055] 1.2 Topology Extraction
[0056] According to one embodiment of the present invention, the method includes performing topology extraction based on a benign traffic sequence to obtain the topology structure of a target containerized cluster in the following manner: each benign traffic in the benign traffic sequence corresponds to a source node and a target node; wherein the source node represents a node that sends the benign traffic in the target containerized cluster, and the target node represents a node that receives the benign traffic in the target containerized cluster; based on the node types of the source node and the target node corresponding to each benign traffic in the benign traffic sequence, a load node set, an API server set, a load node set that communicates with the API server, and a load node set that communicates with an external node are obtained; wherein the load node set, the API server set, the load node set that communicates with the API server, and the load node set that communicates with an external node constitute the topology structure of the target containerized cluster. Wherein, the code shown in representation 2 can be used to obtain the topology structure of the target containerized cluster.
[0057] Table 2
[0058]
[0059] 2. The first testing stage
[0060] 2.1 Maximum value detection
[0061] Before introducing the specific operation of extreme value detection, let's first explain why extreme value detection can be used to identify lateral movement traffic.
[0062] Lateral movement attacks will cause changes in the traffic patterns of the load, host, and API server in the containerized cluster, which will lead to certain differences between the generated lateral movement traffic and benign traffic. This difference mainly exists in two aspects. On the one hand, for a small but critical amount of lateral movement traffic, it will transmit a large amount of data, causing the values of some batch transmission features and packet length features to exceed the value range of benign traffic (for this part of lateral movement traffic, it can be detected by the maximum value). This is because the load is under the control of the attacker's C&C server. The attacker installs malware or scripts on the load and steals key information from the API server. On the other hand, for most lateral movement traffic, although their feature distribution is different from that of benign traffic, their feature values are still within the value range of benign traffic, and lateral movement traffic often has lower data transmission requirements (for this part of lateral movement traffic, it cannot be detected by the maximum value).
[0063] In order to more intuitively understand the difference in the value range between lateral movement traffic and benign traffic, we take active and idle features as an example to analyze the value range distribution of lateral movement traffic and benign traffic in such features, and obtain the following: Figure 3 The box plot shown. Figure 3(a) is a box plot of the maximum continuous active duration, and Figure 3 (a) The upper middle part shows the maximum continuous active duration corresponding to benign traffic. Figure 3 (a) The lower middle layer shows the maximum continuous active duration corresponding to the lateral movement traffic; Figure 3 (b) is a box plot of the standard deviation of continuous active duration, and Figure 3 (b) The upper middle part is the standard deviation of the continuous active duration corresponding to benign traffic. Figure 3 (b) The lower middle part is the standard deviation of the continuous active duration corresponding to the lateral movement traffic. Figure 3 It can be seen that, whether it is benign traffic or lateral movement traffic, the data distribution under this feature presents a long-tail distribution, and the boxes and lines of the box plot are concentrated on the left side of the figure and are almost invisible; at the same time, most of the space in the figure is occupied by outliers in the box plot. Among these outliers, the outlier on the far right of the lateral movement traffic is further to the right than the benign traffic, indicating that the maximum value and standard deviation of the continuous active duration of the lateral movement traffic are larger. This means that after the containerized cluster suffers a lateral movement attack, some nodes in the network may exchange data more frequently, which leads to a larger maximum value of the continuous active duration of the session, and also drives its standard deviation to increase.
[0064] In order to more intuitively understand the data transmission differences between lateral movement traffic and benign traffic, we take the packet length feature as an example to analyze the data transmission of lateral movement traffic and benign traffic in the length feature. Figure 4 The density map shown. Figure 4 (a) is the maximum value density diagram of the forward data packet length; Figure 4 (b) is the density diagram of the average length of forward data packets; Figure 4 (c) is the density diagram of the standard deviation of the forward packet length; and Figure 4 The blue part is benign traffic, and the yellow part is lateral movement traffic. Figure 4 It can be seen that compared with benign traffic, lateral movement traffic has lower data transmission requirements and is hidden among benign traffic.
[0065] According to the above analysis, after the containerized cluster is attacked by mobile, some characteristic values generated are greater than the lateral movement traffic of benign traffic. Based on this, the inventor proposes that the lateral movement traffic detection can be realized by using the maximum value detection method. Among them, the maximum value detection includes: analyzing whether each network traffic in the traffic sequence to be detected is lateral movement traffic, that is, any network traffic whose target characteristic value is not within the value range of the target characteristic is lateral movement traffic; and taking all lateral movement traffic obtained by the maximum value detection of the traffic sequence to be detected as the first detection result. Among them, the code shown in Representation 3 can be used to perform the maximum value detection.
[0066] Table 3
[0067]
[0068] 2.2 Topology Detection
[0069] Before introducing the specific topology detection operation, let's first explain why topology detection can be used to identify lateral movement traffic.
[0070] A containerized cluster is an architecture that combines multiple containers for management and orchestration. There are two common topological structures for containerized clusters. One is the master-slave structure, which includes a master node and multiple slave nodes. The master node is responsible for managing the state of the entire cluster, scheduling tasks and other core functions; the slave node is responsible for receiving instructions from the master node and running containerized applications; the other is the P2P structure, in which all nodes have equal status and there is no master-slave distinction. Each node can independently receive task requests, run containers, and communicate and collaborate with other nodes. Among them, the node types in the containerized cluster include load nodes, API servers, and external nodes.
[0071] Lateral movement attacks will change the topology in the containerized cluster, that is, the connection relationship between nodes in the containerized cluster changes. Such changes include: the increase of load nodes, the communication between load nodes that do not need to communicate, the increase of nodes communicating with the API server or the communication between nodes that have never communicated with the API server, and the access of load nodes that have never accessed external nodes to external nodes. Among them, a new load node appears in the containerized cluster, and the load node is irrelevant to the expansion strategy configured in the cluster, which means that the attacker illegally created a new load through the API server. The load may have special permissions so that the attacker can move from the load to the host. The communication between load nodes that do not need to communicate indicates that the attacker has moved laterally between loads. The increase of nodes communicating with the API server or the communication between nodes that have never communicated with the API server indicates that the load node starts to communicate with the API server after being attacked in order to obtain the cluster configuration file and perform resource operations. The load node that has never accessed external nodes starts to access external nodes, which indicates that the load node is controlled by the attacker and communicates with the attacker's server.
[0072] According to the above analysis, lateral movement attacks will cause the topology of the containerized cluster to change. Based on this, the inventor proposes that lateral movement traffic detection can be achieved by topology detection. The topology detection includes: analyzing whether the transmission process of each network flow in the flow sequence to be detected meets the topology of the target containerized cluster. If not, the network flow is lateral movement traffic, and all lateral movement traffic obtained by topology detection of the flow sequence to be detected is used as the second detection result.
[0073] According to one embodiment of the present invention, the method includes analyzing whether the transmission process of each network flow in the flow sequence to be detected satisfies the topological structure of the target containerized cluster in the following manner: if the source node or the target node corresponding to the network flow is a new load node and is not in the load node set, the network flow is lateral movement flow; if the source node corresponding to the network flow is not in the load node set communicating with the API server, and the target node corresponding to the network flow is the API server, the network flow is lateral movement flow; if the target node corresponding to the network flow is not in the load node set communicating with the API server, and the source node corresponding to the network flow is the API server, the network flow is lateral movement flow; if the source node corresponding to the network flow is not in the load node set communicating with the external node, and the target node corresponding to the network flow is an external node, the network flow is lateral movement flow; if the target node corresponding to the network flow is not in the load node set communicating with the external node, and the source node corresponding to the network flow is an external node, the network flow is lateral movement flow. Wherein, the code shown in Table 4 can be used for topology detection.
[0074] Table 4
[0075]
[0076] In order to verify the feasibility of extreme value detection and topology detection, Kubernetes-dataset is used as the data set to perform extreme value detection and topology detection, and the extreme value detection results are shown in Table 5 and the topology detection results are shown in Table 6. Among them, in the Kubernetes-dataset data set, benign traffic and lateral movement traffic are recorded separately, so the lateral movement traffic is injected into the end of the benign traffic, while retaining the time relationship between the lateral movement traffic and the benign traffic, and aligning the end time of the lateral movement traffic and the benign traffic.
[0077] As shown in Table 5, the maximum detection can detect 5 lateral movement flows, and another 3 benign flows are falsely reported as lateral movement. As shown in Table 6, the topology detection detects 8 lateral movement flows without any false positives.
[0078] Table 5
[0079]
[0080] Table 6
[0081]
[0082] 3. Second Testing Phase
[0083] The traffic sequence to be detected can detect part of the lateral movement traffic through the detection in the first detection stage, but there is still part of the lateral movement traffic that has not been detected. In order to improve the detection rate of the lateral movement traffic, the inventor proposes that a pre-trained lateral movement detection model can be used to further process the traffic sequence to be detected. Based on this, in the second detection stage, the pre-trained lateral movement detection model is used to perform traffic detection on the traffic sequence to be detected to analyze the lateral movement traffic in the traffic sequence to be detected to obtain a third detection result.
[0084] 3.1 Model Detection
[0085] According to one embodiment of the present invention, the pre-trained lateral movement detection model is configured to perform traffic detection on the traffic sequence to be detected in the following manner: using the traffic sequence to be detected as input, generating a target prediction sequence of the traffic sequence to be detected in a recursive prediction manner; wherein the target prediction sequence includes multiple target predicted network flows, and each target predicted network flow corresponds to a network flow in the traffic to be detected; calculating the error value between each target predicted network flow and its corresponding actual network flow, and if the error value is greater than or equal to a threshold, the network flow is judged to be lateral movement traffic. It should be noted that the target prediction sequence of the traffic sequence to be detected is generated in a recursive prediction manner. For example, assuming that there is a traffic sequence to be detected with a length of n , the lateral movement detection model processes the traffic sequence to be detected When the network traffic Predicting network traffic , based on network traffic and Predicting network traffic , based on network traffic to Predicting network traffic Based on this, the target predicted network traffic sequence is It should also be noted that based on the traffic sequence to be detected The predicted target prediction network sequence does not contain network traffic The corresponding target predicts network traffic. There are two situations at this time. One situation is that there is a previous traffic sequence to be detected, and the network traffic has been predicted in the previous traffic sequence to be detected. The corresponding target predicts network traffic; another case is to predict network traffic by padding The corresponding target predicts network traffic.
[0086] 3.2 Lateral Movement Detection Model
[0087] According to one embodiment of the present invention, the pre-trained lateral movement detection model is a model obtained by training in the following manner: obtaining a training benign traffic sequence, wherein the training benign traffic sequence includes a plurality of benign flows that are continuous in time series, and each benign flow includes a plurality of traffic features; performing feature extraction on the training benign traffic sequence to obtain a target feature set corresponding to each benign flow in the training benign traffic sequence; taking the benign traffic sequence after feature extraction processing as input and the target prediction sequence corresponding to the benign traffic sequence after feature extraction processing as output, performing multiple rounds of iterative training until the lateral movement detection model converges.
[0088] According to one embodiment of the present invention, the pre-trained lateral movement detection model includes a Transformer model and a judgment module, wherein: the Transformer model is used to predict the target prediction sequence corresponding to the traffic sequence to be detected; the judgment module is used to calculate the error value between each target predicted network traffic and its corresponding actual network traffic, and when the error value is greater than or equal to a threshold, the network traffic is judged to be lateral movement traffic. It should be noted that the lateral movement traffic detection model can also set other neural network models to replace the Transformer model, such as the RNN model.
[0089] According to one embodiment of the present invention, Figure 5 As shown ( Figure 5 The judgment module is omitted in the figure), the pre-trained lateral movement detection model includes a spatial feature embedding module, a temporal feature embedding module, a coding and decoding module, a sequence output module and a judgment module. The following introduces each module in the lateral movement detection model.
[0090] Among them, the spatial feature embedding module is used to extract the spatial features of the traffic sequence to be detected to obtain the spatial feature embedding vector of each network traffic in the traffic sequence to be detected, and connect the spatial feature embedding vector of each network traffic with the traffic feature of the network traffic to obtain the initial feature of the traffic sequence to be detected. It should be noted that the reason for extracting spatial features from the traffic sequence to be detected is that the network traffic in the containerized cluster obeys certain rules, and certain spatial features can be mined from it. If the network traffic is modeled as a graph, the nodes in the graph represent the combination of IP addresses and ports corresponding to the nodes in the target containerized cluster, and the edges represent the network traffic generated between the nodes.
[0091] According to one embodiment of the present invention, the spatial feature embedding module is a graph convolutional network (GCN). It should be noted that the spatial feature embedding module can also use other network models for processing graph structure data, such as GraphSAGE, GAT, DeepWalk and Node2vec, and the present invention does not impose special restrictions on the spatial feature embedding module.
[0092] In order to better understand the working mode of the spatial feature embedding module, we take the graph convolutional network (GCN) as an example to illustrate how to extract the spatial features of the traffic sequence to be detected. Figure 5 As shown, the spatial feature embedding module takes the graph data constructed by the traffic sequence to be detected (the graph data includes multiple nodes and multiple edges connecting two nodes) as input, and propagates information through multiple network layers to extract the spatial embedding vector of each node in the graph data, wherein in each layer of information propagation process, all neighbor features of each node and the weighted average of its own node are extracted to obtain the spatial embedding vector of each node corresponding to the network layer, and the spatial embedding vector of each node corresponding to the network layer is passed to the next network layer for processing until all network layers complete the information propagation processing to obtain the final spatial embedding vector of each node; after obtaining the spatial embedding vector of all nodes, it is connected with the traffic features of each network traffic in the traffic sequence to be detected to obtain the initial feature output of the traffic sequence to be detected. Among them, the initial features of the traffic sequence to be detected include the initial features corresponding to each network traffic, and the initial features corresponding to each network traffic are formed by connecting the spatial embedding vector of the source node corresponding to the network traffic, the spatial embedding vector of the target node and the traffic features.
[0093] It should be noted that the spatial feature embedding module does not participate in the update during the multi-round iterative training process, and it is trained separately using benign traffic sequences. It should also be noted that when performing lateral movement detection, the traffic sequence to be detected may contain new nodes that the spatial feature embedding module has not seen, so the spatial embedding vectors of these new nodes need to be considered. According to the network structure of the containerized cluster, the new node may be a new API server, a new load, a new host, or a new external node. Among them, the new host needs to be configured by the administrator of the containerized cluster, so the possibility of being created by an attacker is low and is not included in the consideration of the spatial feature embedding module; the new API server depends on the new load, and the new load has been screened out in the topology detection in the previous stage, so it is also not included in the consideration of the spatial feature embedding module; therefore, the new node included in the traffic sequence to be detected can only be a new external node. For the new external node, the spatial feature embedding module randomly generates a vector as the spatial embedding vector of the node.
[0094] The time feature embedding module is used to extract the time feature based on the initial feature of the traffic sequence to be detected to obtain the time feature embedding vector of the traffic sequence to be detected. According to one embodiment of the present invention, the time feature embedding module is a long short-term memory network (LSTM), and the number of network layers of the long short-term memory network can be set to 3. It should be noted that the number of network layers of the long short-term memory network is determined by actual needs, and the present invention does not impose any special restrictions.
[0095] like Figure 5 As shown, the encoding and decoding module includes an encoder, a first decoder and a second decoder; wherein: the encoder is used to extract the dependency relationship between each network flow in the flow sequence to be detected based on the initial features of the flow sequence to be detected, so as to obtain the first latent vector of the flow sequence to be detected; and based on the initial features of the flow sequence to be detected and the error between each initial predicted network flow obtained by the sequence output module and the actual network flow in the flow sequence to be detected corresponding to itself, the dependency relationship between each network flow in the flow sequence to be detected is re-extracted to obtain the second latent vector of the flow sequence to be detected; the first decoder is used to perform decoding processing based on the first latent vector of the flow sequence to be detected and the time feature embedding vector to obtain the first decoded feature vector of the flow sequence to be detected; the second decoder is used to perform decoding processing based on the second latent vector of the flow sequence to be detected and the time feature embedding vector to obtain the second decoded feature vector of the flow sequence to be detected; the sequence output module is used to generate the initial prediction sequence of the flow sequence to be detected based on the first decoded feature vector of the flow sequence to be detected, and to generate the target prediction sequence of the flow sequence to be detected based on the second decoded feature vector of the flow sequence to be detected; wherein the initial prediction sequence includes multiple initial predicted network flows, and each initial predicted network flow corresponds to a network flow in the flow sequence to be detected.
[0096] Among them, Figure 5 As shown, the decoder includes a multi-head attention module, a first addition and regularization module, a feedforward network and a second addition and regularization module, wherein the multi-head attention module is used to mine the correlation and dependency between network flows; the first addition and regularization module is used to integrate the network flow and the network flow processed by the multi-head attention module into a unified representation and pass it to the feedforward network; the feedforward network is used to further process the processing results of the first addition and regularization module to learn the complex semantic relationships in each network flow; the second addition and regularization module is used to integrate the processing results of the feedforward network and the processing results of the first addition and regularization module into a unified representation and pass it to the first decoder or the second decoder.
[0097] like Figure 5 As shown, the first decoder and the second decoder both include a multi-head attention module and an addition and regularization module. The multi-head attention modules in the first decoder and the second decoder are used to mine the correlation between the output of the temporal feature embedding module and the output of the decoder; the addition and regularization module is used to integrate the processing results of the multi-head attention module into a unified representation and pass it to the sequence output module.
[0098] like Figure 5As shown, the sequence output module includes a feedforward network and a sigmoid activation function, wherein the feedforward network is used to further process the output of the first decoder or the second decoder and obtain the corresponding initial prediction sequence and target prediction sequence after processing by the sigmoid function.
[0099] The judgment module is used to calculate the error value between each target predicted network flow and its corresponding actual network flow, and when the error value is greater than or equal to a threshold, the network flow is judged to be lateral movement flow.
[0100] To better understand how the pre-trained lateral movement detection model works, Figure 5 Take the lateral movement detection model shown in the figure as an example and explain it in combination with the specific detection process. Specifically, assume that the traffic sequence to be detected is ( for The network traffic at the moment, for In the detection process, the graph data constructed based on the traffic sequence to be detected is first input into the lateral movement detection model, and the spatial feature embedding module in the lateral movement detection model extracts spatial features from the graph data corresponding to the traffic sequence to be detected to obtain the traffic sequence to be detected. The corresponding spatial feature embedding vector is obtained, and the spatial feature embedding vector of each network flow is connected with the flow feature of the network flow to obtain the initial features of the flow sequence to be detected. ,in, Include network traffic The corresponding source node The spatial embedding vector of the target node The spatial embedding vector and traffic features of Include network traffic The corresponding source node The spatial embedding vector of the target node The spatial embedding vector and traffic features of the traffic sequence to be detected are then Passed to the time feature embedding module and the encoding and decoding module, where the time feature embedding module is used to generate the initial features of the traffic sequence to be detected. Extract the time feature to obtain the time feature embedding vector of the traffic sequence to be detected, and pass the time feature embedding vector of the traffic sequence to be detected to the encoding and decoding module; then, perform two-step prediction, where in the first step of prediction, the encoder, the first decoder and the sequence output module are used to predict the network traffic The corresponding initial predicted network traffic (Simplified description, only based on network traffic Predicting network traffic process), and calculate the initial predicted network traffic With network traffic The error between ; In the second step of prediction, the error The initial characteristics of the traffic sequence to be detected The network traffic is predicted by connecting the encoder, the second decoder and the sequence output module. Corresponding target prediction network traffic ; Finally, the judgment module calculates the target predicted network traffic The corresponding actual network traffic When the error value is greater than or equal to the threshold, the network traffic is judged to be lateral movement traffic, otherwise, it is benign traffic.
[0101] It should be noted that Figure 5 The lateral movement detection model shown can be trained in multiple rounds with the loss function To calculate the loss, and to minimize the loss to update the model parameters, where represents the error between the network traffic and its corresponding initial predicted network traffic, represents the error between the network flow and its corresponding target predicted network flow, represents the weight, and It decreases with the training iterations.
[0102] It should also be noted that Figure 5 The lateral movement detection model shown uses the error value between the target predicted network traffic and its corresponding actual network traffic to determine whether the corresponding network traffic is lateral movement traffic. Since the error value between the target predicted network traffic and its corresponding actual network traffic is the final error formed after two-step prediction, it can minimize the prediction error of benign traffic, thereby making the prediction error of lateral movement traffic relatively large, thereby reducing the false alarm rate of the model and improving the detection rate of the model.
[0103] 4. Test result output stage
[0104] In the detection result output stage, after removing the network traffic that is repeatedly determined to be lateral movement traffic in the first detection result, the second detection result, and the third detection result, all the remaining lateral movement traffic is used as the final detection result of the traffic sequence to be detected.
[0105] In order to improve the performance of the lateral movement detection model proposed in the present invention in a containerized environment, Figure 5The lateral movement detection model shown is compared with the existing lateral movement detection models proposed by other researchers. Among them, the lateral movement detection model proposed by researcher Bowman, Euler model, TGN model, USAD model, OmniAnomaly model and TranAD model are involved in the comparison experiment.
[0106] In the comparative experiment, the parameters of each model are set to the parameters shown in Table 7, and the same training data and the same test data are used to evaluate the TPR, FPR, F1 and AUC scores of each model, and the experimental results shown in Table 8 are obtained.
[0107] In Table 7, the meanings of the parameters are as follows. For the model proposed by Bowman: p represents the possibility of repeatedly visiting a node, q represents the control parameter for interpolation between the breadth-first strategy and the depth-first strategy, walk_length represents the length of the walk, context_size represents the actual context size of the sample, and the effective sampling rate is improved by reusing samples between different source nodes; walks_per_node represents the number of samples of each node. For the Euler model: delta:euler represents the time-related parameters of the model, gnn represents the graph convolutional network, and rnn represents the gated recurrent unit. For the TGN model, batch_size represents the number of data (samples) passed to the program for training at a single time. For the USAD model, n_hidden represents the number of hidden parameters, n_latent represents the dimension of the latent space, and n_window represents the window size. For the OmniAnomaly model: Beta represents the scale parameter of the tail of the fitted probability distribution of the generalized Pareto distribution, n_hidden represents the number of hidden parameters, n_latent represents the dimension of the latent space, and n_window represents the window size. For the TranAD model, n_window represents the window size, rnn represents the three-layer LSTM, and gnn represents the graph convolutional network. The model proposed by the present invention: n_window represents the window size, rnn represents the three-layer LSTM, and gnn represents the graph convolutional network.
[0108] As shown in Table 8, the model proposed by researcher Bowman obtained the lowest score because the model did not utilize time information but only performed link prediction tasks on static graphs. The Euler model has a slightly higher AUC score because it uses a link prediction method for dynamic graphs; however, since the model method requires continuous time to be divided into discrete time slices, the granularity of its time information is too coarse, resulting in its low performance. The TGN model models continuous time without dividing time slices, so it obtains the highest AUC score compared with the model proposed by Bowman and the Euler model. However, the AUC scores of these three models are lower than those of the other models, which shows that graph-based models are not applicable in containerized environments; in addition, the experimental results of the TGN model show that lateral movement detection should be considered from a global perspective, rather than just considering the local information between two adjacent nodes on the graph.
[0109] The AUC scores of the USAD model, OmniAnomaly model, and TranAD model are all above 0.8, indicating that the encoder-decoder structure and adversarial training method can improve the detection performance. Among the three models, the AUC score of the TranAD model is significantly higher than that of the other two models, reaching 0.91, which shows that the self-attention mechanism is particularly suitable for lateral movement detection scenarios.
[0110] The model proposed in this invention has the highest AUC score, reaching 0.96, and its detection rate and false alarm rate are better than those of the TranAD model.
[0111] Regarding the training time overhead, the model proposed by the researcher Bowman and the Euler model have the shortest training time, because they do not consider fine-grained temporal information and their model capacity is insufficient. Among the remaining models, the model proposed by this paper has the shortest training time, because, compared with TGN, this paper does not couple the temporal information with the spatial structure for training; compared with USAD, OmniAnomaly, and TranAD, the method proposed in this paper removes the adversarial training mechanism, so the training time can be shortened.
[0112] Table 7
[0113]
[0114] Table 8
[0115]
[0116] The beneficial effects of the present invention are: (1) extracting packet-level traffic features and session-level traffic features, and eliminating traffic features with lower importance based on the importance ranking results of the traffic features, thereby solving the problem of unclear lateral movement detection features in containerized clusters; (2) setting up two-stage detection to screen lateral movement traffic, thereby improving the detection rate and reducing the false alarm rate; (3) the lateral movement detection model uses two-step prediction to minimize the prediction error of benign traffic and maximize the prediction error of lateral movement traffic, thereby achieving more accurate lateral movement traffic detection.
[0117] It should be noted that although the above describes the various steps in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.
[0118] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0119] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a protruding structure in a groove on which instructions are stored, and any suitable combination thereof.
[0120] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A lateral movement traffic detection method is used to detect whether there is lateral movement traffic in the to-be-detected traffic sequence generated in a target containerized cluster, wherein: The traffic sequence to be detected includes multiple network flows that are continuous in time sequence, and the lateral movement traffic is the traffic data generated when the target containerized cluster is attacked. The method includes: Preprocessing stage: Obtain a benign traffic sequence containing multiple benign traffics generated by a target containerized cluster, perform feature extraction on the benign traffic sequence to obtain a target feature set containing multiple target features, traverse each target feature in each benign traffic in the benign traffic sequence to obtain a value range corresponding to each target feature, and perform topology extraction based on the benign traffic sequence to obtain a topological structure of the target containerized cluster; First testing stage: Maximum value detection: Analyze whether each network flow in the flow sequence to be detected is lateral movement flow; wherein each network flow includes multiple target features, and any network flow whose target feature value is not within the value range of the target feature is lateral movement flow; all lateral movement flows obtained by the maximum value detection of the flow sequence to be detected constitute the first detection result; Topology detection: Analyze whether the transmission process of each network flow in the flow sequence to be detected meets the topology structure of the target containerized cluster. If not, the network flow is lateral movement flow. Among them, all lateral movement flows obtained by topology detection of the flow sequence to be detected constitute the second detection result. Second testing stage: Using a pre-trained lateral movement detection model to perform traffic detection on the traffic sequence to be detected to analyze the lateral movement traffic in the traffic sequence to be detected to obtain a third detection result; Detection result output stage: After removing the network traffic that is repeatedly determined to be lateral movement traffic in the first detection result, the second detection result, and the third detection result, all remaining lateral movement traffic is used as the final detection result of the traffic sequence to be detected.
2. The method according to claim 1, characterized in that: The method comprises extracting features from a benign traffic sequence in the following manner to obtain a target feature set including a plurality of target features: Each benign flow in the benign flow sequence includes multiple data packets, and based on all the data packets of any benign flow, data packet-level feature extraction is performed on the benign flow to obtain multiple data packet-level flow features corresponding to the benign flow; and aggregating all packets in the benign traffic into multiple sessions based on a packet connection protocol to extract multiple session-level traffic features corresponding to the benign traffic; wherein all packet-level traffic features and all session-level traffic features constitute an initial feature set; The importance of each traffic feature in the initial feature set is evaluated using a preset evaluation method to obtain an importance evaluation result of each traffic feature; The importance evaluation results are sorted in descending order, and a preset number of traffic features ranked first are selected as target features.
3. The method according to claim 2, characterized in that The preset evaluation method is: Gradient boosting decision tree, random forest and mutual information are used to evaluate the importance of each traffic feature in the initial feature set, and the mean of the importance of each traffic feature is calculated to obtain the importance evaluation result of each traffic feature.
4. The method according to claim 3, characterized in that The preset number is 57.
5. The method according to claim 4, characterized in that The target features include: standard deviation of continuous idle time, average value of forward bulk data packet transmission bytes, minimum value of forward data packet length, average value of forward bulk data packet rate, forward RST flag, number of forward active data packets, minimum value of backward inter-frame arrival time, backward PSH flag, minimum value of forward message segment size, initial value of backward window, total TCP flow time, PSH flag count, ACK flag count, average value of forward inter-frame arrival time, FIN flag count, backward header length, minimum value of inter-flow arrival time, minimum value of forward inter-frame arrival time, number of forward data packets, average value of data packet arrival time interval, total length of backward data packets, SYN flag count, maximum value of data packet length, average value of backward inter-frame arrival time, ratio of backward to forward data packets, number of backward data packets, forward header length, flow duration in milliseconds, forward data packet arrival time Sum of intervals, Sum of backward packet arrival intervals, Mean forward packet length, Mean forward segment size, Forward PSH flag, Standard deviation of packet arrival intervals, Standard deviation of forward packet arrival intervals, Number of bytes in backward subflow, Forward window initial value, Standard deviation of backward packet lengths, Mean packet lengths, Mean backward packet lengths, Packet length variation, Standard deviation of backward packet arrival intervals, Maximum packet arrival intervals, Mean backward segment size, Maximum backward packet arrival intervals, Maximum forward packet arrival intervals, Standard deviation of packet lengths, Mean packet lengths, Forward flow rate, Total forward packet length, Standard deviation of forward packet lengths, Number of bytes in forward subflow, Total backward packet length, Packet flow rate, Byte flow rate, Backward packet flow rate, and Maximum forward packet length.
6. The method according to claim 5, characterized in that The method includes performing topology extraction based on a benign traffic sequence to obtain a topology structure of a target containerized cluster in the following manner: Each benign flow in the benign flow sequence corresponds to a source node and a target node; wherein the source node represents a node that sends the benign flow in the target containerized cluster, and the target node represents a node that receives the benign flow in the target containerized cluster; Based on the node types of the source node and the target node corresponding to each benign traffic in the benign traffic sequence, a load node set, an API server set, a load node set communicating with the API server, and a load node set communicating with external nodes are obtained; wherein the load node set, the API server set, the load node set communicating with the API server, and the load node set communicating with external nodes constitute the topology structure of the target containerized cluster.
7. The method according to claim 6, characterized in that The method includes analyzing whether the transmission process of each network flow in the flow sequence to be detected satisfies the topological structure of the target containerized cluster in the following manner: If the source node or target node corresponding to the network traffic is a new load node and is not in the load node set, the network traffic is lateral movement traffic; If the source node corresponding to the network traffic is not in the set of load nodes communicating with the API server, and the target node corresponding to the network traffic is the API server, then the network traffic is lateral movement traffic; If the target node corresponding to the network traffic is not in the set of load nodes communicating with the API server, and the source node corresponding to the network traffic is the API server, then the network traffic is lateral movement traffic; If the source node corresponding to the network traffic is not in the set of load nodes that communicate with the external node, and the destination node corresponding to the network traffic is an external node, then the network traffic is lateral movement traffic; If the target node corresponding to the network traffic is not in the set of load nodes that communicate with the external node, and the source node corresponding to the network traffic is an external node, then the network traffic is lateral movement traffic.
8. The method according to claim 7, characterized in that The pre-trained lateral movement detection model is configured to perform traffic detection on the traffic sequence to be detected as follows: Taking the flow sequence to be detected as input, a target prediction sequence of the flow sequence to be detected is generated according to a recursive prediction method; wherein the target prediction sequence includes a plurality of target predicted network flows, and each target predicted network flow corresponds to a network flow in the flow to be detected; The error value between each target predicted network flow and its corresponding actual network flow is calculated. If the error value is greater than or equal to the threshold, the network flow is determined to be lateral movement flow.
9. The method according to claim 8, characterized in that The pre-trained lateral movement detection model is a model trained in the following manner: Acquire a benign traffic sequence for training, wherein the benign traffic sequence for training includes a plurality of benign traffics that are continuous in time sequence, and each benign traffic includes a plurality of traffic features; Perform feature extraction on the benign traffic sequence for training to obtain a target feature set corresponding to each benign traffic in the benign traffic sequence for training; Taking the benign traffic sequence after feature extraction as input and the target prediction sequence corresponding to the benign traffic sequence after feature extraction as output, multiple rounds of iterative training are performed until the lateral movement detection model converges.
10. The method according to claim 9, characterized in that The pre-trained lateral movement detection model includes a Transformer model and a judgment module, wherein: The Transformer model is used to predict the target prediction sequence corresponding to the traffic sequence to be detected; The judgment module is used to calculate the error value between each target predicted network flow and its corresponding actual network flow, and when the error value is greater than or equal to a threshold, the network flow is judged to be lateral movement flow.
11. The method according to claim 9, characterized in that The pre-trained lateral movement detection model includes a spatial feature embedding module, a temporal feature embedding module, a coding and decoding module, a sequence output module and a judgment module; wherein: The spatial feature embedding module is used to extract the spatial features of the traffic sequence to be detected to obtain the spatial feature embedding vector of each network traffic in the traffic sequence to be detected, and connect the spatial feature embedding vector of each network traffic with the traffic feature of the network traffic to obtain the initial features of the traffic sequence to be detected; The time feature embedding module is used to extract the time feature based on the initial feature of the flow sequence to be detected to obtain the time feature embedding vector of the flow sequence to be detected; The encoding and decoding module includes an encoder, a first decoder, and a second decoder, wherein: The encoder is used to extract the dependency relationship between each network flow in the flow sequence to be detected based on the initial features of the flow sequence to be detected, so as to obtain the first latent vector of the flow sequence to be detected; and to re-extract the dependency relationship between each network flow in the flow sequence to be detected based on the initial features of the flow sequence to be detected and the error between each initial predicted network flow obtained by the sequence output module and the actual network flow in the flow sequence to be detected corresponding to itself, so as to obtain the second latent vector of the flow sequence to be detected; The first decoder is used for performing decoding processing based on the first latent vector of the traffic sequence to be detected and the time feature embedding vector to obtain a first decoding feature vector of the traffic sequence to be detected; The second decoder is used for performing decoding processing based on the second latent vector of the traffic sequence to be detected and the time feature embedding vector to obtain a second decoding feature vector of the traffic sequence to be detected; The sequence output module is used to generate an initial prediction sequence of the traffic sequence to be detected based on the first decoded feature vector of the traffic sequence to be detected, and to generate a target prediction sequence of the traffic sequence to be detected based on the second decoded feature vector of the traffic sequence to be detected; wherein the initial prediction sequence includes a plurality of initial predicted network flows, and each initial predicted network flow corresponds to a network flow in the traffic sequence to be detected; The judgment module is used to calculate the error value between each target predicted network flow and its corresponding actual network flow, and when the error value is greater than or equal to a threshold, the network flow is judged to be lateral movement flow.
12. The method according to claim 11, characterized in that The spatial feature embedding module is a graph convolutional network, and the temporal feature embedding module is a long short-term memory network.
13. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of any method described in claims 1-12.
14. An electronic device, characterized in that: include: one or more processors, and A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of any one of the methods of claims 1-12 by executing the executable instructions.
Citation Information
Patent Citations
Lateral movement attack detection method and system based on heterogeneous graph network
CN113094707A
Transverse movement detection method, storage medium and electronic equipment
CN117614735A