A method for analyzing encrypted traffic based on interactive spatiotemporal characteristics
By constructing the FITDect model and using the GNN layer and MLP layer to analyze the spatiotemporal coupling relationship of encrypted traffic, the problem of encrypted traffic analysis relying on shallow statistical features in existing technologies is solved, and efficient identification and detection of encrypted traffic is achieved.
Patent Information
- Application Number
- CN202510828318.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing deep flow inspection technologies rely on shallow statistical features to analyze encrypted traffic, which cannot effectively characterize the potential interaction characteristics between data packets within the flow, making it difficult to identify and detect encrypted traffic.
A FITDect model is constructed, including the input layer, GNN layer and MLP layer. By mining the spatiotemporal interaction relationship of data packets within the encrypted flow, a dynamic traffic interaction graph is constructed. The traffic interaction graph is dynamically represented and the spatiotemporal coupling relationship is analyzed. The MLP layer is used to make classification decisions.
It significantly enhances the ability to characterize the hidden behavior of encrypted traffic, can effectively capture the implicit behavior patterns of encrypted traffic, and provides a new method for detecting malicious encrypted traffic.
Smart Images

Figure CN120378219B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and in particular to an encrypted traffic analysis method based on interactive spatiotemporal features. Background Art
[0002] Many malware or malicious users exploit the concealment provided by encryption protocols to conceal the transmission of malware or remote control commands, reducing the probability of detection. Unlike traditional plaintext traffic, encrypted traffic, because encryption protocols (such as SSL / TLS and IPSec) encrypt data packets at the application layer, the application-layer plaintext is invisible. This poses significant challenges for both QoS (Quality of Service) assessment on the business side and Intrusion Detection and Recognition (IDR) (Intrusion Detection and Recognition) on the security side.
[0003] Existing traffic analysis methods can be categorized as either DPI (Deep Packet Inspection) or DFI (Deep Flow Inspection). DPI technology uses the data packet as the smallest analysis unit and extracts the five-tuple information: protocol type, source IP address, destination IP address, source port, and destination port. This allows it to analyze the packet transmission path and basic attributes. DFI, a significant evolution of DPI in the field of traffic analysis, elevates the analysis dimension to the flow level. This method extracts statistical features such as flow duration, packet length distribution, transmission interval, and traffic cycle, and combines them with the time series patterns of traffic behavior to construct a macroscopic channel characteristic map of traffic. This analysis paradigm based on streaming spatiotemporal features inherits DPI's ability to analyze network behavior. By analyzing the overall behavior patterns of traffic, it effectively avoids the interference of encrypted payloads on content parsing, providing a new technical path for identifying encrypted traffic.
[0004] The problems of the prior art mainly involve:
[0005] Although DFI technology achieves macro-classification of encrypted traffic through statistical features such as flow duration, packet length distribution, and transmission interval, it essentially still regards data packets as independent events with discrete time series and cannot effectively characterize the potential interaction characteristics between data packets within the flow.
[0006] Therefore, it is urgent to develop a solution to solve the above problems. Summary of the Invention
[0007] The purpose of the present invention is to provide an encrypted traffic analysis method based on interactive spatiotemporal features, which addresses the problems that encrypted traffic features are highly concealed, existing deep flow detection technology relies on shallow statistical features, and traditional deep packet inspection technology fails.
[0008] The present invention provides an encrypted traffic analysis method based on interactive spatiotemporal features, which adopts the following technical solutions:
[0009] A method for analyzing encrypted traffic based on interactive spatiotemporal features, comprising the following steps:
[0010] Based on the original encrypted stream, a FITDect model consisting of an input layer, a GNN layer, an MLP layer, and an output layer is constructed;
[0011] The traffic interaction graph is dynamically represented through the input layer of the FITDect model;
[0012] Through hierarchical feature extraction, the GNN layer of the FITDect model is used to analyze the spatiotemporal coupling relationship of the traffic interaction graph;
[0013] The MLP layer of the FITDect model is used to make classification decisions, and the classified results are sent to the output layer of the FITDect model to output the analysis results.
[0014] An encrypted traffic analysis method based on interactive spatiotemporal features is constructed for the original encrypted flow. The method constructs a FITDect model consisting of an input layer, a GNN layer, an MLP layer and an output layer. By mining the spatiotemporal interaction relationship of data packets within the original encrypted flow, a dynamic traffic interaction graph is constructed, and dynamic characterization of the traffic interaction graph is performed. Hierarchical feature extraction is utilized to analyze the spatiotemporal coupling relationship through the GNN layer of the FITDect model, and classification decisions are made using the MLP layer of the FITDect model. The classification results are sent to the output layer, and finally the analysis results are output. The present invention realizes the fusion analysis of the spatiotemporal features of encrypted traffic by constructing the graph structure information and node attribute information of the traffic interaction graph, significantly enhancing the model's ability to characterize the hidden behavior of encrypted traffic, effectively capturing the implicit behavior patterns of encrypted traffic, and providing a new method for malicious encrypted traffic detection.
[0015] Optionally, the input layer of the FITDect model preprocesses the raw traffic data and constructs a graph structure as the input of the GNN layer of the FITDect model; the traffic interaction graph is dynamically represented by the input layer of the FITDect model, which is further expressed as:
[0016] Extract attack type traffic samples from the data set and perform data set sampling; re-represent the labels of different flows in the data set; clear invalid flows in the data set; define interaction actions, and construct a traffic interaction graph based on the defined interaction actions; dynamically characterize the interaction features through the traffic interaction graph.
[0017] Optionally, the defined interaction action is further expressed as:
[0018] A set of data packets with the same flow ID is defined as a Flow. The flow ID is a five-tuple consisting of the source IP address, destination IP address, source port number, destination port number, and protocol type. The sequence of interaction action directions is represented by the flow ID and the data packet set Flow.
[0019] Optionally, the construction of the traffic interaction graph is further expressed as:
[0020] The flows are grouped in chronological order to form a data packet sequence flow; the sequence flow is abstracted into several interactive actions and divided into action slices; the nodes belonging to the same action slice are internally connected; the nodes of adjacent action slices are externally connected.
[0021] Optionally, the invalid flow is a flow with a data packet length less than 5 or a flow without a complete TCP handshake process.
[0022] Optionally, the hierarchical feature extraction is further expressed as:
[0023] Extract direction features and length features from the data packets of the flow; extract interaction features of independent communication behaviors; extract interaction state features.
[0024] Optionally, the GNN layer of the FITDect model is used to analyze the spatiotemporal coupling relationship of the traffic interaction graph, which is further expressed as follows;
[0025] Optionally, the GNN layer of the FITDect model aggregates and extracts the interaction features of the traffic interaction graph; the GNN layer of the FITDect model aggregates and extracts the node attribute features of the traffic interaction graph; the traffic interaction graph is analyzed for spatiotemporal coupling relationships, and an Embedding vector is generated.
[0026] Optionally, a GraphSAGE algorithm is used to parse the spatiotemporal coupling relationship of the traffic interaction graph, and the GraphSAGE algorithm generates node embeddings by sampling and aggregating neighbor node information.
[0027] Optionally, generating an Embedding vector includes adopting a weighted average aggregation strategy to adjust the influence of different nodes by introducing weights; calculating the weighted average of all node vectors after weighting the traffic interaction graph to generate a graph-level Embedding vector.
[0028] Optionally, the MLP layer of the FITDect model consists of three fully connected layers, which perform nonlinear transformation on node embeddings, thereby capturing feature combinations and patterns to make classification decisions; the output layer of the FITDect model maps the classification results processed by the MLP layer to the task space, and outputs the analysis results through the Softmax activator.
[0029] The beneficial effects of the present invention are: an encrypted traffic analysis method based on interactive spatiotemporal features, facing the original encrypted flow, constructing a FITDect model consisting of an input layer, a GNN layer, an MLP layer and an output layer, and constructing a dynamic traffic interaction graph by mining the spatiotemporal interaction relationship of data packets within the original encrypted flow, and dynamically characterizing the traffic interaction graph. By utilizing hierarchical feature extraction, the spatiotemporal coupling relationship is parsed through the GNN layer of the FITDect model, and the MLP layer of the FITDect model is used to make classification decisions, and the classification results are sent to the output layer, and finally the analysis results are output. The present invention realizes the fusion analysis of the spatiotemporal features of encrypted traffic by constructing the graph structure information and node attribute information of the zero-traffic interaction graph, significantly enhancing the model's ability to characterize the hidden behavior of encrypted traffic, and can effectively capture the implicit behavior patterns of encrypted traffic, providing a new method for malicious encrypted traffic detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematic diagram of the FITDect model framework of the encrypted traffic analysis method of the present invention;
[0031] Figure 2 Schematic diagram of action slicing division of the encrypted traffic analysis method of the present invention;
[0032] Figure 3 A schematic diagram of a framework for constructing a traffic interaction graph for the encrypted traffic analysis method of the present invention;
[0033] Figure 4 Schematic diagram of the encrypted traffic communication process of the encrypted traffic analysis method of the present invention;
[0034] Figure 5 Schematic diagram of interactive behavior of the encrypted traffic analysis method of the present invention;
[0035] Figure 6 The figure is a schematic diagram of the encryption analysis process of the encrypted traffic analysis method of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be the common meanings understood by people with ordinary skills in the field to which the invention belongs. The words "including" and similar words used in this article mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0037] The embodiment of the present invention provides an encrypted traffic analysis method based on interactive spatiotemporal features, such as Figure 6 As shown, the following steps are included:
[0038] S1. Based on the original encrypted stream, a FITDect model consisting of an input layer, a GNN layer, an MLP layer, and an output layer is constructed.
[0039] S2,dynamically characterize the traffic interaction graph through the input layer of the FITDect model;
[0040] S3. Through hierarchical feature extraction, the GNN layer of the FITDect model is used to analyze the spatiotemporal coupling relationship of the traffic interaction graph;
[0041] S4. Use the MLP layer of the FITDect model to make classification decisions, and send the classified results to the output layer of the FITDect model to output the analysis results.
[0042] An encrypted traffic analysis method based on interactive spatiotemporal features is constructed for the original encrypted flow. The method constructs a FITDect model consisting of an input layer, a GNN layer, an MLP layer and an output layer. By mining the spatiotemporal interaction relationship of data packets within the original encrypted flow, a dynamic traffic interaction graph is constructed, and dynamic characterization of the traffic interaction graph is performed. Hierarchical feature extraction is utilized to analyze the spatiotemporal coupling relationship through the GNN layer of the FITDect model, and classification decisions are made using the MLP layer of the FITDect model. The classification results are sent to the output layer, and finally the analysis results are output. The present invention realizes the fusion analysis of the spatiotemporal features of encrypted traffic by constructing the graph structure information and node attribute information of the traffic interaction graph, significantly enhancing the model's ability to characterize the hidden behavior of encrypted traffic, effectively capturing the implicit behavior patterns of encrypted traffic, and providing a new method for malicious encrypted traffic detection.
[0043] In the context of the problem that traditional DFI methods have a shallow dependence on encrypted traffic analysis, the FITDect model proposed in this paper uses a graph model to capture and represent the internal interaction patterns between data packets in encrypted traffic.
[0044] Specifically, in some embodiments, Figure 1 As shown in Figure 1, the FITDect model consists of an input layer, a GNN layer, an MLP layer, and an output layer. The input layer first preprocesses the raw traffic data to filter out invalid or short flows, then divides the interaction actions and constructs a homogeneous graph to serve as the input to the GNN layer.
[0045] The core goal of the preprocessing stage is to improve data quality. Raw traffic data often contains a lot of noise, such as invalid flows caused by network jitter or device anomalies, and short flows caused by transmission interruptions or collection errors. If these invalid data are directly input into the model, it will not only take up computing resources, but may also interfere with the model's extraction of normal traffic features. By setting thresholds such as traffic duration and packet integrity, the system can accurately filter out invalid flows, ensuring that subsequent analysis focuses on traffic segments with practical significance. This step is similar to the denoising operation in data cleaning, which can significantly reduce the interference of noise on model training and improve the overall usability of the data set.
[0046] The process of dividing interaction actions and constructing a homogeneous graph is essentially a structured modeling of traffic data. Traffic data is essentially a collection of nodes (such as IP addresses and device IDs) and edges (such as communication records and interaction behaviors), but raw data often exists in an unstructured form. By breaking down traffic interactions into specific actions (such as HTTP requests and DNS queries) and constructing a homogeneous graph based on the interaction relationships between nodes, the system can transform complex traffic patterns into a graph structure. The advantage of this modeling approach is that it simultaneously captures both node attribute characteristics (such as device type and IP location) and edge structural characteristics (such as interaction frequency and temporal relationships), providing richer information input for the GNN layer. For example, in social network analysis, the attention relationships between users can be abstracted as a graph structure, and the device communication relationships in traffic analysis can be treated similarly.
[0047] The GNN layer aggregates and extracts the interaction features of the flow graph and the attribute features of the nodes themselves, generating an embedding vector for each graph. The vector classification results are then sent to the output layer through the MLP layer (Multi-layer Perceptron). The prediction results are finally output through the Softmax activator.
[0048] Specifically, in some embodiments, Figure 1As shown, the input layer of the FITDect model preprocesses the original traffic data and constructs a graph structure as the input of the GNN layer of the FITDect model; the traffic interaction graph is dynamically represented by the input layer of the FITDect model, which is further expressed as:
[0049] Extract attack type traffic samples from the data set and perform data set sampling; re-represent the labels of different flows in the data set; clear invalid flows in the data set; define interaction actions, and construct a traffic interaction graph based on the defined interaction actions; dynamically characterize the interaction features through the traffic interaction graph.
[0050] Specifically, in some embodiments, the training data set sampling is to extract enough attack type traffic samples from a large-scale public data set, such as SQL injection, brute force attack, man-in-the-middle attack, etc. The purpose of this step is to enrich the types of attack traffic as much as possible while ensuring that the scale of the training data set is not too large, so that the model can recognize more attack interaction patterns.
[0051] Label encoding is the re-representation of labels of different streams in a dataset. On the one hand, it converts labels into numerical types to facilitate subsequent model learning and classifier prediction. On the other hand, it also converts multi-classification tasks in the dataset into binary classification tasks.
[0052] The invalid flow is a flow with a data packet length less than 5 or a flow without a complete TCP handshake process.
[0053] The data cleaning step is mainly to remove incomplete flows or flows that have no training significance in the data set, that is, invalid flows. In the present invention, flows with a data packet length less than 5 or flows that do not have a complete TCP handshake process are considered invalid flows, and this step will remove these flows.
[0054] Furthermore, in some embodiments, the definition of the interactive action is further expressed as:
[0055] A set of data packets with the same flow ID is defined as a Flow. The flow ID is a five-tuple consisting of the source IP address, destination IP address, source port number, destination port number, and protocol type. The sequence of interaction action directions is represented by the flow ID and the data packet set Flow.
[0056] Specifically, in order to expose the interaction pattern of the flow in the graph as much as possible, it is necessary to analyze the interaction process of the communication behavior between individuals. In some embodiments, the encryption flow under study is defined as a session, which refers to a session containing several streams with the same flow ID ( ) is a set of data packets, as shown in the following formula:
[0057] ;
[0058] The Flow is a stream (Flow) formed by a sequence of several data packets (N).
[0059] Is a five-tuple consisting of the source IP address ( ), target IP address ( ), source port number ( ), the destination port number ( ), protocol type ( ), as shown in the following formula:
[0060] ;
[0061] Based on the above formula, the data packet set Flow and flow ID ( ) can be extended to the sequence representation of the interaction action direction, and the specific formula is:
[0062] ;
[0063] in, A sequence representing the direction of the interaction action, where the direction of the action is recorded as The value range is {-1,1}. A value of -1 indicates that the interaction is in the downlink direction (server-client), and a value of 1 indicates that the interaction is in the uplink direction (client-server). Indicates the interaction direction of the i-th data packet.
[0064] Based on the sequence representation formula of the interaction action direction, the direction representation of the interaction action in the flow is described. Combined with this representation, the flow interaction process can be further abstracted into a sequence of receiving and sending behaviors. The specific formula is:
[0065] ;
[0066] in, Indicates sending, Acceptance, Represents a sequence of behaviors, Represents the i-th behavior element in the behavior sequence.
[0067] In the present invention, the quintuple is used as the core design of the flow ID to achieve accurate identification of network interactions. The source IP address and destination IP address in the quintuple clearly define the two parties of communication, the source port number and destination port number further refine the specific service or application of the communication, and the protocol type defines the rules and standards followed by the communication. The combination of these five elements makes each Flow an independent and unique interaction instance in the network, which can accurately distinguish different communication sessions. For example, in a complex network environment, multiple devices may communicate with the same server at the same time, but through the distinction of the quintuple, the independent interaction process between each device and the server can be clearly identified, avoiding data confusion and misjudgment. This precise identification provides a reliable basis for subsequent traffic analysis, allowing analysts to conduct in-depth research on specific interaction instances.
[0068] Defining a collection of packets with the same flow ID as a flow enables logical aggregation of network interactions. In real-world network transmission, a complete interaction often consists of multiple packets, which are transmitted across the network according to a specific order and rules. By aggregating packets with the same flow ID into a flow, these scattered packets can be integrated into a coherent whole, providing a more intuitive picture of the entire interaction process. This logical aggregation not only simplifies data complexity but also enables analysts to grasp key information such as the scale, duration, and data volume of the interaction from a macro perspective. For example, when analyzing the transmission of a video stream, by aggregating packets with the same flow ID, the start and end time of the video stream, the amount of data transmitted, and other information can be clearly visualized, enabling assessment of the video stream's transmission quality and user experience.
[0069] In some embodiments, as Figure 2 As shown in Figure 1, the process of associating each action with the interaction process is called an action slice. An action slice consists of the longest continuous segment of action with the same direction during the interaction. At the beginning of an action slice, its direction is always opposite to the previous action slice, and it continues until the direction of the next action slice changes. Except for the initial and final action slices, the action slices during the interaction are characterized by alternating between the same internal direction and the opposite direction.
[0070] Using flow IDs and packet collections (Flows) to represent sequences of interaction directions provides rich semantic information for traffic analysis. These interaction direction sequences reflect the direction of data flow between communicating parties. By analyzing these sequences, we can understand data transmission patterns, service call relationships, and potential abnormal behavior within the network. For example, in normal network communication, a client sends a request packet to a server, and the server returns a response packet upon receiving the request. This request-response interaction pattern is common. However, if analysis reveals a large number of request packets sent from the server to the client within a flow, but no corresponding response from the client, this may indicate abnormal network behavior, such as a server attack or a client vulnerability. Through in-depth analysis of interaction direction sequences, these potential security threats can be promptly identified and prevented.
[0071] In some embodiments, the constructing of the traffic interaction graph is further expressed as:
[0072] The flows are grouped in chronological order to form a data packet sequence flow; the sequence flow is abstracted into several interactive actions and divided into action slices; the nodes belonging to the same action slice are internally connected; the nodes of adjacent action slices are externally connected.
[0073] Furthermore, based on the definition of action slices above, a single-flow interaction graph can be generated by permuting and combining multiple layers of action slices. Edges can be constructed between adjacent slices to express potential connections between adjacent actions. This not only preserves the original time series information but also integrates richer information. It provides an intuitive description of the spatial structure of the interaction process. The following details the construction method of the interactive behavior graph.
[0074] Specifically, in some embodiments, Figure 3 As shown in the figure, based on the analysis of the transition states between interactive actions, the construction of the graph structure is divided into four steps, as follows:
[0075] Step 1, flow grouping step. In this stage, according to Extract the response data packets and group them in chronological order to form a data packet sequence .
[0076] Step 2, divide the action slices. In this stage, first divide the data packet sequence Abstracted as a series of interactive actions, the attribute sequence of these interactive actions is expressed as .by As initial input, the interaction process is expressed as independent slices to represent a single communication behavior.
[0077] Step 3: Intra-slice Connections. In this stage, interactions are first mapped to graph vertices. Based on the above analysis, there are no transitions between interactions within a single slice. Therefore, edges are used to connect these interactions in chronological order. In particular, the attribute of the edge indicating that the transition state has not changed is marked as 0.
[0078] Step 4: Slice External Connection. In this stage, based on the transition state, the interactions within each slice are connected to all interactions in adjacent slices. This step represents the relationship between interactions between different slices. Edges indicating that the transition state has changed are marked as 1.
[0079] Through these steps, the interactive process graph structure of the encryption flow is constructed. It is important to note that the vertices in the graph correspond to the mapping of interactive actions, and each edge represents the transition state between interactive actions. Because the transition state between interactive actions represents a bidirectional process, the edges of the graph are undirected.
[0080] In the actual graph construction process, Figure 4 As shown in the figure, based on the interaction process of a given encrypted stream, an interaction timing diagram between a client and a server is displayed. The diagram presents the interaction process of the two parties in different slices in chronological order. The entire interaction process is divided into six slices, namely Slice 1 to Slice 6, each slice representing a specific stage or time period in the interaction process. This slice-based interaction timing diagram can clearly show the interaction between the client and the server at different stages, including the time point of the interaction, the amount of data that may be involved, and other information. It helps to analyze the timing characteristics, response time, and data transmission status of the interaction between the two parties. It has important reference value for understanding the network interaction process, troubleshooting interaction problems, and optimizing interaction performance.
[0081] like Figure 5 Figure 1 shows an interactive behavior graph of a multi-slice network structure. The graph shows six slices, labeled Slice 1 through Slice 6. Each slice represents a local region in the network. Slices are connected through nodes and edges, forming a complex network topology. Nodes in different slices are interconnected through edges, forming a cross-slice network structure. Node attributes in the graph consist of data packets and time, and graph edge weights are introduced to further represent the degree of static attention paid to edges at different locations. Nodes in Slice 2 are connected to nodes in Slices 3 and 4. These cross-slice connections reflect the interaction and information transfer between different local network regions.
[0082] In this invention, the core purpose of constructing an interactive behavior graph is to reveal the inherent structure of complex interactions. In traffic analysis, data transmission in a network often involves dynamic interactions between multiple entities (such as IP addresses, devices, and users). In the graph construction of this invention, each packet in the flow is considered a node, and edges represent the connections between packets.
[0083] The use of interactive behavior graphs in the present invention can effectively make up for the shortcomings of existing graph learning methods in encrypted traffic analysis methods. Existing graph learning methods have the problem of insufficient granularity in interactive topology modeling. They usually only rely on external metadata such as IP address association or session similarity to build graph structures, but ignore the temporal interaction patterns at the packet level within the encrypted traffic, such as TLS handshake negotiation state transfer, encrypted payload fragmentation behavior sequence, etc. The interactive behavior graph can go deep into the encrypted traffic and capture these packet-level temporal interaction patterns, thereby more accurately characterizing cross-flow collaborative behaviors in APT attacks. In APT attack scenarios, the attacker's activities often involve multiple stages and collaboration between different traffic flows. Ordinary graph learning methods are difficult to effectively identify such complex attack behaviors due to insufficient modeling granularity. The interactive behavior graph, with its ability to capture fine-grained interaction patterns, provides strong support for detecting such attacks.
[0084] Existing graph learning methods also suffer from a disconnect between local and global features, often employing a single-level feature extraction strategy. These methods either focus on the statistical features of a single session or rely on aggregated metrics of global traffic, failing to jointly model fine-grained interaction patterns and macro-behavioral patterns. Fine-grained interaction patterns include the temporal dependencies of packets within a single flow, while macro-behavioral patterns include the logical topology of instructions across multiple flows. This disconnected feature extraction approach limits the model's sensitivity to detecting multi-stage encrypted attack chains. Interaction behavior graphs overcome this limitation by integrating fine-grained interaction patterns and macro-behavioral patterns for modeling. By capturing fine-grained features such as the temporal dependencies of packets within a single flow and depicting macro-features such as the logical topology of instructions across multiple flows, interaction behavior graphs can more comprehensively and accurately reflect the true behavior of network traffic, thereby improving detection capabilities for multi-stage encrypted attack chains and promptly identifying potential security threats.
[0085] Specifically, in some embodiments, the hierarchical feature extraction is further expressed as:
[0086] Extract direction features and length features from the data packets of the flow; extract interaction features of independent communication behaviors; extract interaction state features.
[0087] Furthermore, after processing at the input layer, a single flow is represented by a communication behavior graph, which is achieved based on the state transition of the action slice. Therefore, the communication behavior graph only contains the state transition information of the action slice, that is, the structural characteristics. However, in the action slice transition, the characteristics of the slice itself and the relationship characteristics between adjacent slices are also an important part of describing the communication behavior. Therefore, the present invention proposes the characteristics of the four interaction states of the action slice, specifically:
[0088] The internal state of the slice includes the direction and length extracted from the data packet, which are its basic characteristics; the overall state of the slice describes the characteristics of its action slice as an independent communication behavior in the interaction; the adjacent slice state describes the context characteristics related to the action slice; the adjacent slice comparison is given in the form of a ratio to describe the change of the slice state.
[0089] Specifically, interaction state features are designed to capture various features related to interaction states and consist of a feature name, dimension, and description. The feature names clearly identify different interaction state characteristics, including length, direction, number of slice actions, slice action ratio, slice length, slice length ratio, action length ratio, and average number of bytes in a slice action. These features characterize the interaction state from different perspectives.
[0090] The dimension clearly states the dimension value corresponding to each feature, ranging from 1 to 3. The dimension of length and direction is 1, the dimension of the number of slice actions and slice length is 3, the dimension of the ratio of the number of slice actions and the ratio of slice length is 2, and the dimension of the ratio of action length and the average number of bytes of slice actions is 1. The setting of the dimension reflects the complexity of the feature in the data structure or the number of considerations. The description clearly states the specific meaning of each feature. The length describes the length of each action; the direction represents the direction of each action; the number of slice actions refers to the specific number of actions contained in the slice; the ratio of the number of slice actions reflects the proportional relationship between the number of interactive actions between a slice and its adjacent slices; the slice length represents the total length of all action packets in the slice; the slice length ratio is the ratio of the total length of action packets between a slice and its adjacent slices; the action length ratio represents the proportion of the length of a single action in the total length of the slice; and the average number of bytes of slice actions is the average length of actions within a single slice.
[0091] Specifically, in some embodiments, the GNN layer of the FITDect model is used to analyze the spatiotemporal coupling relationship of the traffic interaction graph, which is further expressed as follows;
[0092] The GNN layer of the FITDect model aggregates and extracts the interaction features of the traffic interaction graph; the GNN layer of the FITDect model aggregates and extracts the node attribute features of the traffic interaction graph; the spatiotemporal coupling relationship of the traffic interaction graph is analyzed, and an Embedding vector is generated.
[0093] In this paper, the GNN layer aggregates and extracts interaction features and node attribute features from the traffic interaction graph. This process allows for deep mining of the rich information contained within the traffic interaction graph. Interaction features reflect the patterns and regularities of interactions between different nodes, such as the frequency and intensity of interactions, while node attribute features characterize the characteristics of each node, such as its type and status. Through the aggregation and extraction operations of the GNN layer, the model can more comprehensively understand the relationships and attributes between the various elements in the traffic interaction graph, providing a more accurate and comprehensive basis for subsequent analysis and judgment.
[0094] Analyzing the spatiotemporal coupling relationships in the traffic interaction graph and generating an embedding vector further enhances the model's ability to process traffic data. This spatiotemporal coupling relationship considers the correlation between traffic interactions in both time and space. The temporal dimension captures the temporal trends of interactions, while the spatial dimension captures the spatial layout and mutual influence between different nodes. Generating an embedding vector represents the complex spatiotemporal coupling relationships obtained through analysis in a concise and efficient vector form. This vector representation offers numerous advantages, including ease of computer processing and computation, significantly reducing the dimensionality and complexity of data while preserving key information from the traffic interaction graph. In subsequent model training and prediction tasks, the embedding vector can be used as input features, enabling the model to better learn and identify patterns and regularities in traffic data, improving its accuracy and generalization, and ultimately enhancing the performance and effectiveness of the FITDect model in relevant application scenarios.
[0095] Furthermore, the spatiotemporal coupling relationship of the traffic interaction graph is parsed using the GraphSAGE algorithm, which generates node embeddings by sampling and aggregating neighbor node information.
[0096] Specifically, in order to effectively extract useful information from graph structure data, the sampling aggregation-based GraphSAGE algorithm is used to generate graph embeddings. The GraphSAGE algorithm is suitable for small and medium-sized graphs and also allows learning embedding representations of unseen graph structures.
[0097] The GraphSAGE algorithm generates node embeddings by sampling and aggregating neighbor node information. This approach not only reduces computational complexity but also allows embedding learning for unseen nodes. However, its performance is highly dependent on the choice of sampling strategy. In some embodiments, GraphSAGE uses uniform sampling, which is expressed as follows:
[0098] ;
[0099] in, represents the node set sampled from the neighbor node set of node v, is a sampling function that selects a specified number of elements from a given set according to certain rules. is the number of samples per layer, which can be set in this paper v represents the central node currently being processed. This set represents all nodes u in the graph that are connected to node v by edges, that is, v's direct neighbors; is the set of all nodes in the graph.
[0100] In addition, the GCN algorithm can also be used to analyze the spatiotemporal coupling relationship of the traffic interaction graph.
[0101] The GCN algorithm learns node representations by applying convolution operations on the adjacency matrix of the graph. This method can effectively capture local structural information, but its computational complexity is high, especially when processing large-scale graphs. The process formula for updating the node representation of the first layer of the GCN algorithm is:
[0102] ;
[0103] in, , is the adjacency matrix with self-connection added, is the adjacency matrix of the graph, is the identity matrix, for The degree matrix of is the learnable weight matrix, is the feature matrix. At the same time, in some embodiments, the features of the stacked two layers of convolution can be enhanced later. The specific formula is expressed as:
[0104] ;
[0105] Among them, the Softmax function is used for the output layer of multi-classification problems. (adding the self-connected adjacency matrix), and are all learnable weight matrices, is the feature matrix of the nodes in the graph, It is an activation function used to introduce nonlinear factors so that the model can learn more complex feature representations.
[0106] Furthermore, the GAT algorithm may be used to analyze the spatiotemporal coupling relationship of the traffic interaction graph.
[0107] The GAT algorithm uses the attention mechanism to dynamically assign weights to the neighbors of each node, thereby better capturing important connection relationships in the graph. Although it is highly flexible and can handle complex graph structures, it may face the risk of overfitting during training.
[0108] In this paper, the GraphSAGE algorithm dynamically captures the spatiotemporal dependencies between network nodes, improving our understanding and prediction of complex interaction patterns. The core purpose of this step is to combine the spatial proximity of nodes in the traffic interaction graph with their temporal evolution characteristics to construct an embedded representation that reflects multidimensional relationships, providing a more expressive data foundation for subsequent analysis.
[0109] In a traffic interaction graph, nodes typically represent network entities (such as servers, user terminals, or routers), while edges represent data transmission or interaction between entities. Traditional methods often analyze spatial structures or time series in isolation, making it difficult to capture the coupling effect between the two. During peak hours, the traffic surge at certain nodes may not only be related to the load of their direct neighbors, but also be affected by the behavior of remote nodes within the historical time window. GraphSAGE, through neighbor sampling and feature aggregation mechanisms, can dynamically integrate information from multi-hop neighbors in a local scope, while combining feature encoding in the time dimension to generate a spatiotemporal-aware embedding of the node. This embedding not only contains the static properties of the node, but also incorporates the dynamic changes of its neighbors within the time window, enabling the model to understand the rule that "spatially close nodes exhibit similar behaviors within a specific time window."
[0110] Furthermore, GraphSAGE's aggregation functions are designed as learnable modules, automatically adjusting feature combinations based on specific tasks. In traffic interaction scenarios, different types of data transmission (such as video streaming and file downloads) may have varying sensitivities to spatiotemporal coupling. By optimizing the aggregation function, the algorithm can more accurately capture these differences. For interactions requiring high real-time performance, the algorithm may prioritize local features along the temporal dimension; whereas for long-term trend analysis, it may prioritize global information integration along the spatial dimension.
[0111] The advantages of the GraphSAGE algorithm lie in its flexibility and scalability. Through a sampling mechanism, the algorithm can handle large-scale traffic interaction graphs, avoiding the memory and computational overhead of full-graph calculations. Furthermore, the diversity of aggregation functions (such as mean aggregation and LSTM aggregation) allows the algorithm to adapt to different types of spatiotemporal coupling patterns. For example, when faced with traffic interactions with periodic characteristics, LSTM aggregation can better capture long-term dependencies in time series; whereas, in scenarios where node features vary widely, mean aggregation or pooling aggregation may be more efficient. Furthermore, GraphSAGE's inductive learning capabilities enable the model to handle unseen nodes or subgraphs, which is particularly important in dynamically changing network environments. When new nodes join or the behavior of existing nodes changes, the algorithm does not need to retrain the entire model; instead, it can generate new embedding representations based on local sampling and aggregation, greatly improving the model's practicality and adaptability.
[0112] Specifically, in some embodiments, generating an Embedding vector includes adopting a weighted average aggregation strategy to adjust the influence of different nodes by introducing weights; calculating the weighted average of all node vectors after weighting the traffic interaction graph to generate a graph-level Embedding vector.
[0113] In this paper, the FITDect model adopts a weighted average aggregation strategy, applying different weights to edges inside and outside the action slice to achieve the best effect. Compared with the four aggregation strategies of mean aggregation, maximum aggregation, weighted average aggregation, and LSTM aggregation, mean aggregation is simple and direct, suitable for capturing global features; maximum aggregation emphasizes the most important neighbor information, which helps to discover anomalies or key nodes; weighted average aggregation combines the advantages of the first two, and by introducing weights, it can more flexibly adjust the influence of different neighbors, which is suitable for most scenarios; LSTM aggregation uses the sequence model LSTM to model the sequential pattern of node embedding to generate a global embedding of the graph.
[0114] The weighted average aggregation strategy used to generate graph-level embedding vectors for the traffic interaction graph aims to integrate node features through differentiated weight assignments, thereby capturing key information and overall patterns within the graph structure. The core purpose of this step is to aggregate local node-level features into a vector representation that reflects global characteristics, providing a foundation for subsequent classification, prediction, or analysis tasks.
[0115] In a traffic interaction graph, the importance of different nodes often varies significantly. Core routers or servers may carry a large amount of critical traffic, and their characteristics have a much greater impact on the overall network state than edge nodes. Traditional unweighted average aggregation methods treat all nodes equally, diluting the information of important nodes and obscuring key patterns through noise from secondary nodes. Weighted average aggregation introduces weights to dynamically adjust the contribution of nodes to the final embedding. Weights can be assigned based on metrics such as degree centrality (number of connections), betweenness centrality (criticality of information transfer), or traffic load, so that the characteristics of core nodes dominate the aggregation process. This mechanism enables the generated graph-level embeddings to more accurately reflect the core characteristics of the traffic interaction graph and avoid interference from irrelevant information.
[0116] Compared to the multiple rounds of message passing in complex graph neural networks, weighted averaging requires only a single aggregation operation to generate a graph-level representation, significantly reducing computational complexity. In real-time traffic monitoring scenarios, rapidly generating graph-level embeddings is crucial for timely responding to network anomalies. Furthermore, the weight distribution mechanism provides interpretability. By analyzing the weight distribution, it is possible to intuitively identify which nodes have the greatest impact on the overall network state. Furthermore, the interpretability of weights enhances model transparency. By analyzing the weight distribution, it is possible to intuitively identify which nodes contribute most to the overall embedding, providing a basis for network optimization or anomaly location. Furthermore, the weight distribution mechanism of weighted averaging is inherently interpretable. By analyzing the relationship between weights and node attributes, it is possible to intuitively understand which factors (such as traffic load and number of connections) contribute most to graph-level embeddings.
[0117] The weighted average aggregation strategy is also robust to noise and outliers. In a traffic interaction graph, abnormal data from individual nodes (such as bursts or measurement errors) can interfere with global analysis. By properly designing weights (such as confidence scores based on historical node behavior), the impact of abnormal nodes on embeddings can be reduced. Nodes with frequent abnormal traffic can have their weights reduced or even excluded, ensuring the stability of the graph-level representation. This mechanism makes the algorithm more practical in real-world network environments.
[0118] Furthermore, in some embodiments, the MLP layer of the FITDect model is composed of three fully connected layers, which perform nonlinear transformation on node embeddings, thereby capturing feature combinations and patterns to make classification decisions; the output layer of the FITDect model maps the classification results processed by the MLP layer to the task space, and outputs the analysis results through the Softmax activator.
[0119] Specifically, the MLP layer of the FITDect model contains a three-layer neural network structure of input layer, hidden layer, and output layer. The task type should be binary classification (malicious or benign), so the softmax activation function will be used in this layer to compress the output to the (0, 1) interval to represent the probability of the positive class. The three-layer forward propagation formula can be expressed as:
[0120] ;
[0121] in, , represents the input feature vector, that is, the graph Embedding vector, , , , represents the weight matrix of the fully connected layer; , , , represents the bias vector; is the ReLU activation function, defined as , the ReLU activation function is used to introduce nonlinear factors, so that the neural network can learn more complex feature representations; is the Sigmoid activation function, expressed as , used to describe the probability of the positive class.
[0122] In the loss function, since the classification scenario is binary, binary cross entropy is used here, and L2 regularization is introduced to prevent overfitting. The loss function formula can be expressed as:
[0123] ;
[0124] in, is the number of batch samples; , represents the true label, 0 represents normal traffic, and 1 represents abnormal traffic; , represents the predicted probability; represents the regularization coefficient, =0.001; is the weight decay term.
[0125] In this paper, the FITDect model architecture focuses on solving encrypted traffic analysis by combining graph methods to extract interactive behavior features (spatiotemporal dimensions). The FITDect model can analyze data at the flow level, focusing on the timing characteristics and payload interaction patterns within a single data flow. Graph learning-based methods often require aggregating multi-flow session information, making real-time detection difficult. Intra-flow interaction analysis, however, can identify traffic at the granularity of a single TCP / UDP flow. The FITDect model also decouples the multidimensional spatiotemporal characteristics of flows for comprehensive analysis. The proposed analysis method improves upon traditional graph learning frameworks by jointly modeling the graph structure (spatial dimension) and node attribute information (temporal dimension) of the traffic interaction graph, enabling a fused analysis of the spatiotemporal characteristics of encrypted traffic. Compared to traditional graph learning methods that rely solely on static topology (such as fixed graphs based on IP relationships) or isolated node features (such as single-flow statistics), this spatiotemporal coupled modeling significantly enhances the model's ability to characterize the hidden behaviors of encrypted traffic.
[0126] In some implementations, dynamic interaction graphs are constructed based on three real-world encrypted traffic benchmark datasets: CIC-IDS 2017, CCCS-CIC-AndMal-2020, and MalwareTraffic Analysis. Experiments are conducted to first compare the optimal aggregation strategy for GNN layers and the optimal GNN algorithm. Finally, traditional DFI methods and recent GNN variants are used as baseline algorithms to demonstrate the model's detection performance.
[0127] In some embodiments, the system is run on a single desktop host with a hardware environment consisting of a 16-core Intel(R) Xeon(R) Gold 6210R CPU, an NVIDIA GeForce 3070TI GPU, and 16GB of RAM. The software environment includes the Ubuntu 22.04 desktop operating system, Python 3.10.14, the deep learning framework Pytorch 2.2.2, and the graph neural network modeling library Torch-Geometric 2.5.3.
[0128] Specifically, in the embodiment, typical classification evaluation indicators in the field of information retrieval and machine learning can be used, including accuracy, precision, recall, F1 score and AUC score:
[0129] ;
[0130] ;
[0131] ;
[0132] ;
[0133] ;
[0134] Among them, TP is the true positive example, which indicates the number of samples that are actually positive and correctly predicted as positive; TN is the true negative example, which indicates the number of samples that are actually negative and correctly predicted as negative; FP is the false positive example, which indicates the number of samples that are actually negative but incorrectly predicted as positive; FN is the false negative example, which indicates the number of samples that are actually positive but incorrectly predicted as negative; Accuracy (AC) represents the ratio of the number of samples correctly classified by the model to the total number of samples, reflecting the overall accuracy of the classification; Precision (PR) represents the ratio of the number of samples correctly classified as positive by the model to the total number of true positive samples, reflecting the completeness of the positive class coverage; It stands for recall rate, which indicates the proportion of samples that are actually positive that are correctly predicted as positive by the model, reflecting the model's ability to find all positive samples; F1 value represents the harmonic mean of precision and recall, which comprehensively reflects precision and recall; FPR stands for false positive rate, which indicates the proportion of samples that are actually negative that are incorrectly predicted as positive, and is often used to draw ROC curves; AUC value refers to the area under the ROC curve. The ROC curve shows the performance of the classifier by plotting the relationship between true positive rate and false positive rate under different thresholds.
[0135] In order to systematically analyze the impact of different graph learning algorithms on model performance, three spatial domain graph neural network algorithms - GCN, GraphSAGE, and GAT are run separately, and their evaluation indicators are comprehensively measured.
[0136] The resulting F1 score comparison curve shows that GraphSAGE's F1 score remains high from the start and converges after 20 epochs of training. However, the F1 scores of both GCN and GAT are unstable at lower epochs, and their F1 convergence levels are not as high as GraphSAGE's. This disadvantage of the GAT method is particularly evident.
[0137] The loss function curves for the training and test sets show that GraphSAGE performs well in controlling prediction error, with the fastest convergence and high stability. GCN also maintains a relatively good prediction error, but the curve on the test set exhibits a peak. The GAT method exhibits significant fluctuations in the first 10 rounds, and while it also converges roughly, its stability cannot be guaranteed.
[0138] Based on the above experiments, it can be determined that GraphSAGE is the best learning algorithm for interactive behavior graphs, thanks to GraphSAGE's good sampling and aggregation strategies. Specifically, the GraphSAGE sampling strategy adopted by the present invention is based on the idea of weighted averaging. For each node, the action slice in which it is located is used as the anchor point, and the information of all nodes is collected from the three slices above and below, all of which are regarded as sampling nodes. At the same time, weights are assigned to the points to be sampled according to the slice distance. Under such logic, the sampling points in adjacent slices have the highest weights, while the weights of the sampling points separated by slices decrease. In this way, the interactive features embodied in the context of the action slices are strengthened as much as possible, and the algorithm effect is optimized.
[0139] In the present invention, for the FITDect model, the optimal aggregation strategy for the interaction behavior graph is WAvg (weighted average aggregation), and the weight distribution should be more biased towards the nodes on the edge of the slice (this paper adopts the distribution strategy of ɛ=0.7, ɵ=0.3). Compared with the GAT and GCN algorithms, GraphSAGE is more suitable for learning the interaction behavior graph constructed in this paper. GraphSAGE has the fastest learning speed and the most stable effect. This may be due to GraphSAGE's multi-layer sampling and reasonable aggregation strategy, which can better predict graph patterns that have not been seen in the training phase and has strong robustness. The FITDect model achieved good results in the mixed data set constructed by the present invention, and the F1 score and AUC comprehensive indicators are on par with or ahead of traditional machine learning methods such as FlowLens and TSCRNN and new graph neural network methods such as GraphDApp and IBGC.
[0140] An encrypted traffic analysis method based on interactive spatiotemporal features is constructed for the original encrypted flow. The method constructs a FITDect model consisting of an input layer, a GNN layer, an MLP layer and an output layer. By mining the spatiotemporal interaction relationship of data packets within the original encrypted flow, a dynamic traffic interaction graph is constructed, and dynamic characterization of the traffic interaction graph is performed. Hierarchical feature extraction is utilized to parse the spatiotemporal coupling relationship through the GNN layer of the FITDect model, and classification decisions are made using the MLP layer of the FITDect model. The classification results are sent to the output layer, and finally the analysis results are output. The present invention realizes the fusion analysis of the spatiotemporal features of encrypted traffic by constructing the graph structure information and node attribute information of the zero-traffic interaction graph, significantly enhancing the model's ability to characterize the hidden behavior of encrypted traffic, effectively capturing the implicit behavior patterns of encrypted traffic, and providing a new method for malicious encrypted traffic detection.
[0141] While the embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations of these embodiments are possible. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as set forth in the claims. Furthermore, the invention described herein is susceptible to other embodiments and may be practiced or implemented in a variety of ways.
Claims
1. A method for analyzing encrypted traffic based on interactive spatiotemporal features, characterized in that: include: Based on the original encrypted stream, a FITDect model consisting of an input layer, a GNN layer, an MLP layer, and an output layer is constructed; The input layer of the FITDect model preprocesses the original traffic data and constructs a graph structure as the input of the GNN layer of the FITDect model, and dynamically represents the traffic interaction graph through the input layer of the FITDect model; Through hierarchical feature extraction, the GNN layer of the FITDect model is used to analyze the spatiotemporal coupling relationship of the traffic interaction graph; Use the MLP layer of the FITDect model to make classification decisions, and send the classified results to the output layer of the FITDect model to output the analysis results; The dynamic characterization of the traffic interaction graph through the input layer of the FITDect model is further expressed as follows: extracting attack type traffic samples from the data set and performing data set sampling; Re-represent the labels of different streams in the dataset; Clear invalid flows in the data set; define interaction actions, and construct a flow interaction graph based on the defined interaction actions; dynamically characterize interaction features through the flow interaction graph; The construction of the traffic interaction graph is further represented as follows: grouping flows in chronological order to form a data packet sequence flow; Abstract the sequence flow into several interactive actions and divide them into action slices; connect the nodes belonging to the same action slice internally; and connect the nodes of adjacent action slices externally; The hierarchical feature extraction is further expressed as: extracting direction features and length features from the data packets of the flow; Extract interaction features of independent communication behaviors; extract interaction state features; The interaction state features include slice internal state, slice overall state, adjacent slice state and adjacent slice comparison.
2. The encrypted traffic analysis method based on interactive spatiotemporal features according to claim 1 is characterized in that: The defined interaction action is further expressed as: A flow is defined as a set of packets with the same flow ID, where the flow ID is a five-tuple consisting of the source IP address, destination IP address, source port number, destination port number, and protocol type. The sequence of interaction action directions is represented by the flow ID and the data packet set Flow.
3. The encrypted traffic analysis method based on interactive spatiotemporal features according to claim 1 is characterized in that: The invalid flow is a flow with a data packet length less than 5 or a flow without a complete TCP handshake process.
4. The encrypted traffic analysis method based on interactive spatiotemporal features according to claim 1 is characterized in that: The GNN layer of the FITDect model is used to analyze the spatiotemporal coupling relationship of the traffic interaction graph, which is further expressed as follows; The GNN layer of the FITDect model aggregates and extracts interaction features of the traffic interaction graph; The GNN layer of the FITDect model aggregates and extracts node attribute features of the traffic interaction graph; The time-space coupling relationship of the traffic interaction graph is analyzed, and an Embedding vector is generated.
5. The encrypted traffic analysis method based on interactive spatiotemporal features according to claim 4 is characterized in that: The GraphSAGE algorithm is used to parse the spatiotemporal coupling relationship of the traffic interaction graph. The GraphSAGE algorithm generates node embeddings by sampling and aggregating neighbor node information.
6. The encrypted traffic analysis method based on interactive spatiotemporal features according to claim 4 is characterized in that: Generating an Embedding vector includes adopting a weighted average aggregation strategy to adjust the influence of different nodes by introducing weights; The weighted average of all node vectors after weighting in the traffic interaction graph is calculated to generate a graph-level embedding vector.
7. The encrypted traffic analysis method based on interactive spatiotemporal features according to claim 1 is characterized in that: The MLP layer of the FITDect model consists of three fully connected layers, which perform nonlinear transformations on node embeddings to capture feature combinations and patterns for classification decisions; The output layer of the FITDect model maps the classification results processed by the MLP layer to the task space and outputs the analysis results through the Softmax activator.
Citation Information
Patent Citations
Method for predicting time sequence network link by adopting GraphSAGE
CN111756587A
Encrypted traffic classification method in IPSec tunnel mode based on spatio-temporal information
CN118656690A