Cross-domain malicious traffic detection method based on graph model

Through the cross-domain malicious traffic detection method based on graph model, the false alarm and missed reporting problems of traditional methods in cross-domain detection are solved, the detection efficiency and accuracy are improved, and the real-time requirements are met.

CN120281556AInactive Publication Date: 2025-07-08SUZHOU ZHUOMING INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510579102.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional malicious traffic detection methods in a single domain are difficult to effectively deal with cross-domain malicious traffic, and are inefficient in large-scale network traffic data processing, which cannot meet real-time requirements.

Method used

A cross-domain malicious traffic detection method based on graph models is adopted, through data acquisition, preprocessing, feature extraction, graph construction, graph embedding and abnormal detection model training, combined with distributed acquisition architecture and optimization strategies, a directed graph model is built and low-dimensional vector mapping is carried out, and a support vector machine, neural network or decision tree model is used for detection.

Benefits of technology

It improves the accuracy and efficiency of cross-domain malicious traffic detection, reduces false alarms and missed reports, and meets the detection requirements in network environments with real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281556A_ABST
    Figure CN120281556A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain malicious traffic detection method based on a graph model, and the method specifically comprises the following steps: S1, data collection: collecting network traffic data from a plurality of different network domains, the network traffic data comprising a source IP address, a destination IP address, a port number, a protocol type, a traffic size and timestamp information; s2, data preprocessing: cleaning the collected network flow data, removing invalid data and repeated data, and performing standardization processing on the data; the invention relates to the technical field of network security, and the cross-domain malicious traffic detection method based on the graph model can collect network traffic data from a plurality of different network domains by adopting a distributed collection architecture, comprehensively considers various characteristics and connection relationships of cross-domain network traffic by constructing the detection method based on the graph model, and improves the detection accuracy of the cross-domain malicious traffic. The problem that cross-domain malicious traffic is difficult to detect is effectively solved, and the detection capability of complex cross-domain network attacks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and particularly to a cross-domain malicious traffic detection method based on a graph model. Background Art

[0002] With the rapid development of the Internet, network attack means have become increasingly complex and diverse, and malicious traffic detection has become an important task to ensure network security.

[0003] In a cross-domain network environment, due to different network domains having different network structures, application scenarios, and security policies, traditional single-domain malicious traffic detection methods are difficult to effectively meet the detection requirements of cross-domain malicious traffic. For example, some malicious traffic will evade detection by jumping between multiple network domains, disguising, etc. Traditional rule-based or single-feature detection methods are prone to false negatives or false positives. Moreover, many existing malicious traffic detection methods are inefficient in processing large-scale network traffic data and cannot meet the requirements of a network environment with high real-time requirements. Therefore, the present invention provides a cross-domain malicious traffic detection method based on a graph model. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a cross-domain malicious traffic detection method based on a graph model, which can effectively process cross-domain network traffic data and improve the accuracy and efficiency of malicious traffic detection.

[0005] To achieve the above object, the present invention is realized through the following technical solutions: A cross-domain malicious traffic detection method based on a graph model, specifically including the following steps:

[0006] S1. Data collection: Collect network traffic data from multiple different network domains, and the network traffic data includes source IP address, destination IP address, port number, protocol type, traffic size, and timestamp information.

[0007] S2. Data preprocessing: Clean the collected network traffic data, remove invalid data and duplicate data, and perform standardization processing on the data to convert various types of data into a unified format for subsequent analysis.

[0008] S3. Feature extraction: Extract features related to traffic behavior from the preprocessed network traffic data, including connection duration feature, connection frequency feature, traffic burst feature, port usage feature, and protocol distribution feature.

[0009] S4. Graph construction: Use the extracted features as nodes, and construct a directed graph model according to the source IP address, destination IP address, and connection relationship in the network traffic data, where the edges between nodes represent network connection relationships, and the weights of the edges represent metrics related to the frequency of connection or traffic size.

[0010] S5. Graph embedding: A graph embedding algorithm is used to map the constructed directed graph model to a low-dimensional vector space to obtain a low-dimensional vector representation of each node. The graph embedding algorithm includes a Node2vec algorithm or a GraphEmbedding algorithm.

[0011] S6. Anomaly detection model training: Use low-dimensional vectors of graph nodes corresponding to known malicious traffic and normal traffic data as training data to train an anomaly detection model, wherein the anomaly detection model includes a support vector machine, a neural network or a decision tree model.

[0012] S7. Detection: The low-dimensional vector of the node obtained after the network traffic data to be detected is input into the trained anomaly detection model after data collection, preprocessing, feature extraction, graph construction, and graph embedding steps to determine whether the network traffic is malicious traffic.

[0013] S8. Result output: Output the detection result. If malicious traffic is detected, an alarm message is generated and relevant traffic information is recorded. If it is normal traffic, it is allowed to be transmitted normally.

[0014] Preferably, in S1, a distributed collection architecture is adopted, and collection probes are deployed at key network nodes in different network domains to realize the collection of large-scale network traffic data.

[0015] Preferably, in S2, invalid data and duplicate data are removed by setting data validity rules and data deduplication algorithms, wherein the data validity rules are set based on network protocol specifications and a reasonable range of traffic data.

[0016] Preferably, in S3, the connection duration feature is obtained by calculating the average value and standard deviation of the connection duration between the same source IP and destination IP; the connection frequency feature is determined by counting the number of connections between the source IP and the destination IP per unit time; and the traffic burst feature is extracted by analyzing the rapid changes in traffic in a short period of time.

[0017] Preferably, in S4, different types of feature nodes are distinguished and marked with different colors or shapes for intuitively observing and analyzing the graph model structure.

[0018] Preferably, in S5, when using the graph embedding algorithm, the algorithm parameters are optimized and adjusted according to the characteristics of the network traffic data to improve the accuracy and efficiency of the graph embedding.

[0019] Preferably, in S6, a cross-validation method is used to train and evaluate the anomaly detection model, and the model parameters and model structure with the best performance are selected.

[0020] Preferably, in S8, the detection results are stored in a database, and a detection report is generated regularly. The report content includes the type, source, time distribution of malicious traffic, and detection accuracy information.

[0021] Advantageous Effects

[0022] The present invention provides a cross - domain malicious traffic detection method based on a graph model. Compared with the prior art, it has the following advantageous effects:

[0023] (1) This cross - domain malicious traffic detection method based on a graph model can collect network traffic data from multiple different network domains by adopting a distributed acquisition architecture. By constructing a detection method based on a graph model, it comprehensively considers various features and connection relationships of cross - domain network traffic, effectively solves the problem of difficult detection of cross - domain malicious traffic, and improves the detection ability for complex cross - domain network attacks.

[0024] (2) This cross - domain malicious traffic detection method based on a graph model extracts multi - dimensional traffic behavior features, maps the graph model to a low - dimensional vector space using a graph embedding algorithm, and then combines with an anomaly detection model for training and detection. It can more accurately identify malicious traffic, reduce false positives and false negatives. Moreover, optimization strategies are adopted in various links such as data acquisition, pre - processing, graph construction and embedding, and model training, such as distributed acquisition, parameter optimization, etc., improving the processing efficiency of the entire detection method and meeting the malicious traffic detection requirements in a network environment with high real - time requirements. Brief Description of the Drawings

[0025] Figure 1 It is a flowchart of the traffic detection method of the present invention. Detailed Embodiments

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0027] Please refer to Figure 1 The present invention provides a technical solution:

[0028] A cross - domain malicious traffic detection method based on a graph model specifically includes the following steps:

[0029] S1. Data acquisition: Collect network traffic data from multiple different network domains. The network traffic data includes source IP address, destination IP address, port number, protocol type, traffic size, and timestamp information.

[0030] S2. Data preprocessing: Clean the collected network traffic data, remove invalid and duplicate data, standardize the data, and convert various types of data into a unified format for subsequent analysis.

[0031] S3. Feature extraction: Extract features related to traffic behavior from the preprocessed network traffic data, including connection duration features, connection frequency features, traffic burst features, port usage features, and protocol distribution features.

[0032] S4. Graph construction: Using the extracted features as nodes, a directed graph model is constructed based on the source IP address, destination IP address, and connection relationship in the network traffic data. The edges between nodes represent the network connection relationship, and the weight of the edge represents the frequency of the connection or a measure related to the traffic size.

[0033] S5. Graph embedding: A graph embedding algorithm is used to map the constructed directed graph model to a low-dimensional vector space to obtain a low-dimensional vector representation of each node. The graph embedding algorithm includes the Node2vec algorithm or the GraphEmbedding algorithm.

[0034] S6. Anomaly detection model training: Use low-dimensional vectors of graph nodes corresponding to known malicious traffic and normal traffic data as training data to train anomaly detection models. Anomaly detection models include but are not limited to support vector machines, neural networks, or decision tree models.

[0035] S7. Detection: The low-dimensional vector of the node obtained after the network traffic data to be detected is input into the trained anomaly detection model after data collection, preprocessing, feature extraction, graph construction, and graph embedding steps to determine whether the network traffic is malicious traffic.

[0036] S8. Result output: Output the detection result. If malicious traffic is detected, an alarm message is generated and relevant traffic information is recorded. If normal traffic is detected, it is allowed to be transmitted normally.

[0037] In the embodiment of the present invention, in S1, a distributed collection architecture is adopted, and collection probes are deployed at key network nodes in different network domains to realize the collection of large-scale network traffic data.

[0038] Specifically, by adopting a distributed collection architecture, it aims to comprehensively and efficiently obtain network traffic data from multiple different network domains. Specifically, specially designed collection probes are deployed at key network nodes in different network domains. These network domains cover diverse network environments such as enterprise internal network domains, cloud service network domains, and public network domains.

[0039] In the enterprise internal network domain, collection probes are deployed at key nodes such as the border router connecting the internal network and the external network and the core switch connecting each department subnet. These probes are carefully configured to deeply capture all network data packets flowing through this node. For the cloud service network domain, collection devices are set at the virtual network border gateway of the cloud platform and the network isolation nodes of different tenants to ensure accurate capture of the traffic within the cloud environment. In the public network domain, collection probes are deployed at positions such as the backbone network nodes of Internet service providers (ISPs) and the access routers of data centers to obtain large-scale public network traffic information.

[0040] After capturing the network data packets, the collection probes immediately start to extract rich information from them, including key information such as source IP address, destination IP address, port number, protocol type, traffic size, and accurate timestamp. For example, when a network data packet passes through the collection probe at the enterprise internal network border router, the probe quickly analyzes the header information of the data packet and accurately extracts the source IP address (such as 192.168.1.10), destination IP address (such as 202.108.35.210), port number (such as port 80), protocol type (such as TCP), traffic size (such as 1024 bytes), and timestamp (accurate to milliseconds, such as 2024-11-25 10:30:25.500).

[0041] The collected data is transmitted to the centralized processing server through a pre-built secure channel. The secure channel uses a high-strength encryption algorithm (such as the AES encryption algorithm) to encrypt and transmit the data to ensure the confidentiality and integrity of the data during transmission. At the same time, the secure channel has a reliable transmission protocol (such as a custom reliable transmission protocol based on TCP), which can automatically retransmit the data in case of network fluctuations or brief interruptions to ensure the stable transmission of the data. For example, even when there is a brief congestion in the enterprise internal network resulting in some data transmission delays or losses, the secure channel can detect and retransmit the lost data in a timely manner to ensure that the centralized processing server can receive all the collected network traffic data completely for subsequent in-depth processing.

[0042] In the embodiment of the present invention, in S2, invalid data and duplicate data are removed by setting data validity rules and data deduplication algorithms, where the data validity rules are set based on network protocol specifications and the reasonable range of traffic data.

[0043] Specifically, after the centralized processing server receives the collected network traffic data, it first performs data cleaning operations. By setting data validity rules, for example, removing data with illegal source IP addresses or destination IP addresses (such as private IP addresses appearing in public network traffic), and abnormal traffic sizes (such as being too large or too small beyond the normal range). At the same time, using a data deduplication algorithm to remove duplicate network connection records. Then, perform standardization processing on the cleaned valid data, uniformly convert timestamps in different formats to the standard time format, and uniformly convert the traffic size to a specific unit (such as bytes), etc., for subsequent analysis and processing.

[0044] In the embodiments of the present invention, in S3, the connection duration feature is obtained by calculating the average value and standard deviation of the connection duration between the same source IP and destination IP; the connection frequency feature is determined by counting the number of connections between the source IP and destination IP per unit time; the traffic burst feature is extracted by analyzing the sharp change of traffic in a short period of time.

[0045] Specifically, extract features related to traffic behavior from the preprocessed network traffic data:

[0046] 1) Connection duration feature: For each connection record between a pair of source IP and destination IP, calculate its connection duration, and count the average value and standard deviation of the connection duration between the same pair of IPs. For example, if it is found that the average value of the connection duration between a certain source IP and a specific destination IP is much higher than the average value of the normal connection duration and the standard deviation is large, there may be an abnormality.

[0047] 2) Connection frequency feature: Count the number of connections between each source IP and different destination IPs per unit time (such as per minute). If a certain source IP establishes connections with a large number of different destination IPs in a short period of time, this may be a characteristic of malicious scanning behavior.

[0048] 3) Traffic burst feature: Analyze the sharp change of traffic in a short period of time (such as within a few seconds). By calculating the traffic change rate, if it is found that the traffic of a certain connection suddenly increases to an abnormal value within an extremely short period of time, it may be a sign of malicious data transmission.

[0049] 4) Port usage feature: Record the port numbers used by each connection, and count the usage frequencies of different port numbers. Some malicious software often uses specific ports for communication. For example, some Trojan programs often use high ports for data transmission.

[0050] 5) Protocol distribution feature: Count the proportions of different protocols (such as TCP, UDP, ICMP, etc.) in the network traffic. If it is found that the usage proportion of a certain protocol is abnormally high, there may be a situation where malicious traffic uses this protocol for attacks.

[0051] In the embodiments of the present invention, in S4, different types of feature nodes are distinguished and marked with different colors or shapes for visually observing and analyzing the graph model structure.

[0052] Specifically, taking various extracted features as nodes, a directed graph model is constructed according to the source IP address, destination IP address, and connection relationship in the network traffic data. For example, the source IP address node, destination IP address node, connection duration feature node, connection frequency feature node, etc. are connected by directed edges. The direction of the edge represents the flow direction of the network traffic, and the weight of the edge can be set according to relevant metrics such as the frequency of connection (the more connections, the greater the weight) or the traffic size (the larger the traffic, the greater the weight). Moreover, different types of feature nodes are distinguished and marked with different colors or shapes. For example, the source IP address node is represented by a circle, the destination IP address node is represented by a square, and the feature node is represented by a triangle, etc., so as to visually observe and analyze the graph model structure.

[0053] In the embodiments of the present invention, in S5, when using the graph embedding algorithm, the algorithm parameters are optimized and adjusted according to the characteristics of the network traffic data to improve the accuracy and efficiency of graph embedding.

[0054] Specifically, the constructed directed graph model is mapped to a low-dimensional vector space by using a graph embedding algorithm (such as the Node2vec algorithm). When using the Node2vec algorithm, the algorithm parameters (such as the walk step size, return probability, exploration probability, etc.) are optimized and adjusted according to the characteristics of the network traffic data. For example, for network traffic data with relatively frequent and complex connections, the walk step size is appropriately increased to better capture the relationships between nodes. Through the graph embedding algorithm, the low-dimensional vector representation of each node is obtained. These low-dimensional vectors can facilitate subsequent model training and analysis while retaining the graph structure information.

[0055] In the embodiments of the present invention, in S6, the cross-validation method is used to train and evaluate the anomaly detection model, and the model parameters and model structure with the optimal performance are selected.

[0056] Specifically, the low-dimensional vectors of the graph nodes corresponding to the known malicious traffic and normal traffic data are used as training data to train the anomaly detection model (such as a support vector machine model). The cross-validation method is used to train and evaluate the anomaly detection model. The training data is divided into a training set, a validation set, and a test set. During the training process, by adjusting the kernel function parameters (such as linear kernel, polynomial kernel, radial basis kernel, etc.) and regularization parameters of the support vector machine, the model parameters and model structure with the optimal performance are selected. For example, through multiple experiments, it is found that for the cross-domain malicious traffic detection scenario of the present invention, the radial basis kernel function can achieve better detection effects in most cases.

[0057] In an embodiment of the present invention, in S8, the detection result is stored in a database, and a detection report is generated regularly. The report content includes the type, source, time distribution of malicious traffic, and detection accuracy information.

[0058] Specifically, if the detected traffic is malicious, an alarm message is generated on the server side, and relevant traffic information (such as source IP address, destination IP address, malicious traffic characteristics, etc.) is recorded. At the same time, the detection result is stored in the database, and a detection report is generated regularly. The detection report content includes the type of malicious traffic (such as DDoS attack traffic, malware traffic, etc.), source (such as from a specific network domain or IP address segment), time distribution (such as in which time periods malicious traffic is more concentrated), and detection accuracy, etc. If it is normal traffic, it is allowed to be transmitted normally without additional processing.

[0059] In summary, by adopting a distributed acquisition architecture, network traffic data can be collected from multiple different network domains. And by constructing a detection method based on a graph model, various characteristics and connection relationships of cross-domain network traffic are comprehensively considered, effectively solving the problem of difficult detection of cross-domain malicious traffic and improving the detection ability for complex cross-domain network attacks. At the same time, by extracting multi-dimensional traffic behavior characteristics, mapping the graph model to a low-dimensional vector space using a graph embedding algorithm, and then combining with an anomaly detection model for training and detection, malicious traffic can be more accurately identified, reducing false positives and false negatives. Moreover, optimization strategies such as distributed acquisition and parameter optimization are adopted in each link of data acquisition, preprocessing, graph construction and embedding, and model training, improving the processing efficiency of the entire detection method and meeting the malicious traffic detection requirements in a network environment with high real-time requirements.

[0060] At the same time, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0061] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0062] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. Cross - domain malicious traffic detection method based on a graph model, characterized in that: The specific steps include: S1. Data collection: Collect network traffic data from multiple different network domains, the network traffic data including source IP address, destination IP address, port number, protocol type, traffic size and timestamp information. S2. Data preprocessing: Clean the collected network traffic data, remove invalid and duplicate data, standardize the data, and convert various types of data into a unified format for subsequent analysis. S3. Feature extraction: Extract features related to traffic behavior from the preprocessed network traffic data, including connection duration features, connection frequency features, traffic burst features, port usage features, and protocol distribution features. S4. Graph construction: Using the extracted features as nodes, a directed graph model is constructed based on the source IP address, destination IP address, and connection relationship in the network traffic data. The edges between nodes represent the network connection relationship, and the weight of the edge represents the frequency of the connection or a measure related to the traffic size. S5. Graph embedding: A graph embedding algorithm is used to map the constructed directed graph model to a low-dimensional vector space to obtain a low-dimensional vector representation of each node. The graph embedding algorithm includes a Node2vec algorithm or a GraphEmbedding algorithm. S6. Anomaly detection model training: Use low-dimensional vectors of graph nodes corresponding to known malicious traffic and normal traffic data as training data to train an anomaly detection model, wherein the anomaly detection model includes a support vector machine, a neural network or a decision tree model. S7. Detection: The low-dimensional vector of the node obtained after the network traffic data to be detected is input into the trained anomaly detection model after data collection, preprocessing, feature extraction, graph construction, and graph embedding steps to determine whether the network traffic is malicious traffic. S8. Result output: Output the detection result. If malicious traffic is detected, an alarm message is generated and relevant traffic information is recorded. If it is normal traffic, it is allowed to be transmitted normally.

2. The cross-domain malicious traffic detection method based on a graph model according to claim 1, characterized in that In S1, a distributed collection architecture is adopted, and collection probes are deployed at key network nodes in different network domains to realize the collection of large-scale network traffic data.

3. The cross-domain malicious traffic detection method based on a graph model according to claim 1, characterized in that In S2, invalid data and duplicate data are removed by setting data validity rules and data deduplication algorithms, wherein the data validity rules are set based on network protocol specifications and the rationality range of traffic data.

4. The cross-domain malicious traffic detection method based on a graph model according to claim 1, characterized in that In S3, the connection duration feature is obtained by calculating the average and standard deviation of the connection duration between the same source IP and destination IP; the connection frequency feature is determined by counting the number of connections between the source IP and the destination IP per unit time; the traffic burst feature is extracted by analyzing the rapid changes in traffic in a short period of time.

5. The cross-domain malicious traffic detection method based on a graph model according to claim 1, characterized in that In S4, different types of feature nodes are distinguished and marked with different colors or shapes to intuitively observe and analyze the graph model structure.

6. The cross-domain malicious traffic detection method based on a graph model according to claim 1, wherein In S5, when using the graph embedding algorithm, the algorithm parameters are optimized and adjusted according to the characteristics of the network traffic data to improve the accuracy and efficiency of the graph embedding.

7. The cross-domain malicious traffic detection method based on a graph model according to claim 1, characterized in that In S6, the cross-validation method is used to train and evaluate the anomaly detection model and select the model parameters and model structure with the best performance.

8. The cross-domain malicious traffic detection method based on a graph model according to claim 1, characterized in that In S8, the detection results are stored in the database, and a detection report is generated regularly. The report content includes the type, source, time distribution of malicious traffic, and detection accuracy information.