Link flooding attack detection method and system based on SDN service priority

By building a mapping relationship between services and links on the SDN platform, and combining machine learning models to detect and mitigate link flood attacks, the problems of blind target link selection, high threshold detection false alarm rate and poor mitigation effect in the existing methods are solved, and efficient malicious host identification and link congestion mitigation are achieved.

CN119995986AInactive Publication Date: 2025-05-13Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145196.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing link flood attack detection methods have problems such as blindness in target link selection, high false alarm rate of threshold detection, and poor mitigation effect.

Method used

The SDN-based service-first method is adopted to construct the mapping relationship between services and links, monitor the target link, and use machine learning models to perform feature extraction and malicious host detection on switch flow table information, and adopt a strategy of fast mitigation and slow detection.

Benefits of technology

It realizes rapid identification of malicious hosts, with high detection accuracy, recall rate and low false positive rate, and can effectively alleviate link congestion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995986A_ABST
    Figure CN119995986A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of network space, and relates to a link flooding attack detection method and system based on SDN (Software Defined Network) service priority. The method comprises the following steps: target link selection: acquiring actual routing paths from other switches in a network to a switch where a target service is located, and selecting a link with a high influence degree from the actual routing paths as a target link; monitoring a target link and judging whether the link is congested or not; when it is detected that the link is congested, the attack detection and mitigation module receives the congestion position information of the link sent by the SDN controller module, acquires flow table information forwarded to the congested link by a switch where the congested link is located, and integrates and extracts the acquired host behavior information as features of a machine learning model; and judging whether the host belongs to a malicious host or not, and then performing link congestion relief. According to the method, the malicious hosts can be quickly identified, and meanwhile, the method has relatively high detection accuracy and recall rate and relatively low false positive rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network space, and in particular relates to a link flooding attack detection method and system based on SDN service priority. Background Art

[0002] Unlike traditional distributed denial of service (DDoS) attacks, link flooding attacks (LFA) target links rather than services directly. The attacker congests the link to the target service, making it impossible for legitimate users to access the relevant services. Link flooding attacks have the following characteristics: using real and legitimate IP addresses, the traffic bandwidth of a single attacking host is small, and the traffic characteristics are highly similar to those of legitimate hosts, making it difficult for victims to detect. LFA has two forms: Coremelt attack (Studer A, Perrig A. The Coremelt Attack [C]. Computer Security-ESORICS 2009 (ESORICS), 2009, 5789: 37-52.) and Crossfire attack (Kang MS, Lee SB, Gligor V D. The Crossfire Attack [C]. 2013 IEEE Symposium on Security and Privacy, 2013: 127-141.). The Coremelt attack defines the target link, and the attacker uses a group of malicious hosts that send data to each other to congest the target link. Crossfire attack is a further development of Coremelt attack, which targets more complex network environments. In this attack, malicious hosts are used as the source and public services are used as the target. It is more common in network attacks. At present, research on LFA is focused on Crossfire attack. Crossfire attack mainly includes three stages: target network detection, target link selection, and rolling attack. In the target network detection stage, the attacker aims to obtain the network topology information of the area where the target service is located; in the target link selection stage, in order to achieve LFA, the attacker selects a batch of specific target links and further selects a batch of bait services that are close to these links, such as public service websites, mail servers, etc.; in the rolling attack stage, the attacker usually selects a group of non-intersecting target links and attacks periodically. The purpose is to avoid changes in the network topology caused by long-term network congestion, making it difficult for defenders to defend effectively.

[0003] Existing research on link flooding attack detection usually starts from the three stages of link flooding attack:

[0004] In terms of target link identification, CFADefense (Rafique W, He X, Liu Z, et al. CFADefense: A Security Solution to Detect and Mitigate Crossfire Attacks in Software-Defined IoT-Edge Infrastructure [C]. 2019 IEEE 21st International Conference on High Performance Computing and Communications (HPCC), 2019.) detects the flow table from the ingress node, analyzes the traffic data between all switch pairs, generates traffic weights for each link, and selects the link with a higher weight as the target link. Similarly, LFADefender (Wang J, Wen R, Li J, et al. Detecting and Mitigating Target Link-Flooding Attacks Using SDN [J]. IEEE Transactions on Dependable and Secure Computing, 2019, 16 (6): 944-956.) analyzes the flow table obtained from the switch, counts the traffic passing through each link, and selects the link with high traffic density as the target link, but in large networks, this will bring huge computational overhead. We believe that it is blind to simply rely on the traffic attributes or graph attributes of the link to determine the target link. These methods have limitations, namely, the selected links may not necessarily pose a high threat to specific services, and even if link congestion occurs, it may not necessarily have a significant impact on certain important services.

[0005] In attack detection, some methods use threshold judgments during congestion periods, such as detecting the occurrence of link flooding attacks by changes in the destination IP entropy value. However, relying solely on threshold detection will cause a high false alarm rate and make it difficult to detect malicious hosts. In the use of machine learning to detect link flooding attacks, existing methods such as CyberPulse++ (Rasool R, Ahmed K, Anwar Z, et al. CyberPulse++: A Machine Learning Based Security Framework for Detecting Link Flooding Attacks in Software Defined Networks [J]. International Journal of Intelligent Systems, 2021.) still rely on the features required for traditional DDoS attack detection. However, due to the similarity between LFA single attack traffic and normal user traffic and the small attack bandwidth, these methods are difficult to apply when defending against LFA.

[0006] In terms of attack mitigation, some methods such as CFADefense (Rafique W, He X, Liu Z, et al. CFADefense: A Security Solution to Detect and Mitigate Crossfire Attacks in Software-Defined IoT-Edge Infrastructure [C]. 2019 IEEE 21st International Conference on High Performance Computing and Communications (HPCC), 2019.) consider using a strategy of mitigation first and then detection, using the rerouting method to alleviate link congestion, and then using multiple rounds of interaction to identify malicious hosts. The disadvantage of this method is that it is difficult to achieve the expected mitigation effect when the network lacks sufficient alternative links. In addition, the detection of malicious hosts requires multiple rounds of interaction. Summary of the invention

[0007] Aiming at the problems in the existing methods that the target link selection is blind, the simple reliance on threshold detection will cause a high false alarm rate, and it is difficult to achieve the expected mitigation effect, the present invention proposes a link flooding attack detection and mitigation method (SPM) based on service priority on software defined network (SDN). The experimental results based on the simulation platform show that compared with the traditional feature method, the method of the present invention can quickly identify malicious hosts, and at the same time has higher detection accuracy, recall rate and lower false positive rate.

[0008] The present invention specifically adopts the following technical solutions:

[0009] The present invention focuses on the target service and monitors the target link by building a mapping relationship between the service and the link. The method deeply analyzes the characteristics of LFA attack behavior, extracts features from the switch flow table information during the congestion period, and uses a machine learning model to judge the features. For the detected malicious host, the method of the present invention adopts a fast mitigation and slow detection defense strategy, which effectively reduces the impact on the legitimate host.

[0010] The present invention provides a link flooding attack detection method based on SDN service priority, comprising:

[0011] Step 1: target link selection: obtain the actual routing path from other switches in the network to the switch where the target service is located, and select the link with high influence, that is, the target link; among them, according to the Pareto principle, the links with the top 20% of the influence ranking are selected as the high-influence links; in addition, the target service is selected by input in the present invention.

[0012] Step 2: Monitor the target link and determine whether the link is congested.

[0013] The present invention uses link delay as an indicator to determine whether a link is congested. The link delay is calculated by formula (1), where T1 and T2 are the delays for two adjacent switches to receive LLDP data packets and report them to the SDN controller, and T A , T B The round-trip delay from the SDN controller to the two switches is:

[0014] One-way link delay: link_delay = (T1 + T2 - T A -T B ) / twenty one)

[0015] Step 3: When link congestion is detected, the attack detection and mitigation module receives the link congestion location information sent by the SDN controller module, obtains the flow table information forwarded by the switch where the congested link is located to the congested link, and integrates the obtained host behavior information as the feature of the machine learning model for extraction.

[0016] The features include: number of destination IPs, number of high-frequency destination IPs, proportion of high-frequency destination IPs, number of bytes at the current moment, number of packets at the current moment, number of bytes per packet, total number of bytes, total number of packets, proportion of bytes at the current moment, proportion of packets at the current moment, and duration. The machine learning model is a random forest algorithm. According to the Pareto principle, the top 20% of the traffic reaching different destination IPs is selected as the high-frequency destination IP.

[0017] Step 4: Determine whether the host is a malicious host, and then alleviate link congestion.

[0018] Specifically, the confidence threshold and identification threshold are set. When the probability that the host is identified as a malicious host by the machine learning model is greater than the given confidence threshold, the sample is identified as a suspicious host and the host traffic is temporarily blocked. When the number of times the host is identified as a suspicious host is greater than the given identification threshold, the sample is defined as a malicious host and is directly blocked.

[0019] The present invention provides a link flooding attack detection system based on SDN service priority, which comprises an SDN controller module and an attack detection and mitigation module.

[0020] The SDN controller module is used for target link selection and monitoring, obtaining target link status information, and sending link congestion location information to the attack detection and mitigation module; the SDN controller implements target link selection and link delay calculation.

[0021] The attack detection and mitigation module is used to receive the congestion location information of the link, request the flow table information from the switch at the congestion location, extract attack features and detect malicious hosts, and then issue mitigation strategies to alleviate the congested link. After receiving the congestion location information, the attack detection and mitigation module directly requests the flow table information from the switch at the congestion location via the southbound interface; for the detected malicious host, the attack detection and mitigation module issues a discard action flow table via the southbound interface to alleviate link congestion.

[0022] The beneficial effects of the present invention are:

[0023] The nature of the LFA mitigation method based on software defined network (SDN) service priority of the present invention determines that it cannot solve the problem locally or on a single link, and it must be detected and defended in the entire network. SDN has good global visibility, and its characteristics such as separation of forwarding and control, customizable control services, and complete north-south interfaces provide defenders with a more convenient way to implement strategies. Service priority means not providing a unified and equal defense mechanism for the network, but focusing on several key services in the network, shifting the defense target from specific links in the network to specific services. The advantage of doing this is that the correspondence between links and services can be clarified, effectively reducing the cost of detection and defense.

[0024] The present invention analyzes the characteristics of LFA and extracts behavioral features from the SDN switch flow table, overcoming the limitation that the current LFA still uses the features required for traditional DDoS detection. By obtaining link statistics in real time during congestion, congestion can be alleviated in a timely and rapid manner.

[0025] The present invention is tested on a variety of machine learning models and found that the random forest (RF) algorithm is more effective. For suspicious hosts detected by the model, the present invention adopts a defense strategy of fast mitigation and slow identification. Only when the host is identified as suspicious multiple times is it determined to be a malicious host, otherwise the discard time is only gradually increased. This can not only quickly alleviate link congestion, but also avoid additional costs such as rerouting. Compared with methods that require multiple rounds of interaction, the method of the present invention can quickly detect malicious hosts, and at the same time has a higher recall rate and a lower false positive rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a link flooding attack detection system based on SDN service priority.

[0027] Figure 2 This is the impact of attack bandwidth ratio on round-trip delay under different bandwidths.

[0028] Figure 3 It is the Abilence topology.

[0029] Figure 4 Comparison of precision, recall and false positive rate of different models; (a) the method of the present invention; (b) CyberPulse++; (c) CFADefense.

[0030] Figure 5 Comparison of recall rate and false positive rate under different total amount of malicious hosts; (a) recall rate; (b) false positive rate. In the figure, SPM represents the method of the present invention.

[0031] Figure 6 Comparison of recall rate and false positive rate under different attack cycles; (a) recall rate; (b) false positive rate. In the figure, SPM represents the method of the present invention. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] Figure 1 A link flooding attack detection system based on SDN service priority is demonstrated, which includes an SDN controller module and an attack detection and mitigation module.

[0034] The SDN controller module is used to select and monitor the target link, obtain the target link status information, and send the link congestion location information to the attack detection and mitigation module; specifically, the target link selection and link delay calculation are implemented in the SDN controller, and the SDN controller periodically obtains the link status information to analyze whether a specific link is congested. If congestion occurs, the congestion location information is sent to the attack detection and mitigation module.

[0035] The attack detection and mitigation module is used to receive the congestion location information of the link, request the flow table information from the switch at the congested location, extract attack features and detect malicious hosts, and then issue mitigation strategies to alleviate the congested link. Specifically, after receiving the congestion location information, the attack detection and mitigation module directly requests the flow table information from the switch at the congested location via the southbound interface, further extracts the attack features, and judges the host through the machine learning model. For the detected malicious host, the discard action flow table is issued through the southbound interface to alleviate link congestion. This avoids the SDN controller from bearing a large load, and obtaining attack features from the flow table data also makes attack detection more lightweight and efficient.

[0036] Table 1 Definition and explanation of symbols

[0037]

[0038] The link flooding attack detection method based on SDN service priority includes:

[0039] Step 1: Target link selection: Obtain the actual routing path from other switches in the network to the switch where the target service is located, and select the link with high influence, which is the target link;

[0040] We believe that it is blind and limited to simply rely on the traffic attributes or graph attributes of the link to determine the target link. That is, the selected link may not be highly threatening to a specific service. Even if the link is congested, it may not have a significant impact on some important services. By inputting the selected target service and clarifying the relationship between the target service and the link, we can clearly understand the impact of a specific link on the target service during congestion, thereby making the detection and defense of attacks more efficient. Here, we assume that all legitimate users are evenly distributed in geographic space, that is, the legitimate users connected to each switch have the same weight. Therefore, our approach is to obtain the actual routing path from other switches in the network to the switch where the target service is located. These paths will overlap on some links. By referring to the Pareto principle (Harvey HB, Sotardi ST. The Pareto Principle [J]. Journal of the American College of Radiology, 2018), we select links with high influence (i.e., links with the top 20% of influence ranking as target links). There are many ways to query the routing path, such as obtaining it from the routing policy of the controller or measuring it hop by hop on the switch, which will not be repeated here. The algorithm for obtaining the target link is shown in Algorithm 1:

[0041]

[0042] Step 2: Monitor the target link and determine whether the link is congested.

[0043] In step 1, we obtain a set of links that have a high impact on the target service, and monitor these links to determine whether LFA occurs. Common indicators for determining whether a link is congested include link delay, remaining bandwidth, packet loss rate, etc. However, in terms of service quality, delay is the largest perceived influencing factor. Therefore, the present invention uses link delay as an indicator for determining whether a link is congested. We simulated link attacks. To simulate the actual link situation, we added a delay of 3ms to each link. Figure 2 As shown in the figure, we give the link delay results for different attack bandwidth ratios under different link bandwidths. In the case of link delay, when the attack bandwidth is less than a specific threshold, the link delay is basically stable; when the attack bandwidth exceeds the threshold, the link delay will increase significantly. As the link bandwidth increases, the turning point of the delay change gradually advances. Therefore, it is not feasible to select a unified threshold for the link, and different choices should be made under different bandwidth conditions. The calculation of the link delay can be obtained by formula (1), where T1 and T2 are the delays for two adjacent switches to receive LLDP data packets and report them to the controller, and T A 、T BThe round-trip delay from the controller to the two switches is:

[0044] One-way link delay: link_delay = (T1 + T2 - T A -T B ) / twenty one)

[0045] We let the switches in the network send LLDP packets periodically, and capture and process these packets in the controller, integrating the calculation of link delay into the SDN controller. Here, we set the link congestion status to be detected every 5 seconds. When link congestion is detected, the controller requests the attack detection and mitigation module to obtain the flow table information forwarded by the switch where the congested link is located to the congested link. When the detection program receives the request, it enables a new thread, requests to obtain the flow table information once per second, and ends the thread after 5 seconds. Obtaining the switch flow table information is implemented through the southbound interface. For example, in the Ryu controller, the switch flow table can be issued, obtained, modified, and other operations by enabling the ryu.app.ofctl_rest.py program.

[0046] Step 3: When link congestion is detected, the attack detection and mitigation module receives the link congestion location information sent by the SDN controller module, obtains the flow table information forwarded by the switch where the congested link is located to the congested link, and integrates the obtained host behavior information as the feature of the machine learning model for extraction.

[0047] The flow table information obtained usually includes the following: priority, timeout, matching domain, number of packets, number of bytes, action, etc. Among these flow table information, the most important is the matching domain. A source IP may have multiple flow tables that reach different destinations. We integrate the information belonging to the same source IP to obtain the number of bytes, number of packets, destination IP list and other information sent by the source IP. In addition, the size of all traffic reaching different destination IPs is counted, and the destination IPs with the top 20% of the traffic are selected as high-frequency destination IPs according to the Pareto principle. We consider the behavioral characteristics of hosts during congestion within a certain historical time window. For source IPs that appear for the first time or that have exceeded the historical window, they are only updated in the historical data; otherwise, we will further process the data and eventually integrate the obtained host behavior information as a judgment feature used by the machine learning model.

[0048] The characteristic information obtained by the final processing is shown in Table 2. The reason for counting the number of destination IPs and the number of high-frequency destination IPs is that in the historical window, the number of destination IPs corresponding to the legitimate host is usually small, and LFA will frequently change the target link to achieve rolling attacks. However, at a certain congestion moment, the destination IPs of these attacks will still be concentrated on a small number of high-frequency destination IPs. Congestion time refers to the time when the source IP appears in the congestion period, and the summary is the sum of the time when the source IP appears on different congested links. This is based on the assumption that the longer the appearance time in the congestion period, the greater the possibility of being a malicious host. Although a single attacking host and a legitimate host have the same traffic characteristics, it is feasible to give priority to suppressing traffic with high bandwidth usage during congestion, so we also consider information such as the number of bytes and the number of packets of the host. The reason for considering the current number of packets / bytes as a percentage of the total number of packets / bytes in the time window is that the attack mode of the attacking host is single, while the traffic mode of the legitimate host is more complex and changeable. Similarly, the reason for considering the number of bytes per packet is that the access behavior of the legitimate host is bursty, and in order to achieve a saturation attack effect, it is difficult for the attacking host under the unified control of the attacker to imitate this behavior.

[0049] Table 2 Extracted features and definitions

[0050]

[0051]

[0052] After data preprocessing, we believe that the choice of machine learning model should be robust. Given the small number of features we have, we select algorithm models with high precision, recall, and low false positive rate from the following classic machine learning methods.

[0053] (1) Logistic Regression (LR)

[0054] Logistic regression is a classification algorithm that is mainly used to solve binary classification problems and can also be extended to multi-classification. Its essence is to map the output of linear regression to a probability interval through a nonlinear activation function (usually a Sigmoid function) based on linear regression to predict the probability of an event occurring or not occurring.

[0055] (2) Random Forest (RF)

[0056] Random forest is a powerful ensemble learning algorithm. Its principle is to build multiple decision trees and combine the results of these decision trees to make predictions. In classification tasks, random forest determines the final classification result by voting on multiple decision trees. This method makes random forest have high accuracy and stability and can effectively avoid overfitting problems.

[0057] (3) Support Vector Machine (SVM)

[0058] Support vector machine is an efficient supervised learning algorithm. Its core idea is to find an optimal hyperplane to separate data points of different categories as much as possible. For linearly separable data, this hyperplane can directly separate the two types of data and maximize the distance between the two types of data and the hyperplane. For nonlinearly separable data, SVM introduces kernel functions to map the data to high-dimensional space, making the data in high-dimensional space linearly separable.

[0059] (4) Naive Bayes (NB)

[0060] Naive Bayes is a classification method based on probability statistics. Its basic principle is to assume that features are independent of each other, and then calculate the conditional probability of each feature under each category based on the training data. When classifying, the Bayesian principle is used to calculate the posterior probability that a given sample belongs to each category, and finally the category with the highest probability is used as the prediction result.

[0061] (5) Decision Tree (DT)

[0062] A decision tree is an intuitive classification and regression model that gradually divides data into different subsets by making a series of conditional judgments on features, eventually forming a tree structure. Starting from the root node, data is divided into different child nodes according to the different values ​​of a certain feature, and each child node is further divided according to another feature until a leaf node is reached. The leaf node represents the final classification or prediction result.

[0063] (6) Gradient Boosting Tree (GBT)

[0064] Gradient boosting tree is an ensemble learning algorithm based on decision trees. Its essence is to gradually reduce the value of the loss function by iteratively training decision trees. In each iteration, the model will fit a new decision tree based on the negative gradient value of the loss function of the current model, and then add this new decision tree to the existing model to update the model's prediction results. By continuously adding decision trees, the model can continuously correct previous errors, thereby improving prediction accuracy.

[0065] The above algorithm models can give the probability that the host is a malicious host at the time of congestion. The specific model comparison results will be described in the experimental part. At the time of congestion, the features corresponding to each source IP will be judged by the model. If the host is judged as a suspicious host, we will further process the host.

[0066] Step 4: Determine whether the host is a malicious host, and then alleviate link congestion.

[0067] Specifically, the confidence threshold and identification threshold are set. When the probability that the host is identified as a malicious host by the machine learning model is greater than the given confidence threshold, the sample is identified as a suspicious host and the host traffic is temporarily blocked. When the number of times the host is identified as a suspicious host is greater than the given identification threshold, the sample is defined as a malicious host and is directly blocked.

[0068] In the early stage of the attack, due to the lack of historical data accumulation, it is easy to misjudge legitimate hosts. Therefore, we use confidence thresholds and identification thresholds, and further adopt a defense strategy of fast mitigation and slow detection to reduce the impact on legitimate hosts while timely alleviating link congestion.

[0069] The specific link congestion relief algorithm is shown in Algorithm 2. When a host is identified as a suspicious host, we do not block it completely directly. Instead, we first query the number of times the source IP is identified as a suspicious host and generate a discard time based on the historical number of times. Through the southbound interface, a high-priority DROP action flow table is directly sent to the source switch to temporarily block the source IP traffic. The discard time gradually increases with the number of times the source IP is identified as malicious. When the number of identifications is greater than the identification threshold, we identify it as a malicious host and block it directly. Although this method may cause some legitimate hosts to be misjudged, the impact on legitimate hosts is limited and it is acceptable for alleviating global congestion.

[0070]

[0071] Table 3 Evaluation experimental parameters

[0072]

[0073] To evaluate the effectiveness of the method of the present invention, we built the Abilence topology of the American education backbone network in the Mininet simulation environment (Laszka A, Gueye A. Network topology vulnerability / cost trade-off: Model, application, and computational complexity [J]. Internet Mathematics, 2015, 11 (6): 588-626.). Figure 3As shown in the figure, the network contains 11 nodes and 15 links. The link bandwidth is set to 20Mbps, the one-way link propagation delay is 3ms, and the one-way propagation delay of 100ms is used as the basis for judging link congestion. At the same time, a customized controller and attack detection and mitigation program are implemented on the Ryu controller. Due to the lack of standard data sets for LFA, we refer to the literature (Kang MS, Lee SB, Gligor V D. The Crossfire Attack [C]. 2013 IEEE Symposium on Security and Privacy, 2013: 127-141.) and give the following attack data set construction process. Table 3 gives the specific parameters used in the attack simulation. All experiments were carried out on the Ubuntu 20.04.6 LTS 64-bit operating system. Core TM The test was performed under the environment of i9-9980XE CPU@3.00GHz processor and 64G memory.

[0074] like Figure 3 As shown in the topology, we connect three hosts under each switch. Among them, M is used to simulate malicious hosts, and N is used to simulate legitimate hosts. On these hosts, many attacks and legitimate behaviors are simulated by enabling multithreading. S host enables a simple http service as bait, and selects S9 as the target service. The actual routing path from the other switches to switch v9 has been marked in the figure. According to the Pareto principle, the selected target links include L1, L2, and L3. To simulate normal user behavior, N periodically randomly selects 10 legitimate IPs from the source IP pool of size 50 to randomly access the bait service. The bandwidth occupied by a single legitimate IP traffic is about 0.15Mbps. Table 4 gives the specific attack information. To achieve a saturated attack, each attack source M needs to select 30 malicious IPs at the same time, that is, each round of attack simulates 90 malicious host behaviors. At the end of each round of attack cycle, the attacker changes the target link, uses the corresponding attack source to randomly select a new malicious IP, and re-initiates the attack. After simulating the legitimate host for 5 minutes, LFA is implemented, and the attack lasts for a total of 10 minutes.

[0075] Table 4 Attack information selection

[0076]

[0077]

[0078] We divided the data set into two parts with a ratio of 0.7 and 0.3, and performed performance tests on the selected machine learning models. The selected evaluations include accuracy, precision, recall, and false positive rate. Their calculation methods are shown in the following formulas (2) to (5), where TP represents the number of samples that are actually attack hosts but detected as attack hosts, FN represents the number of samples that are actually attack hosts but detected as legitimate hosts, FP represents the number of samples that are detected as attack hosts but are actually legitimate hosts, and TN represents the number of samples that are actually legitimate hosts and detected as legitimate hosts. Table 5 shows the evaluation results under several pre-selected machine learning classifiers. The test results show that the random forest method achieves the best results among several comparison methods.

[0079] Accuracy:

[0080] Accuracy:

[0081] Recall:

[0082] False Positive Rate:

[0083] Table 5 Evaluation results of machine learning classifiers

[0084]

[0085] The data given in Table 5 are the results before threshold discrimination. In order to further improve the performance of the method of the present invention and reduce the false positive rate, the probability results output by the model are first compared with the confidence threshold. When the probability that the model output sample belongs to a malicious host is greater than the given confidence threshold, the sample is identified as a suspicious host, and the suspicious host is only temporarily discarded. When the number of times a sample is identified as a suspicious host is greater than the given recognition threshold, the sample is defined as a malicious host. At this time, we are more concerned about the recall rate and false positive rate after threshold discrimination, that is, the detection rate of malicious hosts and the false alarm rate of legitimate hosts. In order to find the optimal confidence threshold and recognition threshold, as shown in Tables 6 and 7, we give the experimental results under different threshold conditions. The experimental results show that when the confidence threshold is less than 0.95, the method has a higher false positive rate, and when the confidence threshold is greater than 0.95, the recall rate decreases significantly. Therefore, we choose the confidence threshold to be 0.95. Similarly, we choose the recognition threshold to be 5.

[0086] Table 6 Recall rate and false positive rate under different confidence thresholds

[0087]

[0088]

[0089] Table 7 Recall rate and false positive rate under different recognition thresholds

[0090]

[0091] To verify the effectiveness of the method of the present invention, we choose to compare it with the CyberPulse++ method (Rasool R, Ahmed K, Anwar Z, et al. CyberPulse++: A Machine Learning Based Security Framework for Detecting Link Flooding Attacks in Software Defined Networks [J]. International Journal of Intelligent Systems, 2021.) which uses traditional DDoS detection features and the CFADefense method (Rafique W, He X, Liu Z, et al. CFADefense: A Security Solution to Detect and Mitigate Crossfire Attacks in Software-Defined IoT-Edge Infrastructure [C]. 2019 IEEE 21st International Conference on High Performance Computing and Communications (HPCC), 2019.) which requires multiple rounds of interaction. Our method obtains feature information from the switch flow table during congestion and makes judgments on the random forest model. The CyberPulse++ method uses traditional 15-dimensional features when detecting LFA. These features include used bandwidth, packet loss rate, total bandwidth, packet size, etc. Its authors achieved the best value on 10 machine models. Similarly, we reproduced these 10 models, extracted 15-dimensional features after attack through traffic capture, and selected the best value among the 10 model results. The CFADefense method identifies the high-frequency host intersection between multiple rounds of attacks as malicious hosts when mitigating LFA. We statistically analyze the detection results of this method after each round of attack. We will compare the performance of the three methods from the following aspects: (1) real-time detection capability comparison; (2) the impact of the total number of malicious hosts on the performance of the methods; and (3) the impact of different attack cycles on the performance of the methods.

[0092] First, we compare the real-time precision, recall, and false positive rate of the three methods. For the method of the present invention and the CyberPulse++ method, the confidence threshold and recognition threshold are applied; for the CFADefense method, we give the results after each round of attack. Figure 4 As shown in the figure, due to the use of thresholds, both the method of the present invention and the CyberPulse++ method have high precision, but the recall rate of our method increases faster and can detect malicious hosts faster. When the attack is carried out for 5 minutes, the recall rate of the detection results of the method of the present invention can reach 94.68%, while the recall rate of the comparison method is 88.35%. In the 10th minute of the attack, the recall rate of the traditional feature method is 94.58%, and the false positive rate is 3.40%, while the method of the present invention has achieved a recall rate of 97.52%, and the false positive rate is only 1.36%. The results of the CFADefense method fluctuate greatly in the early stage of the attack. After the 25th round of attack, the precision rate gradually increases, and the false positive rate gradually decreases. The final precision rate is 91.56%, the recall rate is 46.57%, and the false positive rate is 12.91%. The experimental results show that the method of the present invention can detect malicious hosts faster and has a lower false positive rate.

[0093] In addition, we tested the performance of several methods under different attack parameters. Because the method of the present invention has a high accuracy rate, only the differences in recall rate and false positive rate of several methods are considered here. First, we consider the impact of the total number of attack host source IPs on the method. The parameter here is the ratio of the total number of malicious host IPs to the number of malicious IPs launching attacks at the same time. For example, when the ratio is 5, a single host M enables 30 malicious source IPs at the same time. The experiment uses 8 hosts M in total, and the total number of malicious source IPs is 1200. Figure 5 As shown, as the ratio increases, the recall rate of the proposed method decreases less than that of the CyberPulse++ method, while the CFADefense method performs better only when the ratio is small.

[0094] Finally, if Figure 6 As shown, we also compared the results of different methods under different attack cycles, where the parameter is the period of each attack round. As the attack period increases, the method of the present invention always outperforms the CyberPulse++ method in terms of recall rate and false positive rate. Under different attack periods, the detection rate of malicious hosts is always greater than 90%, and the false positive rate is less than 1.5%; while the CFADefense method requires longer attack rounds to effectively capture attacks. When the attack duration is constant, the false positive rate increases significantly with the increase of the attack period. Comprehensive experimental results show that the method of the present invention can detect malicious hosts faster and has better applicability to different attack environments.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A link flooding attack detection method based on SDN service priority, characterized in that: include: Step 1: Target link selection: Obtain the actual routing path from other switches in the network to the switch where the target service is located, and select the link with high influence, which is the target link; Step 2: Monitor the target link and determine whether the link is congested; Step 3: When link congestion is detected, the attack detection and mitigation module receives the link congestion location information sent by the SDN controller module, obtains the flow table information forwarded by the switch where the congested link is located to the congested link, and integrates the obtained host behavior information as the feature of the machine learning model for extraction; Step 4: Determine whether the host is a malicious host, and then alleviate link congestion.

2. The method according to claim 1, characterized in that In step 1, the links with the top 20% of influence ranking are selected as high-influence links according to the Pareto principle.

3. The method according to claim 1, characterized in that In step 2, the indicator for determining whether the link is congested includes the link delay.

4. The method according to claim 2, characterized in that: In step 2, the link delay is calculated using formula (1), where T1 and T2 are the delays for two adjacent switches to receive the LLDP data packet and report it to the SDN controller, and T A , T B The round-trip delay from the SDN controller to the two switches is: One-way link delay: link_delay = (T1 + T2 - T A -T B ) / twenty one).

5. The method according to claim 1, characterized in that In step 3, the characteristics include: number of destination IPs, number of high-frequency destination IPs, proportion of high-frequency destination IPs, number of bytes at the current moment, number of packets at the current moment, number of bytes per packet, total number of bytes, total number of packets, proportion of bytes at the current moment, proportion of packets at the current moment, and duration; among them, according to the Pareto principle, the top 20% of the traffic reaching different destination IPs are selected as high-frequency destination IPs.

6. The method according to claim 1, characterized in that In step 3, the machine learning model is a random forest algorithm.

7. The method according to claim 1, characterized in that In step 4, the confidence threshold and identification threshold are first set. When the probability that the host is identified as a malicious host by the machine learning model is greater than the given confidence threshold, the sample is identified as a suspicious host and the host traffic is temporarily blocked; when the number of times the host is identified as a suspicious host is greater than the given identification threshold, the sample is defined as a malicious host and is directly blocked.

8. A link flooding attack detection system based on SDN service priority, characterized in that: Includes SDN controller module and attack detection and mitigation module; The SDN controller module is used to select and monitor the target link, obtain the target link status information, and send the link congestion location information to the attack detection and mitigation module; The attack detection and mitigation module is used to receive the congestion location information of the link, request the flow table information from the switch at the congestion location, extract attack features and detect malicious hosts, and then issue mitigation strategies to alleviate the congested link.

9. The system according to claim 8, characterized in that After receiving the congestion location information, the attack detection and mitigation module directly requests flow table information from the switch at the congestion location via the southbound interface; for the detected malicious host, the attack detection and mitigation module sends a discard action flow table through the southbound interface to alleviate link congestion.