Link Flooding Attack Detection and Defense Method and System Based on Reputation and Clustering Algorithm

By constructing an IP reputation model and a two-stage clustering algorithm, detecting link flood attacks, the problem of low detection accuracy in the existing technology is solved, and efficient identification of zombie host IP under long-term strategies is achieved.

CN120090878BActive Publication Date: 2025-07-04GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510571041.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-04
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing link flood attack detection and defense methods have the problem of low detection accuracy, especially when attackers adopt long-term mapping strategies, the detection success rate of traditional methods has decreased, and the existing models rely on a large amount of data to train, and there is a deviation between the simulation environment and the real network environment.

Method used

Build an IP reputation model, update the reputation score value by detecting the IP routing tracking behavior, combine the IP entropy of the access target within the untrusted IP set, and use a two-stage clustering algorithm for detection: the first stage uses self-organized graph clustering to eliminate benign traffic, and the second stage uses Gaussian hybrid clustering to improve accuracy, and block malicious IPs through flow table rules.

Benefits of technology

It improves the detection accuracy under long-term mapping strategies, has better robustness, can effectively identify zombie host IP, solves the problem of lack of link flooding attack traffic data, and achieves efficient detection and defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090878B_ABST
    Figure CN120090878B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for detecting and defending link flooding attacks based on reputation and clustering algorithms. The method includes: constructing an IP reputation model, detecting the route tracing behavior newly added by an IP, and updating the reputation score value of the newly added route tracing behavior according to the IP reputation model; detecting the access destination IP entropy of the IPs within the untrusted IP set to determine the occurrence of an LFA attack and locate the victim link; sampling the traffic of the switch at the entrance end of the victim link and extracting flow features; training a first-stage self-organizing map clustering model to perform a preliminary detection on the sampled flows; training a second-stage Gaussian mixture clustering model to perform a secondary detection on the sampled flows; and issuing flow table rules to block the traffic of the malicious IPs identified by clustering. The present application can improve the detection accuracy and has better robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security protection technologies, and in particular, to a method and system for detecting and defending against link flooding attacks based on reputation and clustering algorithms. Background Art

[0002] With the development of the Internet, DDoS attacks have become increasingly rampant, with the scale of attacks continuously increasing and various new types of DDoS attacks emerging in an endless stream. Among them, a new type of DDoS attack called Link flooding attack has been discovered by the academic community and has occurred in the real network environment. The link flooding DDoS attack incident suffered by NetEase in 2015 is the most typical.

[0003] Different from traditional DDoS attack types, the Link flooding attack (LFA) targets the backbone link of the target network domain. The attacker uses a large number of zombie hosts and public servers or decoy servers in the target network domain to make the zombie hosts communicate with the decoy servers at a low-speed normal traffic. By continuously superimposing these low-speed traffic, the target link bandwidth is exhausted. Since accessing the target area must pass through the attacked backbone link, normal users cannot access the target network domain, so the attacker achieves the purpose of denying service to the target.

[0004] The detection and defense of link flooding attacks have many difficulties. First, the attacker's target is the key link of the target network domain, rather than the target host server. This makes the traditional intrusion detection system (IDS) and intrusion prevention (IDP) means using gateways ineffective against link flooding attacks (LFA) because the attack traffic does not reach the target host or gateway at all. The attack traffic leads to decoy servers, which are public servers or servers without any gateway firewalls and traffic detection and filtering means. Second, the zombie hosts used in the attack send low-speed normal network traffic, such as website browsing (HTTP) or TCP communication traffic. This makes it more difficult for traditional deep packet inspection (DPI) traffic filtering means to take effect. In addition, the attacker will also change the pattern of attack traffic and the attack link to avoid detection, making the detection and defense of link flooding attacks face more severe challenges.

[0005] To detect and defend against link flooding attacks, many detection and defense solutions have been proposed in the academic community. Since an attacker initiating a link flooding attack must first conduct network mapping through the traceroute tool to learn the topology of the target network domain and identify available public servers or decoy machines in the target network domain, some existing work has proposed methods for detecting link flooding attacks and discovering attacked links by monitoring the traceroute traffic behavior. One method is to identify the occurrence of link flooding attacks by monitoring the entropy anomaly of traceroute packets. Another method is to count the number of traceroute packets within a time window and determine the occurrence of attacks and identify bot hosts used to initiate attacks by detecting outlier windows. Both methods can only judge the abnormality of traceroute behavior within a short period. When an attacker adopts a long-term mapping strategy, that is, the attacker controls bot hosts to conduct network topology mapping at a traceroute mapping frequency close to that of normal users, the detection success rate of the short-term instantaneous anomaly detection method will decrease.

[0006] The (mitigation) defense means against link flooding attacks can be mainly divided into the following ideas and means. One is the rerouting technology based on SDN technology, which diverts the traffic of the victim link to a backup link or other non-critical links to share the traffic pressure on the victim target link. However, this means still retains malicious traffic, and the malicious traffic will still cause pressure on the network. Another is to use the means of traffic cleaning. However, since the traffic of link flooding attacks is very similar to normal traffic, the false alarm rate is very high when using deep packet inspection (DPI) to identify traffic characteristics. Therefore, there is not much related work. Moreover, since link flooding attacks rarely occur in the actual network environment and there are very few related events, there is currently no dataset for link flooding attacks (LFA). The method adopted by related work is to use the benign traffic of the botnet detection dataset and the malicious traffic generated in the simulation environment to train the model. Most of the existing work uses supervised models, such as convolutional neural networks (CovNet), long short-term memory models (LSTM), etc. These models rely on a large amount of data for training, and there is always a deviation between the traffic generated in the simulation environment and the traffic in the actual network environment. It is difficult to determine the true detection accuracy of the model. Summary of the Invention

[0007] The purpose of this application is to provide a method and system for detecting and defending link flooding attacks based on reputation and clustering algorithms, aiming to solve the problem of low detection accuracy in traditional detection and defense methods for link flooding attacks.

[0008] In a first aspect, this application provides a method for detecting and defending link flooding attacks based on reputation and clustering algorithms. The method includes:

[0009] Build an IP reputation model, detect the route tracing behavior newly added by the IP, and update the reputation score value of the newly added route tracing behavior according to the IP reputation model;

[0010] Detect the access destination IP entropy of the IPs in the untrusted IP set to judge the occurrence of the LFA attack and locate the victim link;

[0011] Sample the traffic of the switch at the entrance end of the victim link and extract the flow characteristics;

[0012] Train the first-stage self-organizing map clustering model to preliminarily detect the sampled flow;

[0013] Train the second-stage Gaussian mixture clustering model to perform secondary detection on the sampled flow;

[0014] Issue flow table rules to block the traffic of the malicious IPs identified by clustering.

[0015] Furthermore, the steps of building the IP reputation model include:

[0016] Build an IP reputation model according to the following formula:

[0017] ,

[0018] where, represents the density function, x represents the reputation score value, μ represents the average reputation score, and σ represents the variance of the reputation score;

[0019] Calculate the average reputation score according to the following formula:

[0020] ,

[0021] Calculate the variance of the reputation score according to the following formula:

[0022] ,

[0023] where, represents the reputation score record table, represents the number of elements in the reputation score record table, represents the reputation score of the IP.

[0024] Furthermore, the steps of detecting the route tracing behavior newly added by the IP and updating the reputation score value of the newly added route tracing behavior according to the IP reputation model include:

[0025] Collect the traffic data packets of the route tracing behavior of the IP, and judge whether the source address IP in the traffic data packets of the route tracing behavior is recorded in the IP reputation table. If not, record it as , if it has been recorded, it is recorded as ;

[0026] If the IP is , its reputation value is obtained according to the following formula:

[0027] ,

[0028] The average reputation score of all IPs is obtained according to the following formula:

[0029] ,

[0030] where represents the reputation value of the IP, represents the reputation score obtained by the IP from the external threat intelligence library, is the weight of the external reputation score in the initial reputation, is the average reputation score of all IPs, is the sum of the reputation scores of all IPs;

[0031] If the IP is , then is added to the route trace observation event set , and it is judged according to the following formula whether it is trustworthy:

[0032] ,

[0033] If is a trustworthy IP, its observation probability is obtained according to the following formula:

[0034] ,

[0035] If is an untrustworthy IP, its observation probability is obtained according to the following formula:

[0036] ,

[0037] The reputation value of is calculated according to the following formula:

[0038] ,

[0039] where indicates whether the IP is trustworthy. If it is 1, it is trustworthy; if it is 0, it is untrustworthy; represents the observation probability of trustworthy IPs, represents the observation probability of untrustworthy IPs, represents the route trace event observation set, Represents the total number of observed route tracing behavior events, Represents the observed host The number of events for route tracing behavior, Represents the set of trusted IPs, Represents the total number of route tracing behaviors performed in the set of trusted IPs, Represents the set of untrusted IPs, Represents the total number of traceroute behaviors performed in the set of untrusted IPs, Represents the prior probability of a trusted IP, Represents the prior probability of an untrusted IP.

[0040] Furthermore, the step of detecting the newly added route tracing behavior of the IP and updating the reputation score value of the newly added route tracing behavior according to the IP reputation model further includes:

[0041] Obtain the updated reputation score value according to the following formula:

[0042] ,

[0043] where, Represents the updated reputation score value, Represents the record of the last behavior update period of the IP, and t represents the parameter value when the IP has a route tracing behavior in the current observation period.

[0044] Furthermore, the step of detecting the access destination IP entropy of the IPs within the set of untrusted IPs to determine the occurrence of an LFA attack and locate the victim link includes:

[0045] Calculate the access destination IP entropy of the traffic of the set of untrusted IPs according to the following formula:

[0046] ,

[0047] where, Represents the access destination IP entropy of the traffic of the set of untrusted IPs, Represents The destination IP address of the access, Represents the destination address The proportion of the access quantity;

[0048] Locate the victim node according to the following formula:

[0049] ),

[0050] where, Represents the set of victim nodes of the link flooding attack, Represents the source address, Represents the shortest path from the source address to the destination address.

[0051] Further, the step of sampling the traffic of the victim link ingress switch and performing flow feature extraction includes:

[0052] Obtain the set of victim nodes of the link flooding attack, and select at least some switches within the set of victim nodes for traffic sampling.

[0053] Further, the step of training the first-stage self-organizing map clustering model to perform preliminary detection on the sampled flows includes:

[0054] Initialize the weight vector of the neuron node to a random value;

[0055] Input the sample features to each neuron;

[0056] Calculate the distance between the weight of each neuron and the sample feature vector;

[0057] Calculate the influence radius of the neuron;

[0058] Update the weights of all neurons according to the influence radius;

[0059] Repeat the iteration until the maximum number of iterations is reached.

[0060] Further, the step of training the second-stage Gaussian mixture clustering model to perform secondary detection on the sampled flows includes:

[0061] Select the number of Gaussian mixture components, initialize the Gaussian mixture coefficients, initialize the mean vector, and initialize the covariance matrix;

[0062] Calculate the posterior probability of each sample belonging to different Gaussian distributions;

[0063] Update the model parameters;

[0064] Repeat calculating the posterior probability and updating the model parameters until the model converges.

[0065] Further, the step of issuing flow table rules to block the traffic of the malicious IPs identified by clustering includes:

[0066] Calculate the sum of the reputation scores of the source IPs of the samples in each cluster;

[0067] Sort the reputation scores of each cluster to obtain the cluster with the lowest reputation score according to the sorting result.

[0068] In a second aspect, the present application provides a link flooding attack detection and defense system based on reputation and clustering algorithms, and the system includes:

[0069] A model construction module for constructing an IP reputation model, detecting the newly added route tracking behavior of an IP, and updating the reputation score value of the newly added route tracking behavior according to the IP reputation model;

[0070] An IP entropy detection module for detecting the access destination IP entropy of IPs within an untrusted IP set to determine the occurrence of an LFA attack and locate the victim link;

[0071] A flow feature extraction module for sampling the traffic of the switch at the ingress end of the victim link and extracting flow features;

[0072] A first training module for training a first-stage self-organizing map clustering model to perform a preliminary detection on the sampled flows;

[0073] A second training module for training a second-stage Gaussian mixture clustering model to perform a secondary detection on the sampled flows;

[0074] A clustering recognition module for issuing flow table rules to block the traffic of malicious IPs identified by clustering.

[0075] In a third aspect, the present application provides a readable storage medium storing one or more programs, which when executed by a processor implement the above-mentioned link flooding attack detection and defense method based on reputation and clustering algorithms.

[0076] In a fourth aspect, the present application provides a computer device, which includes a memory and a processor, wherein:

[0077] The memory is used for storing a computer program;

[0078] The processor is used for implementing the above-mentioned link flooding attack detection and defense method based on reputation and clustering algorithms when executing the computer program stored on the memory.

[0079] Compared with the prior art, the present application has the following advantages:

[0080] 1. By proposing a reputation model based on the historical route tracking behavior of IPs and a link flooding attack detection method for IP source address entropy anomalies, the method used in the embodiments of the present application has higher detection accuracy and better robustness when the attacker adopts a long-cycle mapping strategy compared with the existing detection methods based on behavior statistics and entropy anomalies.

[0081] 2. The embodiment of the present application designs a link flooding attack traffic detection scheme based on a two-stage clustering algorithm. In the first stage, self-organizing map (SOM) clustering is used, and the main goal is to eliminate benign traffic as much as possible. In the second stage, Gaussian mixture (GMM) clustering is adopted, and the main goal is to improve the detection accuracy as much as possible. Different flow features are designed for the two stages. The designed traffic detection scheme has high detection efficiency and detection accuracy, and can effectively detect link flooding attack traffic and identify zombie host IPs.

[0082] 3. The present application uses a two-stage clustering algorithm to perform anomaly detection on link flooding attack traffic. The training model can use existing public data sets or benign traffic in the real network environment, and has no requirements for attack traffic, effectively solving the problem of lack of link flooding attack traffic data. Brief Description of the Drawings

[0083] Figure 1 is a flowchart of a link flooding attack detection and defense method based on reputation and clustering algorithm proposed in an embodiment of the present application;

[0084] Figure 2 is an overall architecture diagram of a link flooding attack detection and defense method based on reputation and clustering algorithm proposed in an embodiment of the present application;

[0085] Figure 3 is a schematic structural diagram of a link flooding attack detection and defense system based on reputation and clustering algorithm proposed in an embodiment of the present application.

[0086] The following specific embodiments will further illustrate the present application in conjunction with the above-mentioned drawings. Specific Embodiments

[0087] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art in the field to which the present application belongs. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.

[0088] Please refer to Figure 1 , which shows a link flooding attack detection and defense method based on reputation and clustering algorithm provided in an embodiment of the present application. The method includes steps S101 to S105, where:

[0089] Step S101: Build an IP reputation model, detect the newly added route tracing behavior of the IP, and update the reputation score value of the newly added route tracing behavior according to the IP reputation model;

[0090] It should be noted that the built IP reputation model is a Bayesian IP reputation model. This Bayesian IP reputation model calculates the reputation score of the IP by monitoring and statistically analyzing the traceroute behavior (route tracing behavior). In the IP reputation model, the IP can be classified into two categories: trustworthy ( ) and untrustworthy ( ). A trustworthy IP is the IP of a normal host within a network domain (AS), and an untrustworthy IP is the IP of a zombie host (bot) that may be exploited by an attacker to initiate a link flooding attack. Distinguishing between trustworthy ( ) and untrustworthy ( ) IPs is mainly determined according to the reputation score of the IP .

[0091] First of all, during the process of building the IP reputation model, it is set that the distribution of the reputation score follows a normal distribution , and the IP reputation model is built according to the following formula:

[0092] ,

[0093] where, represents the density function, x represents the reputation score value, μ represents the average reputation score, and σ represents the variance of the reputation score.

[0094] Among them, the average reputation score is calculated according to the following formula:

[0095] ,

[0096] The variance of the reputation score is calculated according to the following formula:

[0097] ,

[0098] where, represents the reputation score record table, represents the number of elements in the reputation score record table, represents the reputation score of the IP. In addition, for the average reputation score and the variance of the reputation score , they need to be updated after each observation period to fit the new distribution.

[0099] In some embodiments, the reputation calculation steps of the IP reputation model are specifically as follows:

[0100] (1) The IP traceroute behavior traffic data packets (usually ICMP timeout data packets) are collected through a network traffic collector, and it is determined whether the source address IP in the route tracing behavior traffic data packets is recorded in the IP reputation table. If the source address IP of this data packet is not recorded in the IP reputation table then this IP is a newly emerged IP, denoted as , and the recorded IPs will not be collected repeatedly. If this IP has been recorded in the reputation table then the traceroute behavior event record of this IP is incremented, denoted as ;

[0101] (2) If this IP is , calculate the reputation value of this IP and add it to the reputation table . The formula for calculating the reputation value of this IP is:

[0102] ,

[0103] represents the reputation score obtained by the IP from the external threat intelligence library, which is used to improve the recognition speed of malicious IPs. is normalized to between 0 and 1. is the weight of the external reputation score in the initial reputation. α∈ [0,1] , is the average reputation score of all IPs, and its calculation formula is as follows:

[0104] ,

[0105] is the sum of all IP reputation scores.

[0106] (3) If this IP is , add this IP to the traceroute observation event set , and determine whether this IP is a trusted IP. If this IP is a trusted IP, then update the observation probability , otherwise, update the observation probability , specifically as follows:

[0107] Set the IP with a reputation score less than the lower quartile as an untrusted IP (only as a relative determination of IP untrustworthiness, which does not necessarily mean that the IP is a zombie host IP). The formula for the lower quartile is as follows:

[0108] ,

[0109] wherein, is the lower one - quarter quantile of the standard normal distribution, .

[0110] Specifically, the discrimination formula for setting whether an IP is trustworthy is as follows:

[0111] ,

[0112] That is to say, represents whether the IP is trustworthy. If it is 1, it is trustworthy; if it is 0, it is not trustworthy.

[0113] In addition, the IP reputation score can be obtained by observing and statistically analyzing the traceroute behavior of the IP and updating and calculating the posterior probability through Bayes' formula. Specifically, for a normal host, the probability of performing a traceroute behavior is relatively low because only network administrators may use the traceroute tool to detect network connectivity during network maintenance or fault checking, and usually the frequency is not particularly high. Therefore, in the case where the IP is trustworthy ( ), the probability of observing a traceroute behavior is:

[0114] ,

[0115] wherein, represents the set of observed traceroute event collections, that is, the number of events where the host performs a traceroute behavior, represents the set of trustworthy IPs. The numerator part of the above formula represents the total number of traceroute behaviors performed in the set of trustworthy IPs, and the denominator represents the total number of observed traceroute behavior events. This probability will be dynamically updated according to the current network state and is usually a relatively small value because the probability of a normal and trustworthy IP having a traceroute behavior is relatively low. However, during special periods, such as during a large - scale network failure, network administrators need to use the traceroute tool for fault checking, and this probability will increase, indicating an increase in the number of events where a trustworthy IP performs a traceroute behavior.

[0116] In the case where the IP is not trustworthy ( ), the probability of observing a traceroute behavior event is as follows:

[0117] ,

[0118] Among them, represents the set of observed traceroute event observations, that is, the number of events where the host performs traceroute behavior, represents the set of untrusted IPs, represents the total number of traceroute behaviors performed in the set of untrusted IPs. The denominator represents the total number of observed traceroute behavior events. This prior probability is between 0 and 1. Although in this model, the reduction of reputation is due to the frequent traceroute behavior of the IP, there may also be cases of false detection (frequent operations by network administrators). In addition, the initially identified untrusted IPs (the initial reputation calculation of the IP in this model combines external CIT intelligence) are not necessarily the zombie host IPs controlled by the link flooding attacker. It may also be the host IPs performing other malicious behaviors, such as worms, Trojans, conventional DDoS, spam, etc. But if the IP is untrusted ( ), then the probability of observing it performing traceroute behavior is usually much higher than that of a trusted IP performing traceroute behavior. Therefore, this observation probability is usually relatively high, and the specific probability will be dynamically updated and adjusted according to the observation results of network traceroute events.

[0119] (4) Calculate the reputation score of the IP , and update the reputation score value of this IP in the reputation table .

[0120] Specifically, based on the above two observation probabilities, when observing the traceroute event of the IP, the credible probability of the IP can be updated by calculating the Bayesian posterior probability, that is, the reputation value , and the specific formula is as follows:

[0121] ,

[0122] Among them, that is, the prior probability that the current IP is credible, is the observation probability of the event, and the denominator is the total probability. And the untrusted probability of the IP is:

[0123] ,

[0124] Among them, represents the observation probability of a trusted IP, represents the observation probability of an untrusted IP, Represents a set of observations of route tracing events, Represents the total number of observed route tracing behavior events, Represents observing the host The number of events for which route tracing behavior is performed, Represents a set of trusted IPs, Represents the total number of route tracing behaviors performed within the set of trusted IPs, Represents a set of untrusted IPs, Represents the total number of traceroute behaviors performed within the set of untrusted IPs, Represents the prior probability of a trusted IP, Represents the prior probability of an untrusted IP.

[0125] In addition, it should be noted that the higher the trust probability of an IP, the higher the reputation value. The range of the reputation value of an IP R IP ∈ [0,1] .

[0126] In addition, it should also be explained that the reputation value of a normal IP should gradually recover after a long period without traceroute behavior, and a malicious IP can also be considered to have returned to normal after a long period of inactivity. To ensure that the reputation system correctly evaluates the trust probability of an IP, when an observation period ends, it is necessary to restore the reputation scores of those IPs that have not performed any traceroute behavior. Usually, this observation period is relatively long to prevent link flood attackers from using long-period detection strategies to evade the reputation system detection. This update period is recommended to be associated with the network topology update period (the old topology detected by the attacker using long-period detection cannot be used to launch link flood attacks). The reputation score recovery formula is as follows:

[0127] ,

[0128] where, Represents the updated reputation score value, Represents the record of the last behavior update period of the IP. When the IP has a traceroute behavior in the current observation period, this parameter value is updated to , and the recovery weight function of the second term of this formula is a Sigmoid function, which can make the reputation value recover relatively slowly at the beginning and accelerate the recovery as the observation period increases, making it more difficult for the attacker's long-period detection strategy to evade the reputation system detection.

[0129] Step S102: Detect the access destination IP entropy of the IPs within the set of untrusted IPs to determine the occurrence of LFA attacks and locate the victim link;

[0130] It should be noted that in this step, when a link flooding attack occurs, the zombie hosts (Bots) controlled by the attacker will access the decoy at a low and normal traffic rate, superimposing the normal traffic and blocking the critical link. Since the resources of the decoy available for blocking the critical link are limited, the destination IP addresses of the accesses from the zombie hosts (IPs) will be highly similar, while the destination IP addresses of the accesses from normal host IPs are always uniformly random. Therefore, when a link flooding attack occurs, the entropy of the traffic destination IPs of the zombie hosts (bots) will be abnormal.

[0131] In the IP reputation record table of the aforementioned Bayesian IP reputation model it can be considered that IPs with low reputation values are more likely to be the zombie host IPs participating in the link flooding attack. Therefore, only by calculating the entropy of the access destination IP addresses of the low-reputation IPs in the reputation table and monitoring for abnormalities can the occurrence of the link flooding attack be effectively judged. Through the global characteristics of SDN, the attack traffic path (usually the shortest path) can be calculated, and the overlapping nodes of the attack traffic path can be calculated to locate the ingress end node (switch) of the victim link. The relevant calculation formulas are as follows:

[0132] ,

[0133] where, represents the entropy of the access destination IPs of the traffic of the untrusted IP set, represents the access destination IP address, represents the access quantity ratio of the destination address ;

[0134] The formula for locating the victim node (the ingress end switch of the victim link) is as follows:

[0135] ),

[0136] where, represents the set of victim nodes of the link flooding attack, represents the source address, represents the shortest path from the source address to the destination address, and there may be multiple victim nodes.

[0137] Step S103: Sample the traffic of the ingress end switch of the victim link and perform flow feature extraction;

[0138] In this step, specifically according to the set of victim nodes of the link flooding attack obtained in step S102 , and select all or part of the switches in the set for traffic sampling.

[0139] Step S104: Train the first - stage self - organizing map clustering model to conduct a preliminary detection on the sampled flow;

[0140] It should be noted that in this step, in the first stage, the self - organizing map (SOM) clustering algorithm is adopted to make the first distinction between attack flows and benign flows. The purpose is to eliminate the benign flows as much as possible to improve the detection efficiency.

[0141] In addition, the main features of the self - organizing map (SOM) clustering algorithm adopted in the first stage are strong dynamic learning ability of the model and high training efficiency. The main task of the model in this stage is to learn the characteristics of benign flows. The data can be collected from public data sets or real - network environments. The model in this stage is trained only with benign traffic. To improve the efficiency of detection and feature extraction, each network flow is marked with a tuple (src_ip, dst_ip). The flow features used for training the model in this stage are: dst_ip_speed: the access speed of different destination IPs (counted every 10 s); dst_ip_novelty_speed: the number of new destination IPs compared to the previous time window; flow_duration: the duration of the flow (link attack flows usually have a longer duration); pkt_rate: the rate of packets sent by the source IP (zombie sources always send packets at a low rate); pkt_size: the average size of the packets sent by the source IP (there are significant differences between normal sources and malicious sources).

[0142] In some embodiments, the first training stage specifically includes the following steps:

[0143] Step1 (Initialization): The weight vector ( - ) of the neuron node ( ) is initialized to random values.

[0144] Step2 (Input): Input the sample features into each neuron ( - ).

[0145] Step3 (Competition): Calculate the distance between the weight of each neuron and the sample feature vector , and the calculation formula is as follows:

[0146] = ,

[0147] where D is the feature dimension, is the feature value, is the neuron weight, and select the winning neuron C with the shortest distance from the sample feature vector, denoted as 。

[0148] Step4 (Learning): Calculate the influence radius of the neuron as follows: The calculation formula is as follows:

[0149] ,

[0150] where, is the current iteration number, is a constant related to the maximum iteration number and the initial radius The calculation formula is as follows:

[0151] ,

[0152] Update the weights of all neurons according to the influence radius The formula is as follows:

[0153] ,

[0154] where, is the weight of the corresponding neuron at the t-th round of the current iteration number is a weight function, and the calculation formula is as follows:

[0155] ,

[0156] where, is the Euclidean distance between the updated neuron and the winning neuron ( ), is the learning rate, which decreases as the iteration number increases. The calculation formula is:

[0157] ,

[0158] Setp 5 (Iteration): Repeat Step2 to Step4 until the maximum iteration number is reached.

[0159] Step S105: Train the Gaussian mixture clustering model in the second stage to perform secondary detection on the sampling flow;

[0160] It should be noted that the Gaussian mixture (GMM) clustering algorithm is used in the second stage to further distinguish the flow detection results in the first stage, aiming to improve the detection accuracy of malicious flows as much as possible.

[0161] In addition, the main feature of the Gaussian Mixture Model (GMM) adopted in the second stage is its strong learning ability for the complex feature distribution of sample data. The purpose of the model in this stage is to detect malicious flows of link flooding attacks as precisely as possible. To detect attack flows more accurately, considering that the link flooding attack flows have characteristics significantly different from those of benign flows on the traffic behavior graph, five features on the traffic graph are added in this stage to more precisely distinguish attack flows. The newly added features are described as follows: ID: In-degree of the node, the number of destination IPs accessed by the node (IP). The number of accesses from zombie hosts (Bots) to decoys is valid and generally lower than that of benign IP nodes; OD: Out-degree of the node, the number of source IPs accessed by the node (IP); IDW: In-degree weight of the node, the number of received flow data packets of the node (IP); ODW: Out-degree weight of the node, the number of sent flow data packets of the node (IP); LLC: Local clustering coefficient. Since the number of available decoys for zombie hosts (Bots) is limited and the overlap is high, malicious IP nodes will have a relatively high local clustering coefficient (LLC).

[0162] In some embodiments, the specific process of the second training stage is as follows:

[0163] Step1 (Initialization): Select the number of Gaussian mixture components , initialize the Gaussian mixture coefficients , initialize the mean vector , initialize the covariance matrix .

[0164] Step2 (Calculate posterior probability): Calculate the posterior probability of each sample belonging to different Gaussian distributions:

[0165] ( ) = ,

[0166] where, ( is the posterior probability that the sample belongs to the th Gaussian component, represents the mean vector of the jth distribution, represents the covariance matrix of the jth distribution, is the Gaussian function, and the formula is as follows:

[0167] ,

[0168] where, is the sample feature vector, with a dimension of , is the mean vector, is the covariance matrix, is the determinant of the covariance matrix.

[0169] Step3 (Update model parameters): Calculate the posterior probability ( After that, the model parameters to be updated are as follows:

[0170] ,

[0171] where is the total probability of the sample in the k-th distribution.

[0172] ,

[0173] where is the Gaussian mixture coefficient.

[0174] ,

[0175] where is the mean vector of the k-th distribution.

[0176] ,

[0177] where is the covariance matrix of the k-th distribution.

[0178] Step4 (Iteration): Repeat steps Step2 to Step3 until the model converges.

[0179] Step S106: Issue a flow table rule to block the traffic of the malicious IPs identified by clustering.

[0180] In this step, after extracting the features of the flow by sampling the network flow and clustering the flow, the method for judging the category of the clustering cluster is as follows:

[0181] First, calculate the sum of the reputation scores of the source IPs of the samples in each clustering cluster (if the IP exists in the reputation table ), and the formula is as follows:

[0182] ,

[0183] is the i-th clustering cluster, is the reputation score of the IP.

[0184] Then, sort the reputation scores of each cluster:

[0185] ,

[0186] Finally, select the cluster with the lowest reputation score according to the above sorting result, which is the classification cluster of the zombie host IPs that may be used for link flooding attacks.

[0187] In addition, it should also be pointed out that the above method for judging clustering clusters is applicable to the clustering algorithms in both training stages. After detecting the zombie host IPs that may be used for link flooding attacks, the flow table rules can be issued through the SDN controller module, such as Figure 2 shown, to block the traffic of the corresponding zombie host IPs.

[0188] In summary, according to the above link flooding attack detection and defense method based on reputation and clustering algorithm, it has the following advantages:

[0189] 1. By proposing a reputation model based on IP historical routing tracking behavior and a link flooding attack detection method based on IP source address entropy anomaly in the embodiments of the present application, compared with the existing detection methods based on behavior statistics and entropy anomaly, the method used in the embodiments of the present application has higher detection accuracy and better robustness when the attacker adopts a long-term mapping strategy.

[0190] 2. By designing a link flooding attack traffic detection scheme based on a two-stage clustering algorithm in the embodiments of the present application, the self-organizing map (SOM) clustering is used in the first stage, and the main goal is to eliminate benign traffic as much as possible. The Gaussian mixture (GMM) clustering is used in the second stage, and the main goal is to improve the detection accuracy as much as possible. Different flow features are designed for the two stages. The designed traffic detection scheme has high detection efficiency and detection accuracy, and can effectively detect link flooding attack traffic and identify zombie host IPs.

[0191] 3. Using the two-stage clustering algorithm to detect anomalies in link flooding attack traffic in the present application, the training model can use existing public data sets or benign traffic in the actual network environment, and has no requirements for attack traffic, effectively solving the problem of lack of link flooding attack traffic data.

[0192] Please refer to Figure 3 shown in the structural schematic diagram of the link flooding attack detection and defense system based on reputation and clustering algorithm in an embodiment of the present application. The system includes:

[0193] A model construction module 10, configured to construct an IP reputation model, detect the newly added routing tracking behavior of an IP, and update the reputation score value of the newly added routing tracking behavior according to the IP reputation model;

[0194] An IP entropy detection module 20, configured to detect the access destination IP entropy of the IPs in the untrusted IP set to judge the occurrence of the LFA attack and locate the victim link;

[0195] The flow feature extraction module 30 is used to sample the traffic of the switch at the ingress end of the victim link and extract flow features;

[0196] The first training module 40 is used to train the self-organizing map clustering model in the first stage to preliminarily detect the sampled flows;

[0197] The second training module 50 is used to train the Gaussian mixture clustering model in the second stage to perform secondary detection on the sampled flows;

[0198] The clustering and recognition module 60 is used to issue flow table rules to block the traffic of the malicious IPs identified by clustering.

[0199] On the other hand, the present application also proposes a readable storage medium, on which one or more programs are stored, and when the program is executed by a processor, the above-mentioned link flooding attack detection and defense method based on reputation and clustering algorithm is implemented.

[0200] On the other hand, the present application also proposes a computer device, including a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned link flooding attack detection and defense method based on reputation and clustering algorithm.

[0201] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0202] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0203] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0204] Although the embodiments of the present application have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are all within the scope and spirit of the present application as described in the claims. Moreover, the present application described herein can have other embodiments and can be implemented or realized in various ways.

Claims

1. A method for detecting and defending link flooding attacks based on reputation and clustering algorithms, characterized in that The method includes: Construct an IP reputation model, detect the newly added route tracing behavior of an IP, and update the reputation score value of the newly added route tracing behavior according to the IP reputation model; Detect the access destination IP entropy of the IPs within the untrusted IP set to determine the occurrence of an LFA attack and locate the victim link; Sample the traffic of the switch at the entrance end of the victim link and extract flow features; Train the first-stage self-organizing map clustering model to conduct a preliminary detection of the sampled flows; Train the second-stage Gaussian mixture clustering model to conduct a secondary detection of the sampled flows; Issue flow table rules to block the traffic of the malicious IPs identified by clustering.

2. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 1, characterized in that The steps of constructing the IP reputation model include: Construct the IP reputation model according to the following formula: , Among them, represents the density function, represents the credit score value, represents the average credit score, represents the variance of the credit score; Calculate the average value of the reputation scores according to the following formula: , Calculate the variance of the reputation scores according to the following formula: , Among them, represents the credit score record table, represents the number of elements in the credit score record table, represents the credit score of the IP.

3. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 2, wherein, The steps of detecting the newly added route tracing behavior of an IP and updating the reputation score value of the newly added route tracing behavior according to the IP reputation model include: Collect the routing trace behavior traffic data packets of the IP, and determine whether the source address IP in the routing trace behavior traffic data packets is recorded in the IP reputation score record table. If it is not recorded, record it as , if it has been recorded, record it as ; If the IP is , then obtain its credit score according to the following formula: , Obtain the average reputation score of all IPs according to the following formula: , Among them, represents the reputation score of the IP, represents the external reputation score of the IP obtained from the external threat intelligence library, is the weight of the external reputation score in the initial reputation, is the average reputation score of all IPs, is the sum of the reputation scores of all IPs; If the IP is , then add to the routing trace observation event set , and determine whether is credible according to the following formula: , If is a trusted IP, its observation probability is obtained according to the following formula: , If is an untrusted IP, obtain its observation probability according to the following formula: , Calculated according to the following formula 's credit score: , Among them, indicates whether the IP is trustworthy. If it is 1, it is trustworthy; if it is 0, it is not trustworthy; represents the observation probability of trustworthy IPs, represents the observation probability of untrustworthy IPs, represents the set of traceroute observation events, represents the total number of observed traceroute behavior events, represents observing the host the number of events of performing traceroute behavior, represents the set of trustworthy IPs, represents the total number of traceroute behaviors performed in the set of trustworthy IPs, represents the set of untrustworthy IPs, represents the total number of traceroute behaviors performed in the set of untrustworthy IPs, represents the prior probability of trustworthy IPs, represents the prior probability of untrustworthy IPs.

4. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 3, characterized in that, The steps of detecting the newly added route tracing behavior of an IP and updating the reputation score value of the newly added route tracing behavior according to the IP reputation model further include: Obtain the updated reputation score according to the following formula: , Among them, represents the updated credit score, represents the record of the last behavior update period of the IP, and t represents the parameter value when the IP has a traceroute behavior in the current observation period.

5. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 4, characterized in that The steps of detecting the access destination IP entropy of the IPs within the untrusted IP set to determine the occurrence of an LFA attack and locate the victim link include: Calculate the access destination IP entropy of the traffic of the untrusted IP set according to the following formula: , Among them, represents the access destination IP entropy of the untrusted IP set traffic, represents the destination IP address of the access, represents the destination address of the access quantity ratio; Locate the victim node according to the following formula: ), Among them, represents the set of victim nodes of the link flooding attack, represents the source address, represents the shortest path from the source address to the destination address.

6. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 1, characterized in that The steps of sampling the traffic of the switch at the entrance end of the victim link and extracting flow features include: Obtain the set of victim nodes of the link flooding attack and select at least some switches within the set of victim nodes for traffic sampling.

7. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 1, characterized in that The steps of training the first-stage self-organizing map clustering model to conduct a preliminary detection of the sampled flows include: Initialize the weight vectors of the neuron nodes to random values; Input the sample features to each neuron; Calculate the distance between the weight of each neuron and the sample feature vector; Calculate the influence radius of the neuron; Update the weights of all neurons according to the influence radius; Repeat the iteration until the maximum number of iterations is reached.

8. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 7, characterized in that The steps of training the second-stage Gaussian mixture clustering model to conduct a secondary detection of the sampled flows include: Select the number of Gaussian mixture components, initialize the Gaussian mixture coefficients, initialize the mean vectors, and initialize the covariance matrices; Calculate the posterior probability of each sample belonging to different Gaussian distributions; Update the model parameters; Repeat the calculation of the posterior probability and update the model parameters until the model converges.

9. The method for detecting and defending link flooding attacks based on reputation and clustering algorithms according to claim 8, characterized in that, The steps of issuing flow table rules to block the traffic of the malicious IPs identified by clustering include: Calculate the sum of the reputation scores of the source IPs of the samples in each clustering cluster; Sort the sums of the reputation scores of each cluster to obtain the cluster with the lowest sum of the reputation scores according to the sorting result.

10. A link flooding attack detection and defense system based on reputation and clustering algorithms, characterized in that The system includes: A model construction module for constructing an IP reputation model, detecting the newly added route tracing behavior of an IP, and updating the reputation score value of the newly added route tracing behavior according to the IP reputation model; The IP entropy detection module is used to detect the access destination IP entropy of the IPs within the untrusted IP set, so as to judge the occurrence of the LFA attack and locate the victim link; The flow feature extraction module is used to sample the switch traffic at the entrance end of the victim link and extract flow features; The first training module is used to train the self-organizing map clustering model in the first stage to conduct a preliminary detection of the sampled flows; The second training module is used to train the Gaussian mixture clustering model in the second stage to conduct a secondary detection of the sampled flows; The clustering identification module is used to issue flow table rules to block the traffic of the malicious IPs identified by clustering.

Citation Information

Patent Citations

  • Link flooding attack defense method based on selective real-time rerouting

    CN118432912A

  • Architecture, systems and methods to detect efficiently DoS and DDoS attacks for large scale internet

    US7584507B1