DNS attack detection and defense system based on adversarial network
Through the DNS attack detection and defense system based on adversarial networks, the generation of adversarial networks is used to generate adversarial samples, efficient identification and automated defense of DNS attacks are achieved, and the problem of insufficient adaptability and accuracy in the existing technology is solved, and the adaptability and detection capabilities of the DNS defense system are improved.
Patent Information
- Application Number
- CN202510617964.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing DNS attack detection methods have insufficient adaptability and accuracy, and it is difficult to deal with rapidly changing attack patterns, resulting in high false positive rates and detection delays.
DNS attack detection and defense system based on adversarial network is adopted, and through data acquisition and feature engineering modules, adversarial sample generation modules, multimodal detection engine modules and dynamic defense and feedback systems, adversarial attack samples are generated in combination with Generation Adversarial Networks (GANs), for real-time analysis and automated defense.
It improves the adaptability and detection capabilities of the DNS defense system, reduces the false alarm rate, and can quickly identify new attacks and perform adaptive optimization.
Smart Images

Figure CN120301683A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of DNS (Domain Name System), and specifically relates to a DNS attack detection and defense system based on adversarial networks. Background Art
[0002] There are currently two common DNS attack detection methods:
[0003] One is the rule-based detection method, which identifies abnormal behaviors in DNS traffic through predefined rules. For example, by setting thresholds to detect abnormal query frequencies or specific types of DNS requests. However, this method has the following disadvantages:
[0004] 1. Poor adaptability: The rule-based detection method relies on predefined rules and is difficult to cope with constantly changing attack patterns.
[0005] 2. High false positive rate: In a complex network environment, the rules may not be able to accurately distinguish normal traffic from attack traffic, resulting in a relatively high false positive rate.
[0006] The other is the statistics-based detection method, which identifies abnormal behaviors by analyzing the statistical characteristics of DNS traffic, such as request frequency, response time, etc.; this method has the following disadvantages:
[0007] 1. High detection latency: The statistics-based detection method requires collecting a large amount of data for accurate analysis, resulting in a relatively high detection latency.
[0008] 2. Limited ability to detect new attacks: For new DNS attack patterns, the statistical method may not be able to identify them in a timely manner.
[0009] With the continuous evolution of network attacks, DNS attacks have increasingly become an important threat to network security. Traditional DNS attack detection methods are difficult to adapt to rapidly changing attack patterns, so a new type of adversarial detection mechanism is needed to enhance the robustness of the DNS defense system. Summary of the Invention
[0010] Aiming at the problems existing in the existing DNS security protection system, such as low detection accuracy, weak adversarial ability, and inability to dynamically respond to new attack behaviors, the present invention proposes a DNS attack detection and defense system based on adversarial networks, which dynamically optimizes the accuracy and efficiency of DNS attack detection by simulating the game process between the attacker and the detection system.
[0011] The DNS attack detection and defense system based on adversarial networks consists of a data collection and feature engineering module, an adversarial sample generation module, a multi-modal detection engine module, and a dynamic defense and feedback system;
[0012] During the system operation, first, the data collection and feature engineering module captures DNS traffic in real time and parses the protocol, extracts multi-dimensional behavioral features, and combines the rule engine, threat intelligence, historical samples, and trap results to preliminarily label the DNS traffic, generating structured feature data, which is divided into two categories: normal samples and attack samples.
[0013] Subsequently, the attack samples and normal samples are respectively sent into the adversarial sample generation module. By introducing a generative adversarial network (GAN) optimized with the Wasserstein distance, strongly camouflaged adversarial attack samples are constructed for the attack samples.
[0014] The normal samples and adversarial attack samples are used as mixed samples and input into the multi-modal detection engine module. Relying on the attention mechanism, convolutional neural network, and time series model, the input mixed samples are fused and analyzed, and the attack probability and classification results are output.
[0015] Finally, the classification results are transmitted to the dynamic defense and feedback system, which automatically executes response strategies such as BGP blocking and domain name redirection. At the same time, the mixed samples and detection results are fed back into the database for system update and policy optimization, forming a closed-loop adaptive defense mechanism.
[0016] The data collection and feature engineering module is deployed at the DNS server entrance or network boundary, used to collect DNS traffic data in real time, and perform structured feature extraction and preliminary annotation on it; it includes a data collection layer, a protocol parsing unit, a three-level feature extraction unit, and a mixed annotation unit.
[0017] Among them, the data collection layer is based on an FPGA hardware probe to achieve 10Gbps line speed packet capture.
[0018] The protocol parsing unit uses the DPDK high-performance framework to parse protocols such as DNS, DoT, and DoH, supports the decryption of encrypted traffic, and relies on a pre-set CA certificate library.
[0019] The three-level feature extraction unit extracts three types of features by dimension: basic features, advanced features, and payload features.
[0020] The mixed annotation unit, based on the historical sample library and threat intelligence source, provides IOC context information, and at the same time combines the attack traffic captured by the disguised DNS server to perform preliminary annotation to obtain a 48-dimensional structured feature vector, which is divided into normal samples and attack samples.
[0021] The adversarial sample generation module is designed based on the generative adversarial network (GAN), including a generator and a discriminator, and uses the Wasserstein distance as the optimization target to generate adversarial attack samples.
[0022] Among them, the generator takes the attack samples as input and learns from the attack sample distribution Ps Optimal projection mapping relationship to the normal sample distribution P t to generate adversarial samples with a structure approximately normal but retaining attack features.
[0023] Generator A θ (·) has a training objective of minimizing the Wasserstein distance between the adversarial attack samples and the normal sample distribution, and its loss function is defined as:
[0024]
[0025] where x ∼ P s represents the input sampled from the attack sample distribution, and A θ (x) represents the projected output of the generator for the attack samples. By adjusting θ, the Wasserstein distance between the generated adversarial attack samples and the normal sample distribution is minimized. D(·) is a real-valued function output by the discriminator.
[0026] The training objective of the discriminator is to maximize the Wasserstein distance between the normal sample distribution P t and the adversarial attack sample distribution, so as to evaluate the "disguise" of the adversarial samples, and its loss function is defined as:
[0027]
[0028] where x′ ∼ P t represents the input sampled from the normal sample distribution, and D w (·) is a discriminant function that satisfies the 1-Lipschitz constraint, and the output value is used to measure the distance between the adversarial attack samples and the normal samples. By adjusting w, the Wasserstein distance between the normal sample distribution and the adversarial attack sample distribution is maximized. A(x) represents the output result of the generator for the attack samples.
[0029] The parameters w and θ of the discriminator and the generator are optimized independently. In each round of training, first fix the generator parameter θ, and perform multiple rounds of gradient ascent updates on the discriminator parameter w to maximize the Wasserstein distance; then fix the discriminator parameter w and minimize the generator loss function to guide the generator to learn a mapping strategy that approaches the normal sample distribution. When the Wasserstein distance is less than 0.05, generate multiple types of adversarial attack samples including DNS tunnel attack samples and DDoS samples; and after mixing the generated adversarial attack samples with the normal samples, input them into the multi-modal detection engine module.
[0030] The multimodal detection engine module performs fusion recognition and attack classification judgment on the input mixed samples; specifically, it is divided into an input layer, a feature fusion layer, and an output layer, and is equipped with a dynamic threshold adjustment mechanism;
[0031] The input layer receives the 48-dimensional normalized feature vector of the sample and inputs it to the feature fusion layer. The key context information is extracted through a four-head attention mechanism, and at the same time, the spatial structure feature F extracted by the CNN t and the time series behavior feature F extracted by the GRU s are cross-fused in space and time to obtain the fused feature F fusion ; the fusion process formula is as follows:
[0032]
[0033] This fused feature is then input into the Softmax classifier to calculate the attack probability score P attack ; the value range of this attack probability is [0,1], indicating the confidence that the sample is an attack.
[0034] In the real-time detection process, the classification threshold is dynamically adjusted for the sliding window data according to the KL divergence, and the adjustment formula is:
[0035] T = T0 + α·D KL (P||Q)
[0036] T0 is the initially set threshold, α is the adjustment factor, and D KL (P||Q) represents the divergence between the current model output distribution and the normal distribution.
[0037] If the attack probability P attack ≥ T, it is determined that the mixed sample is abnormal; and the detection result is synchronously transmitted to the dynamic defense and feedback system.
[0038] The dynamic defense and feedback system executes an automated defense strategy according to the detection result and realizes the self-learning and continuous optimization of the detection system. After an attack is confirmed, the BGP FlowSpec rule is automatically issued to block malicious domain name resolution; at the same time, the adversarial attack samples are injected into the training set through an online learning mechanism, and the detection model is incrementally updated weekly.
[0039] The advantages of the present invention are:
[0040] 1. A DNS attack detection and defense system based on the adversarial network integrates data collection, feature extraction, adversarial training, intelligent detection, and automatic defense through multi-module collaboration. It has the advantages of automated processing flow, strong detection ability, and good adaptability. It can effectively improve the security protection level of DNS infrastructure and is applicable to attack detection and response scenarios in operator networks, enterprise intranets, and various public DNS service platforms.
[0041] 2. A DNS attack detection and defense system based on the adversarial network combines hardware and software: integrating FPGA acceleration and DPDK optimization to solve the bottleneck of high-bandwidth traffic capture and parsing.
[0042] 3. A DNS attack detection and defense system based on the adversarial network adopts an adversarial training mechanism based on the Wasserstein distance: using the optimal transport theory to generate adversarial samples with higher camouflage, thus enhancing the defense ability of the model.
[0043] 4. A DNS attack detection and defense system based on the adversarial network uses a detection model that fuses spatio-temporal features: applying a multi-modal deep learning model to DNS traffic analysis to improve the recognition ability for complex attack patterns.
[0044] 5. A DNS attack detection and defense system based on the adversarial network has an automated and self-learning defense mechanism: introducing an online learning and feedback mechanism to enable the system to continuously optimize when facing new attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of a DNS attack detection and defense system based on the adversarial network of the present invention;
[0046] Figure 2 It is a structural diagram of the multi-modal detection engine module adopted by the system of the present invention;
[0047] Figure 3 It is an actual application flowchart of a DNS attack detection and defense system based on the adversarial network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following further elaborates in detail the specific implementation method of the present invention with reference to the drawings.
[0049] The present invention proposes a DNS attack detection and defense system based on the adversarial network, which uses a generative adversarial network (GAN) to generate diverse DNS attack samples and combines a real-time feedback and optimization mechanism to enable the system to quickly adapt to new attack patterns and enhance the detection ability for unknown attacks. The present invention brings the following benefits:
[0050] 1. Introduce GAN to generate adversarial samples: Simulate new types of attacks through adversarial training to improve the detection system's ability to recognize unknown attacks.
[0051] 2. Dynamically optimize the detection model: Optimize GAN and the detection system based on real-time feedback to enhance adaptability.
[0052] 3. Conduct multi-dimensional analysis of DNS traffic using deep learning: Combine multiple DNS features (such as query frequency, TTL value, query domain name pattern, etc.) to improve detection accuracy and reduce false alarm rates.
[0053] The DNS attack detection and defense system based on the adversarial network, as Figure 1 shown, consists of a data collection and feature engineering module, an adversarial sample generation module, a multi-modal detection engine module, and a dynamic defense and feedback system; through the collaboration of real-time traffic analysis, deep learning models, and automated defense strategies, it realizes the accurate identification and dynamic disposal of attack behaviors.
[0054] During the operation of the system, first, the data collection and feature engineering module captures DNS traffic in real-time and parses the protocol, extracts multi-dimensional behavioral features including request frequency, response latency, domain name entropy value, TTL fluctuation, and packet length, and combines rule engines, threat intelligence, historical samples, and trap results to preliminarily label the DNS traffic, generate structured feature data, and divide it into two categories: normal samples and attack samples for subsequent modules to use respectively. This module adopts multi-source fusion combined with annotation strategies, supports semi-automatic classification under large-scale traffic, and continuously optimizes the label accuracy and the system's self-learning ability through the model feedback mechanism.
[0055] Subsequently, the attack samples and normal samples are respectively sent into the adversarial sample generation module, and a generative adversarial network (GAN) optimized by introducing the Wasserstein distance is used to construct highly camouflaged adversarial attack samples to enhance the robustness of the downstream detection model;
[0056] The normal samples and adversarial attack samples are jointly input into the multi-modal detection engine module, which relies on the attention mechanism, convolutional neural network, and time series model to perform fusion analysis on the input mixed sample data, and outputs the attack probability and classification results to train the detection ability of the module.
[0057] Finally, the classification results are transmitted to the dynamic defense and feedback system, which automatically executes response strategies such as BGP blocking and domain name redirection. At the same time, the identified adversarial attack samples and detection results are fed back into the database for system update and policy optimization, forming a closed-loop adaptive defense mechanism.
[0058] The data collection and feature engineering module is deployed at the DNS server entrance or network boundary, and is used to collect DNS traffic data in real time and perform structured feature extraction and preliminary annotation on it; it includes a data collection layer (traffic probe), a protocol parsing unit, a three-level feature extraction unit, and a mixed annotation unit.
[0059] Among them, the data collection layer deploys a hardware probe (traffic mirroring based on FPGA) at the DNS server entrance, supporting 10Gbps line speed packet capture.
[0060] The protocol parsing unit uses the DPDK high-performance framework to parse protocols such as DNS, DoT, and DoH, optimizes the parsing efficiency, supports the decryption of DoT / DoH encrypted traffic, and relies on a pre-set CA certificate library.
[0061] The three-level feature extraction unit extracts three types of features by dimension: basic features (such as request frequency QPS, response latency, record type ratio), advanced features (such as domain name information entropy, TTL value standard deviation), and payload features (such as message length, sub-domain depth).
[0062] o Request frequency: Count the number of queries per second (QPS) for a single client
[0063] o Response latency: Calculate the time difference between the request and the response (precision at the μs level)
[0064] o Record type distribution: Proportion of A / AAAA / MX records (threshold: proportion of A records in normal traffic > 80%)
[0065] o Domain name entropy value: Calculate the information entropy based on character distribution (formula: H(X) = -P(x i )log2P(x i ))), threshold: > 3.5 triggers an alarm)
[0066] o TTL fluctuation analysis: Detect the standard deviation of TTL values (standard deviation in abnormal scenarios > 300 seconds)
[0067] Among them, the domain name information entropy value is calculated based on character distribution, and the formula is:
[0068]
[0069] Among them, H(x) represents the information entropy value of the domain name x; x represents the complete domain name string as input; x i represents the i-th unique character in the domain name; P(x i ) represents the probability that the character x i appears in this domain name, that is, the number of times this character appears divided by the total length of the domain name.
[0070] The information entropy value reflects the degree of dispersion of the domain name character distribution. The higher the entropy value, the more evenly distributed and disordered the characters are, usually indicating that the domain name may be automatically generated by a program or embedded with encoded content (such as Base64). The threshold of this system is set to 3.5. If the domain name entropy value H(x) > 3.5, it will be marked as a high-risk domain name and used as one of the features for judging attack behavior. In addition, abnormal indicators such as TTL fluctuations greater than 300 seconds will also trigger suspicious traffic marking.
[0071] o Payload length: Statistic the DNS packet length distribution (tunnel attack length > 512B)
[0072] o Subdomain depth: Analyze the hierarchical structure (e.g., the depth of a.b.c.malicious.com is 4)
[0073] The hybrid annotation unit is based on the historical sample library (such as loading the CIC-DNS2021 annotation data, and the malicious samples cover three categories: spam, phishing, and malware) and threat intelligence sources (such as the AlienVault OTX API, which updates IOC indicators hourly), provides IOC context information, and combines the attack traffic captured by the disguised DNS server to perform preliminary annotation to obtain a 48-dimensional structured feature vector, which is divided into normal samples and attack samples. As shown in Table 1:
[0074] Table 1
[0075]
[0076]
[0077] The adversarial sample generation module is designed based on the generative adversarial network (GAN), aiming to generate highly disguised adversarial attack samples to enhance the robustness of the detection model. It includes a generator and a discriminator, and uses the Wasserstein distance as the optimization objective to generate adversarial attack samples; improving the alignment of the distribution of the generated samples and normal samples.
[0078] Among them, the generator adopts a multi-layer perceptron structure, which consists of 3 fully connected networks, namely the input layer (128-dimensional random noise), two hidden layers (with sizes of 256 and 128 in sequence), the activation function uses LeakyReLU (α = 0.2), the output layer is mapped to a 48-dimensional structured vector, and is normalized by the Tanh function.
[0079] The generator takes the labeled attack samples as input and learns the optimal projection mapping relationship from the attack sample distribution P s to the normal sample distribution P t so as to generate adversarial samples with a structure similar to normal but retaining attack characteristics.
[0080] The generator A θ (·)'s training objective is to minimize the Wasserstein distance between the generated adversarial samples and the normal sample distribution, and its loss function is defined as:
[0081]
[0082] where x ∼ P s represents the input sampled from the attack sample distribution, and A θ (x) represents the projected output of the generator for the attack samples. By adjusting θ, the Wasserstein distance between the generated adversarial samples and the normal sample distribution is minimized. D(·) is a real-valued function output by the discriminator.
[0083] The discriminator is composed of a convolutional neural network (using 3×3 and 5×5 parallel convolutional kernels) combined with an attention mechanism, and outputs the distance score between the adversarial attack samples and the normal sample distribution.
[0084] The discriminator is a real-valued function that satisfies 1-Lipschitz continuity. Its training objective is to maximize the Wasserstein distance between the normal sample distribution P t and the adversarial sample distribution generated by the generator, so as to evaluate the "disguise" of the adversarial samples. Its loss function is defined as:
[0085]
[0086] where x' ∼ P t represents the input sampled from the normal sample distribution. D(·) is a discriminant function that satisfies the 1-Lipschitz constraint, and the output value is used to measure the distance between the adversarial samples and the normal samples. By adjusting w, the Wasserstein distance between the normal sample distribution and the adversarial sample distribution is maximized.
[0087] The parameters w and θ of the discriminator and the generator are optimized independently, and an alternating training strategy is adopted during the training process. In each round of training, the system first fixes the generator parameter θ and performs multiple rounds of gradient ascent updates on the discriminator parameter w to maximize the Wasserstein distance; then fixes the discriminator parameter w and minimizes the generator loss function to guide the generator to learn the mapping strategy closer to the normal sample distribution.
[0088] Different from traditional adversarial example construction methods that focus on perturbing individual samples, the present invention adopts an adversarial example generation strategy based on distribution mapping. Its core idea is: by learning the optimal mapping relationship from the source domain attack sample distribution to the target domain normal sample distribution, data that is close to normal traffic in the overall feature space is generated, thereby enhancing the authenticity and deception of the generated samples. To achieve the above goal, the Wasserstein distance is introduced as a metric function for evaluating the difference between the original distribution and the target distribution during the training process. The training objective of the generator is to minimize the Wasserstein distance between the generated samples and the normal sample distribution, while the training objective of the discriminator is to try to maximize this distance, constituting a game process.
[0089] The adversarial examples generated through the above mechanism approach normal samples in the overall distribution but retain some key attack features, thus being able to effectively break through the detection capabilities of static rules and shallow models. The training convergence criterion is that the Wasserstein distance is less than 0.05. Multiple types of adversarial examples are generated, including DNS tunnel attack samples (such as super-long subdomains: b64.a1b2.c3d4.[...].exfil.com (length > 63 characters), frequent simulation of TXT record queries at 600 times per minute) and DDoS samples (such as forged source IPs: generating random IP addresses conforming to the / 24 network segment distribution, and large-scale concurrent queries: burst traffic reaching 10,000 QPS). After mixing the generated adversarial attack samples with normal samples, they are input into the multi-modal detection engine module.
[0090] The multi-modal detection engine module is used to perform fusion recognition and attack classification judgment on the input mixed samples. The model structure is as Figure 2 shown, specifically divided into an input layer, a feature fusion layer, and an output layer, and equipped with a dynamic threshold adjustment mechanism; it can achieve spatio-temporal feature fusion and dynamic threshold adjustment to achieve high-precision attack classification.
[0091] The input layer receives a 48-dimensional feature vector of the sample (normalized to the [0, 1] interval) and inputs it into the feature fusion layer. The key context information is extracted through a four-head attention mechanism (4 heads, Key dimension = 64), and at the same time, the spatial structure feature F extracted by the CNN is combined t with the time series behavior feature F extracted by the GRU s , and spatio-temporal feature cross-fusion is performed to obtain the fusion feature F fusion ; the fusion expression is as follows:
[0092]
[0093] This fusion feature is then input into the Softmax classifier to calculate the attack probability score P attackThe attack probability ranges from [0, 1], representing the confidence level that the system believes the sample is an attack. The calculation formula is as follows:
[0094]
[0095] where z i represents the original score (logit) of the sample on the i-th class (such as the attack class), and z j represents the model output scores of all candidate classes j. The Softmax function transforms the model output into values that satisfy the probability distribution constraint through exponential mapping and normalization operations, ensuring that the sum of the probabilities of all classification outputs is 1.
[0096] In the real-time detection process, this module processes the sliding window data (window length 300 seconds) every 5 seconds and dynamically adjusts the classification threshold according to the KL divergence. The adjustment formula is:
[0097] T = T0 + α·D KL (P||Q)
[0098] T0 is the initially set threshold, α is the adjustment factor, and D KL (P||Q) represents the divergence between the current model output distribution and the normal distribution.
[0099] If the attack probability P output by the model attack ≥ T and conforms to the behavior pattern defined in the rule engine (such as the NXDOMAIN response ratio being greater than 50%), then the mixed sample is determined to be abnormal; and the detection result is synchronously transmitted to the dynamic defense and feedback system.
[0100] The dynamic defense and feedback system executes automated defense strategies according to the detection results and realizes the self-learning and continuous optimization of the detection system. After an attack is confirmed, BGP FlowSpec rules are automatically issued to block malicious domain name resolution; at the same time, adversarial attack samples are injected into the training set through an online learning mechanism, and the detection model is incrementally updated weekly.
[0101] This module includes a defense execution unit and a model optimization mechanism: In terms of defense execution, the system supports IP reputation library management, and IPs that trigger more than 3 alarms within 10 minutes will be marked as malicious; access filtering policies based on the target IP or UDP port 53 can be quickly issued through the BGP FlowSpec protocol; for example:
[0102] deny udp any any eq 53to 192.168.1.0 / 24;
[0103] Combine with the DNS Sinkhole mechanism to resolve suspicious domain names to a black hole IP (such as 10.0.0.1), thus achieving fast interception. In terms of model optimization, the system supports an online learning mechanism, and newly detected attack samples can be injected into the training dataset in real time within 1 second; at the same time, the system automatically generates 10,000 new adversarial attack samples per week for incremental training to further improve the model's recognition ability under new attacks. In addition, this module also introduces a feature importance adjustment mechanism based on SHAP values to dynamically calculate feature weights, thereby enhancing the interpretability and continuous optimization of the model's discrimination logic. The feature weight calculation formula is:
[0104]
[0105] where f i represents the i-th input feature, and SHAP(f i ) is the SHAP score corresponding to this feature in the prediction of the current sample. ∑ j SHAP(f j ) represents the sum of the SHAP values of all input features under this sample, and w i represents the relative importance weight of feature f i in the model's discrimination logic. Through the above mechanism, the system can dynamically adjust the feature sensitivity during the model operation, improving the interpretability, controllability, and training stability of the model under different traffic scenarios.
[0106] Example 1:
[0107] In an actual application scenario, a large enterprise network deploys the DNS attack detection and defense system based on an adversarial network described in the present invention, and the process is as Figure 3 shown; install this system in front of the enterprise export DNS server, and obtain the DNS request traffic in the network in real time through mirroring as the data source for training and detection.
[0108] In the initial stage when the system is launched, the platform is in the pre-training stage. The data acquisition module uses the FPGA hardware probe deployed at the network boundary to continuously capture the DNS request behaviors initiated by each terminal, and uses the DPDK protocol parsing engine to decode the protocol and extract fields from the traffic, refining a 48-dimensional structured feature vector including query frequency, response delay, domain name structure, TTL change, encoding features, etc. After the samples are preliminarily labeled and denoised, they are input into the adversarial sample generation module. The generator based on the generative adversarial network (GAN) constructs adversarial samples and optimizes the training process through the Wasserstein distance to enhance the model's recognition ability for complex camouflaged attack samples. The above normal and adversarial samples are both used to train the multi-modal detection engine to construct a preliminary detection model. After the system completes offline training and model convergence, it switches to the online detection mode.
[0109] In a certain attack event, the attacker attempts to bypass the firewall restrictions and achieve data exfiltration by implanting malicious programs into enterprise office terminals and using DNS tunneling technology.
[0110] The specific behavior is as follows: The terminal continuously sends a large number of TXT queries to a disguised domain name such as b64.a1b2.c3d4.exfil.com per second, with a request frequency as high as 600 times per second. The domain name structure is complex (the depth of subdomains exceeds 5 levels), the requests contain a large amount of Base64-encoded content, and the TTL value is fixed while the response message is empty.
[0111] This system first captures the DNS behavior of this terminal in real time through the FPGA hardware probe deployed at the network boundary. After extracting the corresponding structured features, it directly injects them into the trained multi-modal detection engine module. The multi-modal detection engine uses CNN, GRU, and attention mechanisms to extract the spatial structure features and time series behaviors of the input samples, combines with the Softmax classifier to output an attack probability of 0.91, and triggers subsequent rule matching. This behavior simultaneously meets multiple preset attack feature conditions such as "TXT record ratio greater than 50%, suspected Base64 encoding, high-frequency query, deep subdomain level", and is determined by the system as a DNS tunneling attack. The detection result is fed back to the dynamic defense and feedback module in real time by the system, triggering the following automatic defense measures: blocking the UDP port 53 access permission of the terminal IP (such as 192.168.1.104) through the BGP FlowSpec protocol; at the same time, redirecting the target malicious domain name *.exfil.com to the black hole address 10.0.0.1 to interrupt the data channel; this attack event is recorded in the system log and automatically archived to the visualization management platform.
[0112] In addition, this attack sample and its generated adversarial samples are returned to the training set in real time to support the incremental learning of the model. In subsequent training rounds, through the online learning mechanism and the SHAP value update process, the weights of features such as "Base64_Pattern" and "Subdomain_Depth" in the model are dynamically increased to further enhance the detection ability for similar attacks.
[0113] Through the above closed-loop process, the system of the present invention successfully identifies and blocks a typical DNS tunneling attack, realizing the full-process processing from abnormal behavior recognition, adversarial sample construction, multi-modal detection, to automatic defense and model update, reflecting the intelligence, self-adaptability, and real-time protection ability of the system in the actual network environment.
[0114] The present invention uses GAN to generate attack samples to improve the adaptability of the detection system to unknown attacks; introduces a real-time detection and feedback mechanism to ensure that the system can respond to new attacks in a timely manner. It also combines various traffic feature analyses to reduce the false alarm rate and improve the detection accuracy. It can effectively improve the detection rate, reduce false alarms, and enhance the defense ability against new attacks in various DNS attack scenarios.
Claims
1. A DNS attack detection and defense system based on an adversarial network, characterized in that, It is composed of a data collection and feature engineering module, an adversarial sample generation module, a multi-modal detection engine module, and a dynamic defense and feedback system; During the operation of the system, first, the data collection and feature engineering module captures DNS traffic in real time and parses the protocol, extracts multi-dimensional behavioral features, and combines the rule engine, threat intelligence, historical samples, and trap results to perform preliminary annotation on the DNS traffic, generates structured feature data, and divides it into two categories: normal samples and attack samples; Subsequently, the attack samples and normal samples are respectively sent into the adversarial sample generation module. By introducing a generative adversarial network GAN optimized by the Wasserstein distance, strongly camouflaged adversarial attack samples are constructed for the attack samples; Taking the normal samples and adversarial attack samples as mixed samples, they are input into the multi-modal detection engine module. Relying on the attention mechanism, convolutional neural network, and time series model, the input mixed samples are fused and analyzed, and the attack probability and classification results are output; Finally, the classification results are transmitted to the dynamic defense and feedback system, which automatically executes the BGP blocking or domain name redirection response strategy. At the same time, the mixed samples and detection results are fed back into the database for system update and policy optimization, forming a closed-loop adaptive defense mechanism.
2. The DNS attack detection and defense system based on the adversarial network according to claim 1, characterized in that, The data collection and feature engineering module is deployed at the DNS server entrance or network boundary, and is used to collect DNS traffic data in real time and perform structured feature extraction and preliminary annotation on it.
3. The DNS attack detection and defense system based on the adversarial network according to claim 1 or 2, characterized in that, The data collection and feature engineering module includes a data collection layer, a protocol parsing unit, a three-level feature extraction unit, and a mixed annotation unit; Among them, the data collection layer is based on an FPGA hardware probe to achieve 10Gbps line speed packet capture; The protocol parsing unit uses the DPDK high-performance framework to parse DNS, DoT, and DoH protocols, supports the decryption of encrypted traffic, and depends on a pre-set CA certificate library; The three-level feature extraction unit extracts three types of features according to dimensions: basic features, advanced features, and payload features; The mixed annotation unit is based on the historical sample library and threat intelligence source, provides IOC context information, and at the same time combines the attack traffic captured by the camouflaged DNS server to perform preliminary annotation to obtain a 48-dimensional structured feature vector, which is divided into normal samples and attack samples.
4. The DNS attack detection and defense system based on the adversarial network according to claim 1, characterized in that The adversarial sample generation module is designed based on the generative adversarial network GAN, includes a generator and a discriminator, and uses the Wasserstein distance as the optimization target to generate adversarial attack samples.
5. The DNS attack detection and defense system based on the adversarial network according to claim 4, characterized in that, The generator takes attack samples as input and learns the optimal projection mapping relationship from the attack sample distribution P s to the normal sample distribution P t so as to generate adversarial samples with a structure approximately normal but retaining attack features; Generator A θ (·) has a training objective of minimizing the Wasserstein distance between the distribution of adversarial attack samples and that of normal samples, and its loss function is defined as: where x ∼ P s denotes the input sampled from the attack sample distribution, A θ (x) represents the projected output of the generator for the attack sample, and the Wasserstein distance between the generated adversarial attack sample and the normal sample distribution is minimized by adjusting θ. D(·) is a real-valued function output by the discriminator.
6. The DNS attack detection and defense system based on the adversarial network according to claim 4, characterized in that The training objective of the discriminator is to maximize the Wasserstein distance between the normal sample distribution P t and the adversarial attack sample distribution, so as to evaluate the "disguise" of adversarial samples. Its loss function is defined as: where x′ ∼ P t denotes the input sampled from the normal sample distribution, D w (·) is a discriminant function satisfying the 1-Lipschitz constraint, and the output value is used to measure the distance between the adversarial attack sample and the normal sample. By adjusting w, the Wasserstein distance between the normal sample distribution and the adversarial attack sample distribution is maximized; A(x) represents the output result of the generator for the attack sample.
7. The DNS attack detection and defense system based on adversarial network according to claim 4 or 5 or 6, characterized in that, The parameters w and θ of the discriminator and generator are independently optimized. In each round of training, first, the parameters θ of the generator are fixed, and the parameters w of the discriminator are updated by multi-round gradient ascent to maximize the Wasserstein distance; then, the parameters w of the discriminator are fixed, and the generator loss function is minimized to guide the generator to learn a mapping strategy close to the normal sample distribution; When the Wasserstein distance is less than 0.05, multiple types of adversarial attack samples including DNS tunnel attack samples and DDoS samples are generated.
8. The DNS attack detection and defense system based on the adversarial network according to claim 1, characterized in that, The multi-modal detection engine module performs fusion recognition and attack classification judgment on the input mixed samples; specifically, it is divided into an input layer, a feature fusion layer, and an output layer, and is equipped with a dynamic threshold adjustment mechanism; The input layer receives the 48-dimensional normalized feature vector of the sample and inputs it into the feature fusion layer. The key context information is extracted through the four-head attention mechanism, and at the same time, the spatial structure feature F extracted by the CNN is combined t with the time series behavior feature F extracted by the GRU s , and spatio-temporal feature cross-fusion is performed to obtain the fused feature F fusion ; The formula for the fusion process is as follows: This fused feature is then input into a Softmax classifier to calculate the attack probability score P attack ; the value range of this attack probability is [0, 1], representing the confidence that the sample is an attack; In the real-time detection process, the classification threshold is dynamically adjusted for the sliding window data according to the KL divergence, and the adjustment formula is: T = T0 + α·D KL (P||Q) T0 is the initial set threshold, α is the adjustment factor, D KL (P||Q) represents the divergence between the current model output distribution and the normal distribution; If the attack probability P attack ≥ T, then it is determined that the mixed sample is abnormal.
9. The DNS attack detection and defense system based on the adversarial network according to claim 1, characterized in that, The dynamic defense and feedback system executes automated defense strategies according to the detection results, and realizes the self-learning and continuous optimization of the detection system; after an attack is confirmed, the BGP FlowSpec rule is automatically issued to block malicious domain name resolution; at the same time, adversarial attack samples are injected into the training set through the online learning mechanism, and the detection model is incrementally updated weekly.
Citation Information
Patent Citations
Network attack traffic generation method based on auxiliary classification type generative adversarial network
CN113158390A
Attack resisting system for deep intrusion detection
CN113392932A
DDoS attack distinguishing method and system based on CVAE-WGAN-GP
CN118631562A
Artificial intelligence enhanced distributed denial of service attack defense method and system
CN119865343A
Method for training a generative adversarial network (GAN), generative adversarial network, computer program, machine-readable memory medium, and device
US20200372297A1
Cited By
Multi-modal large model confrontation safety detection method and system
CN120639526A
A multimodal large model confrontation security detection method and system
CN120639526B
Deep firewall intrusion detection method based on behavior characteristic and packet header field joint modeling
CN120956459A
LLM routing system-oriented rerouting attack detection method
CN121598088A