DNS tunnel detection method

Through the anomaly detection method of multi-level aggregation and deep unsupervised training on DNS traffic, the problem of difficulty in detecting distributed DNS tunnels in traditional technologies is solved, and efficient detection of low-throughput and hidden DNS tunnels is achieved.

CN119996065AActive Publication Date: 2025-05-13HARBIN INST OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510401906.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-13
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The prior art is difficult to detect hidden distributed DNS tunnels, especially low-throughput DNS tunnels, and traditional detection methods have difficulties in aggregating scattered data packets.

Method used

By obtaining the training traffic data set, traffic preprocessing and packet-level metadata extraction are performed, first-order aggregation and detection are performed based on IP-Domain, and second-order aggregation is performed based on IP and Domain. In-depth unsupervised training is used to form an abnormality detection model at the session level, domain name level and communication level, and comprehensive detection is performed through a nonlinear voting machine.

Benefits of technology

Effective detection of conventional DNS tunnels, low-throughput DNS tunnels and distributed DNS tunnels with good concealment are achieved, which improves the balance between detection performance and detection time, and enhances the comprehensiveness of detection range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996065A_ABST
    Figure CN119996065A_ABST
Patent Text Reader

Abstract

The invention discloses a DNS tunnel detection method, and belongs to the technical field of secure communication. The problem that in the prior art, a traditional DNS tunnel detection method is difficult to detect a low-throughput DNS tunnel and a distributed DNS tunnel is solved. The method comprises the following steps: aggregating data packets by using a set first aggregation key, carrying out preliminary detection and scoring by using an auto-encoder, and detecting a conventional DNS tunnel; aggregating the first-order aggregated metadata again by using a set second aggregation key and a third aggregation key, and detecting a low throughput DNS tunnel and a distributed DNS tunnel from a communication object and a domain name dimension; and the scores of the three auto-encoders are sent to an auto-encoder serving as a nonlinear voting mechanism, and finally, the nonlinear voting mechanism makes a final judgment according to the scores of the data under the three aggregation keys to obtain a detection result. According to the method, the comprehensiveness of DNS tunnel detection is effectively improved, and the method can be applied to detection of distributed DNS tunnels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a DNS tunnel detection method and belongs to the technical field of secure communications. Background Art

[0002] Traffic control at the edge of the network is crucial to ensure the security of information systems, and it can effectively suppress malicious traffic attacks. However, some attackers use tunneling technology to disguise malicious traffic as legitimate traffic to bypass security policies at the edge of the network and conduct illegal activities. Tunneling is a data encapsulation technology that can encapsulate the original data packet in another data packet of a different protocol for transmission. Traditional tunnels are mainly based on transport layer or network layer protocols. Currently, tunneling technology is gradually beginning to be implemented using application layer protocols, such as Hypertext Transfer Protocol (HTTP), Domain Name System (DNS), and Secure Shell Protocol (SSH).

[0003] The Domain Name System (DNS) plays an important role in the operation of today's Internet. It provides a two-way mapping service between IP addresses and domain names. Since it was not originally designed for data transmission, traditional network security software or equipment often allow DNS services to use User Datagram Protocol (UDP) port 53 by default. In addition, hosts usually unconditionally trust the response information of DNS servers, so DNS traffic can often spread unimpeded at the edge of the network without strict security policies to constrain it. Attackers exploit the above vulnerabilities and use a covert communication technology called DNS tunneling for data transmission. Figure 2 Attackers usually encode the data to be leaked and put it into a subdomain, and then establish communication with the controlled domain name server through normal DNS query.

[0004] DNS tunnels for data transmission were originally designed to bypass the edge of the network to obtain free Internet access. Due to the effectiveness of bypassing network security mechanisms, more and more services tend to use DNS tunnels. For example, a company proposed a method to distribute updated malicious code signatures through DNS tunnels, the purpose of which is to provide signature update services for antivirus clients. DNS-based remote control malware is considered the most dangerous network attack. In addition, DNS tunnels can also be used for remote command and control information transmission of malware, data leakage, etc. Although DNS tunnels only provide low bandwidth for data transmission, attackers can still steal data through the tunnel or maintain communication with malware. Therefore, DNS tunnel detection technology is urgently needed in academic and industrial fields.

[0005] DNS tunnel detection technology is of great significance for ensuring the security of information systems. DNS tunnel detection technology can effectively identify DNS tunnel traffic mixed in normal traffic and promptly control channels where data leakage is occurring. For example, when facing double ransomware using DNS tunnels, the DNS tunnel function can effectively identify DNS tunnel traffic used for data leakage and promptly notify security personnel to stop the loss.

[0006] In the prior art, DNS tunnel detection can be divided into load-based detection methods and flow-based detection methods according to different detection subjects; the detection object of the load-based detection method is often a single data packet, while the detection object of the flow-based detection method is a session composed of data packets within a period of time. Load-based detection is usually better than flow-based detection in terms of time cost, but the detection capability of low-throughput DNS tunnel data packets similar to normal data packets is poor; therefore, if you want to detect low-throughput DNS tunnels, you usually need to adopt a flow-based detection method, and the first problem faced by flow-based detection research is how to aggregate scattered data packets. In existing research, aggregation keys exist in the form of (src_ip, src_port, protocol, dst_port, dst_ip), (src_ip, protocol, dst_ip), (domain), etc. Researchers usually conduct research based on the assumption that the DNS tunnel server has only one IP and only uses one domain name, and usually only use one aggregation key for aggregation in the research. However, in reality, attackers can effectively reduce the amount of information on a single domain name or address by interleaving multiple IPs and domain names, thereby reducing the risk of being detected. This DNS tunnel technology that uses multiple IPs and domain names to disperse information to different channels is called a distributed DNS tunnel; distributed DNS tunnels can be well combined with botnets. Through the alternating use of different domain names and DNS resolution load balancing technology, DNS query requests are distributed to different zombie hosts, and then forwarded by the zombie hosts to the controlled servers, which aggregate and reassemble the messages to obtain leaked data. Therefore, although DNS tunnel detection technology plays a key role in improving the security of information systems, it still needs to be continuously optimized and developed to cope with increasingly complex attack methods and challenges in practical applications. How to detect low-throughput DNS tunnels and distributed DNS tunnels that are more hidden than conventional low-throughput DNS tunnels has become an issue that cannot be ignored.

[0007] In summary, a DNS tunnel detection method is needed. Summary of the invention

[0008] A brief overview of the present invention is provided below in order to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to a more detailed description discussed later.

[0009] In view of this, in order to solve the problem that the traditional DNS tunnel detection method in the prior art is difficult to detect the hidden distributed DNS tunnel, the present invention provides a DNS tunnel detection method.

[0010] The technical solution is as follows: A DNS tunnel detection method includes the following steps:

[0011] S1. Preprocess the traffic by acquiring a training traffic data set, filtering the traffic data set, extracting and storing packet-level metadata, and obtaining a data packet-level database;

[0012] S2. According to the packet-level database, first-order aggregation and detection based on IP-Domain are performed to obtain a session-level metadata database, a session-level feature library, and a trained session-level anomaly detection model;

[0013] S3. According to the session-level metadata database, based on the second-order aggregation and detection of IP and Domain, the trained domain-level anomaly detection model and communication-level anomaly detection model are obtained;

[0014] S4. Based on the trained session-level anomaly detection model, domain-level anomaly detection model, and communication-level anomaly detection model, a nonlinear voter is used for comprehensive detection, and a high-level alarm is generated if a DNS tunnel is detected.

[0015] Furthermore, the step S1 includes the following steps:

[0016] S11. Obtain training traffic data set;

[0017] Specifically: the process of obtaining the training traffic data set includes the normal traffic acquisition stage of virtual machine A as the client that initiates the DNS query, the tunnel traffic acquisition stage of the client that deploys the DNS tunnel tool in virtual machine A, and the deployment traffic collection script stage of virtual machine B as the local DNS server;

[0018] S12. Filter the DNS data packets in the traffic data set to obtain a traffic data set after traffic filtering;

[0019] S13. Read the traffic data set after traffic filtering, extract and store packet-level metadata, and obtain a data packet-level database;

[0020] In S11, the normal traffic acquisition stage includes the following steps:

[0021] S1111. Set program related parameters;

[0022] S1112. Select browser driver;

[0023] S1113. Select a search engine as a basis for subsequent searches;

[0024] S1114. Select a keyword word library as a search seed;

[0025] S1115. Using the interruption and continuation mechanism, determine whether there is an archive file in the keyword thesaurus, the archive file includes a total set of keywords and a complete set of keywords, respectively, the keyword archive records all the keywords used in each round of crawling, and the keyword complete archive records the keywords that have been used;

[0026] S1116. Randomly extract keywords. If there are no archived files left, this operation is regarded as a new round of crawling, and a number of data are randomly extracted from keyword word library files in various fields;

[0027] S1117. Archiving keywords, that is, storing the keywords extracted in this round of operation into the total keyword set;

[0028] S1118. Read the archived files for keywords. If there are archived files left, this operation is regarded as a continuation of the previous round of crawling. The total set of keywords and the complete set of keywords are read, and the difference between them is taken to obtain the keywords that have not been used yet;

[0029] S1119. Perform a traversal search on the keywords used in this round, and treat the keywords as seed keywords. After the keyword search, store them in the keyword completion set;

[0030] S11110. Take anti-crawler measures to avoid them;

[0031] S11111. Determine whether the crawler is countered, that is, determine whether the anti-crawler measures of the search engine are effective by detecting the structure of the web page obtained by the search;

[0032] S11112. Perform recommendation jump avoidance, that is, when the current keyword can no longer produce search results, jump through the search engine's related recommendations;

[0033] S11113. Traverse the search results;

[0034] S11114. Jump to search page, that is, select sequential page jump or random page jump according to the parameters set by the user;

[0035] S11115. When all the keywords used in this round of operation are traversed, the archive files recording the total set of keywords and the completed set of keywords are deleted;

[0036] In S11, the tunnel traffic acquisition stage includes the following steps:

[0037] S1121. Perform parameter configuration for the DNS tunnel tool;

[0038] S1122. Select the file for DNS tunnel testing;

[0039] S1123 sends a DNS tunnel packet to build a DNS tunnel through the DNS tunnel server;

[0040] S1124. Read the response returned by the DNS tunnel tool server;

[0041] In S11, the stage of deploying the traffic collection script includes the following steps:

[0042] S1131. Select monitoring protocol;

[0043] S1132. Set BPF filtering conditions to quickly filter the underlying protocols of the transport layer, network layer, data link layer, and physical layer;

[0044] S1133. Sniff all packets of the selected network port, that is, use Scapy to monitor the traffic of the specified network adapter;

[0045] S1134. Rapidly filter the traffic according to the BPF filtering conditions;

[0046] S1135. Perform deep filtering at the application layer for traffic that meets the requirements of the underlying protocol;

[0047] S1136. Save the deep filtered traffic by adding append;

[0048] The S12 includes the following steps:

[0049] S121. Read the traffic data set, that is, read the saved DNS tunnel original traffic file;

[0050] S122. Traverse the DNS data packets in the traffic data set;

[0051] S122. If the primary domain name queried in the DNS data packet is the same as the domain name set for DNS tunnel construction, it is saved in a new traffic data set in an incremental manner;

[0052] S123. If the primary domain name queried in the DNS packet is different from the domain name set for DNS tunnel construction, it is discarded;

[0053] S124. Obtaining the traffic data set after traffic filtering;

[0054] The S13 includes the following steps:

[0055] S131 reads the traffic data set after traffic filtering, that is, reads the saved DNS tunnel original traffic file after traffic filtering;

[0056] S132. Initialize the database based on the traffic data set information after traffic filtering and the selected packet-level metadata, so that each traffic data set after traffic filtering corresponds to a data table;

[0057] S133. Traverse the DNS data packets in the traffic data set after traffic filtering;

[0058] S134. According to the set packet-level metadata, extract the packet-level metadata in the DNS data packet in the traffic data set after traffic filtering;

[0059] S135. Save the extracted packet-level metadata into the corresponding data packet-level metadata table of the database storing metadata, and obtain a data packet-level database.

[0060] Furthermore, the step S2 includes the following steps:

[0061] S21. According to the packet-level database, a session-level aggregation key represented by src_ip, domain, and dst_ip is generated to obtain a session-level metadata library and a session-level feature library, and complete the first-order aggregation based on IP-Domain;

[0062] S22. According to the session-level feature library, tensor conversion is performed, the obtained session feature data tensor is standardized, and the session-level autoencoder is trained in deep unsupervised manner using the standardized session-level feature data to obtain a trained session-level anomaly detection model, and complete the first-order detection based on IP-Domain;

[0063] The S21 includes the following steps:

[0064] S211. Read the data packet level metadata table in the data packet level database;

[0065] S212. Create a session-level database based on the packet-level metadata table information of the packet-level database, the selected session-level metadata database and the session-level features, where each data table in the packet-level database corresponds to a session-level metadata table and a session-level feature table;

[0066] S213 traverses each packet-level metadata table of the packet-level database;

[0067] S214. Read the packet-level metadata in the data packet-level metadata table;

[0068] S215. Extract the primary domain name from the domain name field in the package-level metadata;

[0069] S216. Generate a session-level aggregation key, that is, extract the source IP represented as src_ip and the destination IP represented as dst_ip in the packet-level metadata, and combine them with the primary domain name domain extracted in step S215 to obtain a session-level aggregation key represented as src_ip, domain, dst_ip;

[0070] S217. Create session-level metadata entries, that is, if the currently processed aggregate key does not have a related record in the session-level metadata dictionary, it is regarded as a new session, and related entries are created in the session-level metadata dictionary and the session-level time dictionary, and the session-level metadata entries are integrated into session-level metadata;

[0071] S218. Update the session-level metadata, that is, if the currently processed aggregation key has a related record in the session-level metadata dictionary, update the session-level metadata dictionary and the session-level time dictionary;

[0072] S219. Obtaining session-level features by performing aggregation mean calculation on session-level metadata;

[0073] S2110. The aggregation key and session-level metadata are saved in the session-level metadata table to obtain a session-level metadata database, and the aggregation key and session-level feature data are saved in the session-level feature table to obtain a session-level feature database;

[0074] The S22 includes the following steps:

[0075] S221. Read the session-level feature data table in the session-level feature library;

[0076] S222. Traverse each session-level feature data table in the session-level feature library;

[0077] S223 reads the characteristic data in each session-level characteristic data table. To avoid memory explosion, read the set N characteristic data each time;

[0078] S224. Perform tensor conversion to convert the read feature data from a tuple type into a tensor type suitable for the deep learning framework Pytorch;

[0079] S225. Process the session feature data tensor using the minimum-maximum normalization method to obtain the normalized session feature data, and map each type of metadata value to the interval [0, 1];

[0080] S226. Use the normalized session-level feature data as input data of the session-level autoencoder to perform deep unsupervised training;

[0081] S227. Save the trained session-level anomaly detection model as a .pth file.

[0082] Furthermore, the step S3 includes the following steps:

[0083] S31. According to the session-level metadata library, a communication-level aggregation key represented by src_ip and dst_ip is generated to obtain a communication-level feature library to complete communication-level feature aggregation;

[0084] S32. Generate a domain-level aggregation key based on the session-level metadata library to obtain a domain-level feature library to complete domain-level feature aggregation;

[0085] S33. Use the standardized domain-level features and communication-level feature data to perform deep unsupervised training on the domain-level autoencoder and the communication-level autoencoder, respectively, to obtain a trained domain-level anomaly detection model and a communication-level anomaly detection model, and complete the second-order detection;

[0086] The S31 includes the following steps:

[0087] S311. Read all data tables in the session-level metadata database;

[0088] S312. Create a communication-level feature database based on the session-level metadata table information and the selected communication-level features in the data session-level metadata database, where each session-level metadata table corresponds to a communication-level feature table;

[0089] S313. Traverse each session-level metadata table in the session-level metadata database;

[0090] S314. Read the session-level metadata in the session-level metadata table;

[0091] S315. Generate a communication-level aggregation key, extract the source IP represented as src_ip and the destination IP represented as dst_ip in the session-level metadata, and combine them to obtain a communication-level aggregation key represented as src_ip, dst_ip;

[0092] S316. Create a communication-level metadata entry, that is, if the currently processed aggregation key does not have a related record in the communication-level metadata dictionary, it is regarded as a new session, and a related entry is created in the communication-level metadata dictionary to be integrated into the communication-level metadata;

[0093] S317. Update the communication-level metadata, that is, if the currently processed aggregation key has a related record in the communication-level metadata dictionary, update the communication-level metadata dictionary;

[0094] S318. Obtain communication-level features by calculating the communication-level metadata;

[0095] S319. The communication-level aggregation key and the communication-level feature are saved in the communication-level feature table to obtain a communication-level feature library;

[0096] The S32 includes the following steps:

[0097] S321. Read all session-level metadata tables in the session-level metadata database;

[0098] S322. Based on the session-level metadata table information in the data session-level metadata database and the selected domain-level features, a domain-level feature database is created, where each session-level metadata table corresponds to a domain-level feature table;

[0099] S323. Traverse each session-level metadata table in the session-level metadata database;

[0100] S324. Read the session-level metadata in the session-level metadata table;

[0101] S325. Generate a domain-level aggregation key, that is, extract the primary domain name domain in the session-level metadata as the domain-level aggregation key;

[0102] S326. Create a domain-level metadata entry. If the currently processed aggregation key has no related record in the domain-level metadata dictionary, it is regarded as a new session, and a related entry is created in the domain-level metadata dictionary to be integrated into the domain-level metadata;

[0103] S327. Update the domain-level metadata, that is, if the currently processed aggregation key has a related record in the domain-level metadata dictionary, update the domain-level metadata dictionary;

[0104] S328. Obtain domain-level features by calculating domain-level metadata;

[0105] S329. Save the domain-level aggregation key and domain-level features into the domain-level feature table to obtain a domain-level feature library.

[0106] Furthermore, the step S4 includes the following steps:

[0107] S41. Read the reconstruction errors of the trained session-level anomaly detection model, domain-level anomaly detection model, and communication-level anomaly detection model;

[0108] S42. Calculate the reconstruction errors output by the three anomaly detection models through a voting machine to obtain anomaly scores;

[0109] S43. If the abnormal score output by the voting machine is greater than the set threshold, an alarm is issued;

[0110] S44. If the anomaly score output by the voter is less than the set threshold, the current traffic is not processed.

[0111] The beneficial effects of the present invention are as follows: the present invention proposes a DNS tunnel detection method, which achieves a good balance between detection performance and detection time through two aggregations and detections, and the detection range is more comprehensive; the present invention uses session-level aggregation keys to aggregate DNS data packets in the first stage, and uses autoencoders for preliminary detection and scoring. This process is carried out in real time, and the detection target is a conventional DNS tunnel; the present invention then uses communication-level aggregation keys and domain-level aggregation keys to aggregate the metadata of the first-order aggregation again in the second stage, and uses autoencoders for scoring respectively. Then we send the scores of the three autoencoders to an autoencoder as a nonlinear voting mechanism, and finally the nonlinear voting mechanism makes a final judgment based on the scores of the data under the three aggregation keys; the entire processing process of the second stage is near real-time, and its detection target is a low-throughput DNS tunnel and a distributed DNS tunnel with good concealment; based on the anomaly detection principle, the present invention uses unsupervised learning to train and test three autoencoders to achieve detection of different types of DNS tunnels. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0113] Figure 1 A schematic diagram of a flow chart of a DNS tunnel detection method;

[0114] Figure 2 This is a schematic diagram of the principle of DNS tunnel;

[0115] Figure 3 A schematic diagram of an embodiment of a DNS tunnel detection method;

[0116] Figure 4 Schematic diagram of the deployment method for traffic preprocessing;

[0117] Figure 5 This is a schematic diagram of the workflow of the Selenium crawler;

[0118] Figure 6 This is a schematic diagram of the DNS tunnel workflow;

[0119] Figure 7 It is a schematic diagram of the workflow of virtual machine B;

[0120] Figure 8 This is a flow chart of traffic filtering;

[0121] Fig. 9 A schematic diagram of the process of extracting and saving package-level metadata;

[0122] Fig.10 A schematic diagram of the process of aggregating and extracting session-level metadata based on session-level aggregation keys;

[0123] Fig.11 This is a diagram of the training process of the session-level anomaly detection model;

[0124] Fig.12 A schematic diagram of a flow chart of communication-level session feature aggregation calculation based on a communication-level aggregation key;

[0125] Fig.13 A flowchart of domain-level session feature aggregation calculation based on domain-level aggregation keys;

[0126] Fig.14 It is a schematic diagram of the principle of comprehensive detection and alarm;

[0127] Fig.15 Schematic diagram of the workflow of comprehensive detection and alarm. DETAILED DESCRIPTION

[0128] In order to make the technical solutions and advantages of the embodiments of the present invention more clearly understood, the exemplary embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0129] refer to Figure 1-Figure 15 The present embodiment is described in detail, a DNS tunnel detection method specifically includes the following steps:

[0130] S1. Preprocess the traffic by acquiring a training traffic data set, filtering the traffic data set, extracting and storing packet-level metadata, and obtaining a data packet-level database;

[0131] S2. According to the packet-level database, first-order aggregation and detection based on IP-Domain are performed to obtain a session-level metadata database, a session-level feature library, and a trained session-level anomaly detection model;

[0132] S3. According to the session-level metadata database, based on the second-order aggregation and detection of IP and Domain, the trained domain-level anomaly detection model and communication-level anomaly detection model are obtained;

[0133] S4. Based on the trained session-level anomaly detection model, domain-level anomaly detection model, and communication-level anomaly detection model, a nonlinear voter is used for comprehensive detection. If a DNS tunnel is detected (the anomaly score generated by the nonlinear voter is higher than the threshold), an advanced alarm is generated.

[0134] Specifically, refer to Figure 3 The present invention first acquires and filters the traffic in the traffic preprocessing stage, and obtains the packet-level metadata database after extraction; then performs first-order extraction, and aggregates and calculates the packet-level metadata with (src_ip, domain, dst_ip) as the aggregation key to obtain the IP-Domain metadata database; next, the IP-Domain metadata database is used for primary detection and early warning on the one hand, and for data aggregation under the (src_ip, dst_ip) and (domain) aggregation keys on the other hand; finally, the detection results of the metadata under the three aggregation keys are summarized to the voting organization, which conducts a comprehensive evaluation and makes corresponding early warnings.

[0135] Furthermore, the step S1 includes the following steps:

[0136] S11. Obtain training traffic data set;

[0137] Specifically: the process of obtaining the training traffic data set includes the normal traffic acquisition stage of virtual machine A as the client that initiates the DNS query, the tunnel traffic acquisition stage of the client that deploys the DNS tunnel tool in virtual machine A, and the deployment traffic collection script stage of virtual machine B as the local DNS server;

[0138] S12. Filter the DNS data packets in the traffic data set to obtain a traffic data set after traffic filtering;

[0139] S13. Read the traffic data set after traffic filtering, extract and store packet-level metadata;

[0140] In S11, reference Figure 5 ,The normal traffic acquisition phase includes the following steps:

[0141] S1111. Set program-related parameters, such as the number of search pages, waiting time after recommended jump, waiting time after turning pages, web browsing time, sleep time before clicking, and the number of keywords extracted from each vocabulary;

[0142] S1112. Select a suitable browser driver, such as ChromeDriver;

[0143] S1113. Select an appropriate search engine as the basis for subsequent searches, such as Baidu, Sohu, Bing, etc.;

[0144] S1114. Select an appropriate keyword word library as the search seed. To ensure that the collected traffic is comprehensive, it is recommended to use a word library covering various fields. Keywords in a single field may cause websites related to the field to appear repeatedly, reducing the diversity of traffic;

[0145] S1115. Adopt the interrupt-resume mechanism to determine whether there is an archive file in the keyword thesaurus. The archive file includes a total keyword set and a complete keyword set. The keyword archive records all keywords used in each round of crawling, and the keyword complete archive records the keywords that have been used. In order to cope with complex and changeable network conditions and ensure the overall integrity of the data generated by each program run, the interrupt-resume mechanism is adopted to determine whether there is an archive file remaining and whether the program crashes, and then decide whether this run is a new round of crawling or a continuation of the previous round of crawling;

[0146] S1116. Randomly extract keywords. If there are no archived files left, this operation is regarded as a new round of crawling, and a number of data are randomly extracted from keyword word library files in various fields;

[0147] S1117. Archiving keywords, that is, storing the keywords extracted in this round of operation into the total keyword set;

[0148] S1118. Read the archived files for keywords. If there are archived files left, this operation is regarded as a continuation of the previous round of crawling. The total set of keywords and the complete set of keywords are read, and the difference between them is taken to obtain the keywords that have not been used yet;

[0149] S1119. Perform a traversal search on the keywords used in this round, and treat the keywords as seed keywords. After the keyword search, store them in the keyword completion set;

[0150] S11110. Pre-circumvent anti-crawler measures, that is, pre-circumvent possible crawler detection measures by disabling the automatic control feature of the Blink rendering engine, setting the User-Agent option, modifying the experimental options excludeSwitches, useAutomationExtension, using JavaScript to override the navigator.webdriver property, etc.;

[0151] S11111. Determine whether the crawler is countered, that is, determine whether the anti-crawler measures of the search engine are effective by detecting the structure of the web page obtained by the search;

[0152] S11112. Avoid recommended jumps, that is, when the current keyword can no longer produce search results, jump through the related recommendations of Baidu search engine. Due to the existence of anti-crawler mechanisms and the limitation of entry search results, it is impossible to guarantee that each keyword can obtain a sufficient number of search results. Therefore, a recommended jump strategy is introduced. Since related recommendations are often keywords in the same field, the number of browsed topics in each field can be guaranteed to be roughly the same;

[0153] S11113. After traversing the search results, in order to simulate user behavior and circumvent anti-crawler measures, the Selenium crawler will randomly stay on the page for a period of time;

[0154] S11114. Jump to search page, i.e., select sequential page jump or random page jump according to the parameters set by the user. Sequential page jump is more in line with the user's usage habits. Users often only pay attention to the web pages with higher rankings when using search engines, i.e., sequential page browsing. Random page jump can expand the web page coverage of traffic data collection and solve the problem of low ranking of web pages with less visits in search results, making the traffic more comprehensive;

[0155] S11115. When all the keywords used in this round of operation are traversed, the archive files recording the total set of keywords and the completed set of keywords are deleted;

[0156] In S11, reference Figure 6 ,The tunnel traffic acquisition phase includes the following steps:

[0157] S1121. Perform parameter configuration, configure appropriate parameters for the DNS tunnel tool, such as encoding method, DNS resource type, sending time interval, etc.;

[0158] S1122. Select a suitable file for DNS tunnel testing. In this embodiment, several randomly generated files of 1KB size are selected;

[0159] S1123 sends a DNS tunnel packet to build a DNS tunnel through the DNS tunnel server;

[0160] S1124. Read the response returned by the DNS tunnel tool server;

[0161] In S11, reference Figure 7 ,The traffic collection script deployment phase includes the following steps:

[0162] S1131. Select the monitoring protocol, that is, select the type of network protocol to be collected, such as DNS, HTTP, etc.;

[0163] S1132. Set BPF filtering conditions to quickly filter the underlying protocols such as the transport layer, network layer, data link layer, and physical layer;

[0164] S1133. Sniff all packets of the selected network port, that is, use Scapy to monitor the traffic of the specified network adapter;

[0165] S1134. Perform a relatively shallow and rapid filtering of the traffic according to the BPF filtering conditions;

[0166] S1135. Perform deep filtering at the application layer for traffic that meets the requirements of the underlying protocol;

[0167] S1136. Save the deep filtered traffic by adding append;

[0168] In S12, reference Figure 8 , including the following steps:

[0169] S121. Read the traffic data set, that is, read the saved DNS tunnel original traffic file;

[0170] S122. Traverse the DNS data packets in the traffic data set;

[0171] S122. If the primary domain name queried in the DNS data packet is the same as the domain name set for DNS tunnel construction, it is saved in a new traffic data set in an incremental manner;

[0172] S123. If the primary domain name queried in the DNS packet is different from the domain name set for DNS tunnel construction, it is discarded;

[0173] S124. Obtaining the traffic data set after traffic filtering;

[0174] Specifically, the main content of traffic filtering is to filter the normal traffic mixed in the DNS tunnel traffic file through a specific domain name to ensure that the test set is pure enough and the test results are more accurate;

[0175] In S13, reference Fig. 9 , including the following steps:

[0176] S131 reads the traffic data set after traffic filtering, that is, reads the saved DNS tunnel original traffic file after traffic filtering;

[0177] S132. Initialize the database based on the traffic data set information after traffic filtering and the selected packet-level metadata, so that each traffic data set after traffic filtering corresponds to a data table;

[0178] S133. Traverse the DNS data packets in the traffic data set after traffic filtering;

[0179] S134. According to the set packet-level metadata, extract the packet-level metadata in the DNS data packet in the traffic data set after traffic filtering;

[0180] S135. Save the extracted packet-level metadata into the corresponding data packet-level metadata table of the database storing metadata, and obtain a data packet-level database.

[0181] Specifically, referring to Table 1, it represents the packet-level metadata. Compared with the original traffic file, the packet-level metadata has the advantages of occupying less storage space and high usage efficiency. The disadvantage is that it needs to be inferred based on the subsequent aggregated metadata. Therefore, when the subsequent aggregated metadata changes, the packet-level metadata needs to be re-formulated and extracted. The packet-level metadata database stores the packet-level metadata extracted from each traffic data set in the form of a data table. In this embodiment, the lightweight database SQLite is used as the database for storing metadata.

[0182] refer to Figure 4 The client part consists of two virtual machines and a host machine. In the normal traffic acquisition stage, virtual machine A runs the Selenium crawler program, generates a DNS request by imitating the user's browsing behavior, and sends it to virtual machine B. Virtual machine B runs as a local DNS server and runs a traffic monitoring program on it to capture all DNS traffic flowing through virtual machine B. The host machine sends commands to virtual machine A to control the operation of its crawler program and regularly receives the traffic files captured by virtual machine B; the server side is composed of several DNS servers, among which the root domain name server and the top-level domain name server are public servers. When obtaining normal traffic, the authoritative server is the server of the crawler target website, and when obtaining tunnel traffic, the DNS tunnel server as the authoritative domain name server is autonomously controlled, that is, in the tunnel traffic acquisition stage, the DNS tunnel tool client is run on virtual machine A, and the server side of various open source DNS tunnel tools is run on the DNS tunnel server, and DNS tunnel traffic is captured on virtual machine B.

[0183]

[0184] Table 1

[0185] Furthermore, the step S2 includes the following steps:

[0186] S21. According to the packet-level database, a session-level aggregation key represented by src_ip, domain, and dst_ip is generated to obtain a session-level metadata library and a session-level feature library, and complete the first-order aggregation based on IP-Domain;

[0187] S22. Reference Fig.11, according to the session-level feature library, tensor conversion is performed, the obtained session feature data tensor is standardized, and the standardized session-level feature data is used to perform deep unsupervised training on the session-level autoencoder to obtain a trained session-level anomaly detection model and complete the first-order detection based on IP-Domain;

[0188] In S21, reference Fig.10 , including the following steps:

[0189] S211. Read the data packet level metadata table in the data packet level database;

[0190] S212. Create a session-level database based on the packet-level metadata table information of the packet-level database, the selected session-level metadata database and the session-level features, where each data table in the packet-level database corresponds to a session-level metadata table and a session-level feature table;

[0191] S213 traverses each packet-level metadata table of the packet-level database;

[0192] S214. Read the packet-level metadata in the data packet-level metadata table;

[0193] S215. Extract the primary domain name from the domain name field in the package-level metadata;

[0194] S216. Generate a session-level aggregation key, that is, extract the source IP represented as src_ip and the destination IP represented as dst_ip in the packet-level metadata, and combine them with the primary domain name domain extracted in step S215 to obtain a session-level aggregation key represented as src_ip, domain, dst_ip;

[0195] S217. Create session-level metadata entries, that is, if the currently processed aggregate key does not have a related record in the session-level metadata dictionary, it is regarded as a new session, and related entries are created in the session-level metadata dictionary and the session-level time dictionary, and the session-level metadata entries are integrated into session-level metadata;

[0196] S218. Update the session-level metadata, that is, if the currently processed aggregation key has a related record in the session-level metadata dictionary, update the session-level metadata dictionary and the session-level time dictionary;

[0197] S219. Obtaining session-level features by performing aggregation mean calculation on session-level metadata;

[0198] S2110. The aggregation key and session-level metadata are saved in the session-level metadata table to obtain a session-level metadata database, and the aggregation key and session-level feature data are saved in the session-level feature table to obtain a session-level feature database;

[0199] Specifically, refer to Fig.10 and Fig.11 ,The goal of first-order aggregation and detection is to detect common single-IP single-domain DNS tunnels. In these DNS tunnels, there is only one IP as the DNS tunnel server, and the primary domain name is the same when building the DNS tunnel. Therefore, the session-level aggregation key can effectively aggregate and highlight the anomalies of the above DNS tunnels; Under normal circumstances, the IPs of both parties in a single DNS query session are fixed, that is, there is only the client IP that initiates the DNS query and the server IP that responds to the DNS, and there is only one queried domain name, so the metadata calculated based on the session-level aggregation key is called session-level metadata;

[0200] The work of the first-order aggregation stage revolves around the aggregation calculation of the session-level aggregation key to calculate the data packet-level metadata and obtain the session-level metadata. By aggregating and calculating the packet-level metadata in the packet-level metadata database, the IP-Domain metadata database and the IP-Domain feature database are obtained. Each data table in the above two databases corresponds to a traffic file; since the start time of the same session under the aggregation key is the same, the update and insertion of data will cause a large amount of duplicate data to be generated. Therefore, in order to save memory and speed up processing efficiency, the time-related metadata is separated from other metadata during the session-level metadata aggregation calculation process. The IP-Domain metadata database and the time database are used for the calculation of the second-order aggregation, while the IP-Domain feature database is used for primary detection and early warning;

[0201] Refer to Table 2, which represents session-level metadata;

[0202] Refer to Table 3, which shows session-level features.

[0203]

[0204] Table 2

[0205]

[0206] Table 3

[0207] In S22, reference Fig.11 , including the following steps:

[0208] S221. Read the session-level feature data table in the session-level feature library;

[0209] S222. Traverse each session-level feature data table in the session-level feature library;

[0210] S223. Read the characteristic data in each session-level characteristic data table. To avoid memory explosion, read the set N characteristic data each time. In this embodiment, N is set to 10000.

[0211] S224. Perform tensor conversion to convert the read feature data from a tuple type into a tensor type suitable for the deep learning framework Pytorch;

[0212] S225. Process the session feature data tensor using the minimum-maximum normalization method to obtain the normalized session feature data, and map each type of metadata value to the interval [0, 1];

[0213] S226. Use the normalized session-level feature data as input data of the session-level autoencoder to perform deep unsupervised training;

[0214] S227. Save the trained session-level anomaly detection model as a .pth file;

[0215] Specifically, PyTorch is used to implement the deep learning part of the anomaly detection model. In the first-order detection stage, based on the anomaly detection principle, the session-level autoencoder is used as the anomaly detection model and RMSE is used as the loss function to perform feature detection on the traffic.

[0216] The detection model is composed of an encoder and a decoder. The input data x, i.e. the standardized session-level feature data, is input into the encoder and encoded into an intermediate value m. The decoder reconstructs the original input from the intermediate value m to obtain the output data x'.

[0217] The intermediate value m is expressed as:

[0218] m = encoder(x)

[0219] The output data x' is expressed as:

[0220] x′=decoder(m)

[0221] The loss value e is expressed as:

[0222] e = loss(x′-x)

[0223] Among them, loss is the loss value calculation function;

[0224] The encoder decoder is composed of a 12×10 fully connected layer, a 10×8 fully connected layer, an 8×6 fully connected layer, and a 6×4 fully connected layer, and each fully connected layer is a ReLU activation function layer. The decoder is composed of a 4×6 fully connected layer, a 6×8 fully connected layer, an 8×10 fully connected layer, and a 10×12 fully connected layer. The first three fully connected layers of the decoder are ReLU activation function layers, and the last fully connected layer is followed by a Sigmod activation function layer.

[0225] The detection model uses the Adam optimizer during training;

[0226] The loss RMSE(x,y) is expressed as:

[0227]

[0228] Among them, x i is the i-th element of the input data, y i is the i-th element of the output data, and k is the total number of elements of the input / output data.

[0229] Furthermore, the step S3 includes the following steps:

[0230] S31. According to the session-level metadata library, a communication-level aggregation key represented by src_ip and dst_ip is generated to obtain a communication-level feature library to complete communication-level feature aggregation;

[0231] S32. Generate a domain-level aggregation key based on the session-level metadata library to obtain a domain-level feature library to complete domain-level feature aggregation;

[0232] S33. Use the standardized domain-level features and communication-level feature data to perform deep unsupervised training on the domain-level autoencoder and the communication-level autoencoder, respectively, to obtain a trained domain-level anomaly detection model and a communication-level anomaly detection model, and complete the second-order detection;

[0233] In S31, reference Fig.12 , including the following steps:

[0234] S311. Read all data tables in the session-level metadata database;

[0235] S312. Create a communication-level feature database based on the session-level metadata table information and the selected communication-level features in the data session-level metadata database, where each session-level metadata table corresponds to a communication-level feature table;

[0236] S313. Traverse each session-level metadata table in the session-level metadata database;

[0237] S314. Read the session-level metadata in the session-level metadata table;

[0238] S315. Generate a communication-level aggregation key, extract the source IP represented as src_ip and the destination IP represented as dst_ip in the session-level metadata, and combine them to obtain a communication-level aggregation key represented as src_ip, dst_ip;

[0239] S316. Create a communication-level metadata entry, that is, if the currently processed aggregation key does not have a related record in the communication-level metadata dictionary, it is regarded as a new session, and a related entry is created in the communication-level metadata dictionary to be integrated into the communication-level metadata;

[0240] S317. Update the communication-level metadata, that is, if the currently processed aggregation key has a related record in the communication-level metadata dictionary, update the communication-level metadata dictionary;

[0241] S318. Obtain communication-level features by calculating the communication-level metadata;

[0242] S319. The communication-level aggregation key and the communication-level feature are saved in the communication-level feature table to obtain a communication-level feature library;

[0243] In S32, reference Fig.13 , including the following steps:

[0244] S321. Read all session-level metadata tables in the session-level metadata database;

[0245] S322. Based on the session-level metadata table information in the data session-level metadata database and the selected domain-level features, a domain-level feature database is created, where each session-level metadata table corresponds to a domain-level feature table;

[0246] S323. Traverse each session-level metadata table in the session-level metadata database;

[0247] S324. Read the session-level metadata in the session-level metadata table;

[0248] S325. Generate a domain-level aggregation key, that is, extract the primary domain name domain in the session-level metadata as the domain-level aggregation key;

[0249] S326. Create a domain-level metadata entry. If the currently processed aggregation key has no related record in the domain-level metadata dictionary, it is regarded as a new session, and a related entry is created in the domain-level metadata dictionary to be integrated into the domain-level metadata;

[0250] S327. Update the domain-level metadata, that is, if the currently processed aggregation key has a related record in the domain-level metadata dictionary, update the domain-level metadata dictionary;

[0251] S328. Obtain domain-level features by calculating domain-level metadata;

[0252] S329. Save the domain-level aggregation key and the domain-level features into the domain-level feature table to obtain a domain-level feature library;

[0253] Specifically, refer to Fig.12 and Fig.13 , the goal of second-order aggregation and detection is to detect distributed DNS tunnels. Distributed DNS tunnels can be used in conjunction with botnets. Attackers can deploy the DNS tunnel tool server on bot hosts, so that the IP of the DNS tunnel server is not the same. At the same time, attackers can register multiple domain names, or use the domain name generation algorithm DGA to generate multiple domain names. The characteristics of multiple IP × multiple domain names allow attackers to disperse the DNS tunnels that were originally concentrated in the same session and easily highlighted by session-level aggregation to different domain name sessions on different hosts; the present invention solves the problem of distributed DNS tunnels by combining the communication-level aggregation key and the domain-level aggregation key. In a communication process, the IPs of both communicating parties are fixed, so the data aggregated based on the (src_ip, dst_ip) aggregation key is called the communication level. The role of the aggregation key (domain) is to aggregate all data under the same primary domain name, so we call the data aggregated based on the aggregation key (domain) the domain level.

[0254] In the second-order aggregation stage, since the data packets aggregated by the first-order aggregation key (src_ip, domain, dst_ip) are sub-groups of the second-order aggregation keys (src_ip, dst_ip) and (domain), the detection features under the second-order aggregation key can be directly obtained by aggregating and calculating the metadata under the first-order aggregation key;

[0255] The principle and process of the second-order detection stage are similar to those of the first-order detection stage, but the reconstruction error output by the detection model is not used as a warning reference alone, but is uniformly summarized to the voting agency as the input data for comprehensive detection;

[0256] Refer to Table 4, which shows the communication level characteristics;

[0257] Refer to Table 5, which shows the domain level features.

[0258]

[0259] Table 4

[0260]

[0261]

[0262] Table 5

[0263] Furthermore, the step S4 includes the following steps:

[0264] S41. Read the reconstruction errors of the trained session-level anomaly detection model, domain-level anomaly detection model, and communication-level anomaly detection model;

[0265] S42. Calculate the reconstruction errors output by the three anomaly detection models through a voting machine to obtain anomaly scores;

[0266] S43. If the abnormal score output by the voting machine is greater than the set threshold, an alarm is issued;

[0267] S44. If the anomaly score output by the voting device is less than the set threshold, the current traffic is not processed, that is, the traffic being tested is considered safe;

[0268] Specifically, refer to Fig.14 and Fig.15 , Threshold Checker represents threshold checker, Alarm represents alarm, and the comprehensive detection stage comprehensively analyzes the detection results of the three aggregation features to obtain the final result of whether the current DNS traffic is DNS tunnel traffic. Since the second-order aggregation requires a certain amount of running time, the real-time performance of the comprehensive detection is weaker than that of the first-order detection, but the comprehensive detection is stronger than the first-order detection in detection performance and can detect distributed DNS tunnels. In this embodiment, the voter used is a small autoencoder, which is trained using the reconstruction errors output by the above three anomaly detection models in the test phase, and the training process is similar to that of the anomaly detection model.

[0269] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is selected primarily for readability and teaching purposes, rather than for explaining or defining the subject matter of the present invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is illustrative, not restrictive, with respect to the scope of the present invention, which is defined by the appended claims.

Claims

1. A DNS tunnel detection method, characterized in that: The following steps are involved: S1. Preprocess the traffic by acquiring a training traffic data set, filtering the traffic data set, extracting and storing packet-level metadata, and obtaining a data packet-level database; S2. According to the packet-level database, first-order aggregation and detection based on IP-Domain are performed to obtain a session-level metadata database, a session-level feature library, and a trained session-level anomaly detection model; S3. According to the session-level metadata database, based on the second-order aggregation and detection of IP and Domain, the trained domain-level anomaly detection model and communication-level anomaly detection model are obtained; S4. Based on the trained session-level anomaly detection model, domain-level anomaly detection model, and communication-level anomaly detection model, a nonlinear voter is used for comprehensive detection, and a high-level alarm is generated if a DNS tunnel is detected.

2. A DNS tunnel detection method according to claim 1, characterized in that: The S1 comprises the following steps: S11. Obtain training traffic data set; Specifically: the process of obtaining the training traffic data set includes the normal traffic acquisition stage of virtual machine A as the client that initiates the DNS query, the tunnel traffic acquisition stage of the client that deploys the DNS tunnel tool in virtual machine A, and the deployment traffic collection script stage of virtual machine B as the local DNS server; S12. Filter the DNS data packets in the traffic data set to obtain a traffic data set after traffic filtering; S13 reads the traffic data set after traffic filtering, extracts and stores packet-level metadata, and obtains a data packet-level database; In S11, the normal traffic acquisition stage includes the following steps: S1111. Set program related parameters; S1112. Select browser driver; S1113. Select a search engine as a basis for subsequent searches; S1114. Select a keyword word library as a search seed; S1115. Using the interruption and continuation mechanism, determine whether there is an archive file in the keyword thesaurus, the archive file includes a total set of keywords and a complete set of keywords, respectively, the keyword archive records all the keywords used in each round of crawling, and the keyword complete archive records the keywords that have been used; S1116. Randomly extract keywords. If there are no archived files left, this operation is regarded as a new round of crawling, and a number of data are randomly extracted from keyword word library files in various fields; S1117. Archiving keywords, that is, storing the keywords extracted in this round of operation into the total keyword set; S1118. Read the archived files for keywords. If there are archived files left, this operation is regarded as a continuation of the previous round of crawling. The total set of keywords and the complete set of keywords are read, and the difference between them is taken to obtain the keywords that have not been used yet; S1119. Perform a traversal search on the keywords used in this round, and treat the keywords as seed keywords. After the keyword search, store them in the keyword completion set; S11110. Take anti-crawler measures to avoid them; S11111. Determine whether the crawler is countered, that is, determine whether the anti-crawler measures of the search engine are effective by detecting the structure of the web page obtained by the search; S11112. Perform recommendation jump avoidance, that is, when the current keyword can no longer produce search results, jump through the search engine's related recommendations; S11113. Traverse the search results; S11114. Jump to search page, that is, select sequential page jump or random page jump according to the parameters set by the user; S11115. When all the keywords used in this round of operation are traversed, the archive files recording the total set of keywords and the completed set of keywords are deleted; In S11, the tunnel traffic acquisition stage includes the following steps: S1121. Perform parameter configuration for the DNS tunnel tool; S1122. Select the file for DNS tunnel testing; S1123 sends a DNS tunnel packet to build a DNS tunnel through the DNS tunnel server; S1124. Read the response returned by the DNS tunnel tool server; In S11, the stage of deploying the traffic collection script includes the following steps: S1131. Select monitoring protocol; S1132. Set BPF filtering conditions to quickly filter the underlying protocols of the transport layer, network layer, data link layer, and physical layer; S1133. Sniff all packets of the selected network port, that is, use Scapy to monitor the traffic of the specified network adapter; S1134. Rapidly filter the traffic according to the BPF filtering conditions; S1135. Perform deep filtering at the application layer for traffic that meets the requirements of the underlying protocol; S1136. Save the deep filtered traffic by adding append; The S12 includes the following steps: S121. Read the traffic data set, that is, read the saved DNS tunnel original traffic file; S122. Traverse the DNS data packets in the traffic data set; S122. If the primary domain name queried in the DNS data packet is the same as the domain name set for DNS tunnel construction, it is saved in a new traffic data set in an incremental manner; S123. If the primary domain name queried in the DNS packet is different from the domain name set for DNS tunnel construction, it is discarded; S124. Obtaining the traffic data set after traffic filtering; The S13 includes the following steps: S131 reads the traffic data set after traffic filtering, that is, reads the saved DNS tunnel original traffic file after traffic filtering; S132. Initialize the database based on the traffic data set information after traffic filtering and the selected packet-level metadata, so that each traffic data set after traffic filtering corresponds to a data table; S133. Traverse the DNS data packets in the traffic data set after traffic filtering; S134. According to the set packet-level metadata, extract the packet-level metadata in the DNS data packet in the traffic data set after traffic filtering; S135. Save the extracted packet-level metadata into the corresponding data packet-level metadata table of the database storing metadata, and obtain a data packet-level database.

3. A DNS tunnel detection method according to claim 2, characterized in that: The S2 comprises the following steps: S21. According to the packet-level database, a session-level aggregation key represented by src_ip, domain, and dst_ip is generated to obtain a session-level metadata library and a session-level feature library, and complete the first-order aggregation based on IP-Domain; S22. According to the session-level feature library, tensor conversion is performed, the obtained session feature data tensor is standardized, and the session-level autoencoder is trained in deep unsupervised manner using the standardized session-level feature data to obtain a trained session-level anomaly detection model, and complete the first-order detection based on IP-Domain; The S21 includes the following steps: S211. Read the data packet level metadata table in the data packet level database; S212. Create a session-level database based on the packet-level metadata table information of the packet-level database, the selected session-level metadata database and the session-level features, where each data table in the packet-level database corresponds to a session-level metadata table and a session-level feature table; S213 traverses each packet-level metadata table of the packet-level database; S214. Read the packet-level metadata in the data packet-level metadata table; S215. Extract the primary domain name from the domain name field in the package-level metadata; S216. Generate a session-level aggregation key, that is, extract the source IP represented as src_ip and the destination IP represented as dst_ip in the packet-level metadata, and combine them with the primary domain name domain extracted in step S215 to obtain a session-level aggregation key represented as src_ip, domain, dst_ip; S217. Create session-level metadata entries, that is, if the currently processed aggregate key does not have a related record in the session-level metadata dictionary, it is regarded as a new session, and related entries are created in the session-level metadata dictionary and the session-level time dictionary, and the session-level metadata entries are integrated into session-level metadata; S218. Update the session-level metadata, that is, if the currently processed aggregation key has a related record in the session-level metadata dictionary, update the session-level metadata dictionary and the session-level time dictionary; S219. Obtaining session-level features by performing aggregation mean calculation on session-level metadata; S2110. The aggregation key and session-level metadata are saved in the session-level metadata table to obtain a session-level metadata database, and the aggregation key and session-level feature data are saved in the session-level feature table to obtain a session-level feature database; The S22 includes the following steps: S221. Read the session-level feature data table in the session-level feature library; S222. Traverse each session-level feature data table in the session-level feature library; S223 reads the characteristic data in each session-level characteristic data table. To avoid memory explosion, read the set N characteristic data each time; S224. Perform tensor conversion to convert the read feature data from a tuple type into a tensor type suitable for the deep learning framework Pytorch; S225. Process the session feature data tensor using the minimum-maximum normalization method to obtain the normalized session feature data, and map each type of metadata value to the interval [0, 1]; S226. Use the normalized session-level feature data as input data of the session-level autoencoder to perform deep unsupervised training; S227. Save the trained session-level anomaly detection model as a .pth file.

4. A DNS tunnel detection method according to claim 3, characterized in that: The S3 comprises the following steps: S31. According to the session-level metadata library, a communication-level aggregation key represented by src_ip and dst_ip is generated to obtain a communication-level feature library to complete communication-level feature aggregation; S32. Generate a domain-level aggregation key based on the session-level metadata library to obtain a domain-level feature library to complete domain-level feature aggregation; S33. Use the standardized domain-level features and communication-level feature data to perform deep unsupervised training on the domain-level autoencoder and the communication-level autoencoder, respectively, to obtain a trained domain-level anomaly detection model and a communication-level anomaly detection model, and complete the second-order detection; The S31 includes the following steps: S311. Read all data tables in the session-level metadata database; S312. Create a communication-level feature database based on the session-level metadata table information and the selected communication-level features in the data session-level metadata database, where each session-level metadata table corresponds to a communication-level feature table; S313. Traverse each session-level metadata table in the session-level metadata database; S314. Read the session-level metadata in the session-level metadata table; S315. Generate a communication-level aggregation key, extract the source IP represented as src_ip and the destination IP represented as dst_ip in the session-level metadata, and combine them to obtain a communication-level aggregation key represented as src_ip, dst_ip; S316. Create a communication-level metadata entry, that is, if the currently processed aggregation key does not have a related record in the communication-level metadata dictionary, it is regarded as a new session, and a related entry is created in the communication-level metadata dictionary to be integrated into the communication-level metadata; S317. Update the communication-level metadata, that is, if the currently processed aggregation key has a related record in the communication-level metadata dictionary, update the communication-level metadata dictionary; S318. Obtain communication-level features by calculating the communication-level metadata; S319. The communication-level aggregation key and the communication-level feature are saved in the communication-level feature table to obtain a communication-level feature library; The S32 includes the following steps: S321. Read all session-level metadata tables in the session-level metadata database; S322. Based on the session-level metadata table information in the data session-level metadata database and the selected domain-level features, a domain-level feature database is created, where each session-level metadata table corresponds to a domain-level feature table; S323. Traverse each session-level metadata table in the session-level metadata database; S324. Read the session-level metadata in the session-level metadata table; S325. Generate a domain-level aggregation key, that is, extract the primary domain name domain in the session-level metadata as the domain-level aggregation key; S326. Create a domain-level metadata entry. If the currently processed aggregation key has no related record in the domain-level metadata dictionary, it is regarded as a new session, and a related entry is created in the domain-level metadata dictionary to be integrated into the domain-level metadata; S327. Update the domain-level metadata, that is, if the currently processed aggregation key has a related record in the domain-level metadata dictionary, update the domain-level metadata dictionary; S328. Obtain domain-level features by calculating domain-level metadata; S329. Save the domain-level aggregation key and domain-level features into the domain-level feature table to obtain a domain-level feature library.

5. A DNS tunnel detection method according to claim 4, characterized in that: The S4 comprises the following steps: S41. Read the reconstruction errors of the trained session-level anomaly detection model, domain-level anomaly detection model, and communication-level anomaly detection model; S42. Calculate the reconstruction errors output by the three anomaly detection models through a voting machine to obtain anomaly scores; S43. If the abnormal score output by the voting machine is greater than the set threshold, an alarm is issued; S44. If the anomaly score output by the voter is less than the set threshold, the current traffic is not processed.

Citation Information

Patent Citations

  • DNS (Domain Name System) tunnel Trojan detection method based on communication behavior analysis

    CN107733851A

  • DNS hidden tunnel detection method and system

    CN111953673A

  • DNS (Domain Name Server) tunnel detection method and device and electronic equipment

    CN113347210A

  • Low-throughput DNS covert channel detection method and device

    CN113810372A

  • Web log abnormal behavior identification method based on knowledge graph

    CN114328962A