DNS Abnormal Detection Method, Device, Equipment and Storage Medium

By performing multi-level abnormality detection on the domain names and traffic of DNS data packets, combined with the XGBoost model and time window algorithm, the problem of low efficiency and accuracy of DNS data abnormality detection is solved, and efficient and comprehensive network security defense is achieved.

CN116192417BActive Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211084762.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-07-18
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

The existing DNS data anomaly detection is not efficient and accurate in industrial-grade large-scale data, making it difficult to effectively deal with complex network anomaly behaviors.

Method used

By obtaining the domain name of the DNS packet, determining the number of visits, and combining domain name anomaly detection and traffic anomaly detection, the XGBoost model, time window algorithm and association analysis algorithm of large-scale parallel computing are used to comprehensively evaluate the abnormal probability and score of the DNS packet to achieve multi-level abnormal detection.

Benefits of technology

It improves the recognition efficiency and accuracy of DNS data abnormality detection, reduces the domain name abnormality error rate, realizes all-round defense of multi-band DNS traffic, reduces the workload of network security personnel and saves costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116192417B_ABST
    Figure CN116192417B_ABST
Patent Text Reader

Abstract

This application relates to data processing, and provides a DNS anomaly detection method, device, equipment and storage medium. The method includes: obtaining the domain names of multiple DNS data packets, and determining the access volume of each domain name according to the domain names of the multiple DNS data packets; performing anomaly detection on the domain names of each DNS data packet to obtain the first anomaly probability of each DNS data packet; performing anomaly detection on each domain name according to the access volume of each domain name to obtain the second anomaly probability corresponding to each domain name; determining the target anomaly evaluation information of each DNS data packet based on the first anomaly probability of each DNS data packet and the second anomaly probability corresponding to each domain name; and determining the target DNS data packets with anomalies from the multiple DNS data packets according to the target anomaly evaluation information of each DNS data packet. This application also relates to blockchain, aiming to improve the recognition efficiency and accuracy of DNS anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a DNS anomaly detection method, device, equipment, and storage medium. Background Art

[0002] In the Internet field, DNS (Domain Name System) is used to provide services such as load balancing and permission verification, and is an important network facility in the Internet. During the communication transmission process, DNS will record a lot of flow data. In order to provide users with a better network usage experience, network defense personnel usually do not perform too many detections on DNS data. Therefore, in the face of malicious attack behaviors, or phenomena such as network hijacking and network fraud, it is easy to cause unpredictable losses to enterprises.

[0003] To solve the above problems, it is necessary to perform anomaly detection based on DNS data. However, during the detection process of industrial-scale massive DNS data, most network security personnel only perform targeted detection and defense on a certain type or category of DNS data. For the intricate network anomaly behaviors, the current recognition efficiency and accuracy of DNS data anomaly detection are not high. Summary of the Invention

[0004] The main purpose of this application is to provide a DNS anomaly detection method, device, equipment, and storage medium, aiming to improve the recognition efficiency and accuracy of DNS data anomaly detection.

[0005] In a first aspect, this application provides a DNS anomaly detection method, including:

[0006] Obtain the domain names of multiple DNS data packets, and determine the access volume of each domain name according to the domain names of the multiple DNS data packets;

[0007] Perform anomaly detection on the domain name of each DNS data packet to obtain the first anomaly probability of each DNS data packet;

[0008] Perform anomaly detection on each domain name according to the access volume of each domain name to obtain the second anomaly probability corresponding to each domain name;

[0009] Based on the first anomaly probability of each DNS data packet and the second anomaly probability corresponding to each domain name, determine the target anomaly evaluation information of each DNS data packet;

[0010] According to the target anomaly evaluation information of each DNS data packet, determine the target DNS data packets with anomalies from the multiple DNS data packets.

[0011] In a second aspect, the present application further provides a DNS anomaly detection device, which includes:

[0012] A domain name acquisition module, configured to acquire the domain names of multiple DNS data packets, and determine the access volume of each domain name according to the domain names of the multiple DNS data packets;

[0013] A first anomaly detection module, configured to perform anomaly detection on the domain name of each DNS data packet to obtain a first anomaly probability of each DNS data packet;

[0014] A second anomaly detection module, configured to perform anomaly detection on each domain name according to the access volume of each domain name to obtain a second anomaly probability corresponding to each domain name;

[0015] A target anomaly evaluation module, configured to determine target anomaly evaluation information of each DNS data packet based on the first anomaly probability of each DNS data packet and the second anomaly probability corresponding to each domain name;

[0016] An anomaly data determination module, configured to determine target DNS data packets with anomalies from the multiple DNS data packets according to the target anomaly evaluation information of each DNS data packet.

[0017] In a third aspect, the present application further provides a computer device, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the DNS anomaly detection method described above are implemented.

[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the DNS anomaly detection method described above are implemented.

[0019] The present application provides a DNS anomaly detection method, apparatus, device, and storage medium. The present application obtains the domain names of multiple DNS data packets, and determines the access volume of each domain name based on the domain names of the multiple DNS data packets; performs anomaly detection on the domain names of each DNS data packet to obtain the first anomaly probability of each DNS data packet; performs anomaly detection on each domain name based on the access volume of each domain name to obtain the second anomaly probability corresponding to each domain name; determines the target anomaly evaluation information of each DNS data packet based on the first anomaly probability of each DNS data packet and the second anomaly probability corresponding to each domain name; and determines the target DNS data packet with anomalies from the multiple DNS data packets according to the target anomaly evaluation information of each DNS data packet. By combining the anomaly detection of the domain names of DNS data packets and the anomaly detection of the access volume of domain names, the target DNS data packet with anomalies can be accurately identified, thereby greatly reducing the misjudgment rate of domain name anomalies and significantly improving the identification efficiency and accuracy of DNS data anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of the steps of a DNS anomaly detection method provided by an embodiment of the present application;

[0022] Figure 2 For Figure 1 a sub-step flowchart of the DNS anomaly detection method in

[0023] Figure 3 For Figure 1 another sub-step flowchart of the DNS anomaly detection method in

[0024] Figure 4 It is a schematic block diagram of a DNS anomaly detection apparatus provided by an embodiment of the present application;

[0025] Figure 5 For Figure 4 a schematic block diagram of a sub-module of the DNS anomaly detection apparatus in

[0026] Figure 6 For Figure 4 a schematic block diagram of another sub-module of the DNS anomaly detection apparatus in

[0027] Figure 7A schematic block diagram of a computer device provided by an embodiment of the present application.

[0028] The realization of the purpose of the present application, functional characteristics and advantages will be further described in combination with the embodiments with reference to the accompanying drawings. Specific embodiments

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0030] The flowchart shown in the accompanying drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation. In addition, although the functional modules are divided in the device schematic diagram, in some cases, it can be different from the module division in the device schematic diagram.

[0031] The embodiments of the present application provide a DNS anomaly detection method, device, device and storage medium. Among them, the DNS anomaly detection method can be applied to a terminal device or a server. The terminal device can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant and a wearable device; the server can be a single server or a server cluster composed of multiple servers. The following takes the application of the DNS anomaly detection method to the server as an example for explanation.

[0032] Next, in conjunction with the accompanying drawings, some embodiments of the present application will be described in detail. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0033] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the steps of a DNS anomaly detection method provided by an embodiment of the present application.

[0034] As Figure 1 shown, the DNS anomaly detection method includes steps S101 to S105.

[0035] Step S101, obtain the domain names of multiple DNS data packets, and determine the access volume of each domain name according to the domain names of the multiple DNS data packets.

[0036] The DNS (Domain Name System) is a distributed database that maps domain names and IP addresses to each other, enabling people to access the Internet more conveniently. DNS packets can be the streaming data recorded during the communication and transmission of DNS. For example, when network attacks, phishing frauds, abnormal intrusions, etc. occur, streaming data will be generated in the DNS resolution service. Specifically, DNS packets can include parameters such as domain names, source IP and port, destination IP and port, sending and receiving rates, packet sizes, data lengths, and various response times.

[0037] Among them, multiple DNS packets can be the DNS packets obtained within a preset time period. Each DNS packet can include the domain name it accesses, and the server can obtain the domain names of multiple DNS packets within the preset time period. The preset time period can be set according to the actual situation. The preset time period is, for example, 30 seconds or 1 minute. For example, the server obtains 500 DNS packets within 30 seconds.

[0038] In one embodiment, after obtaining the domain names of multiple DNS packets, the access volume of each domain name is determined according to the domain names of the multiple DNS packets. Among them, there can be multiple domain names. The domain name of each DNS packet can represent the access volume of a domain name once, and this access volume can also be called the access frequency. The domain names of different DNS packets can be the same or different. Therefore, through the domain names of multiple DNS packets, the access volume of each domain name can be accurately determined.

[0039] In one embodiment, before obtaining the domain names of multiple DNS packets, it further includes: obtaining multiple DNS packets, and performing rule filtering on the multiple DNS packets by establishing a domain name blacklist; and, by analyzing the characteristics of similar type domain names, reverse-cracking the DGA generation mechanism to perform rule filtering on the multiple DNS packets to obtain abnormal target DNS packets. Among them, reverse-cracking the DGA generation mechanism can be achieved by rainbow table collision of the domain name generation algorithm. It should be noted that the DGA generation mechanism refers to generating a large number of alternative domain names through the DGA algorithm and querying them. When an attack needs to be launched, a small number of domain names are selected from the alternative domain name list for registration, and the fast-changing IP technology can be applied to the registered domain names to quickly change the domain names and IPs. By establishing a domain name blacklist and reverse-cracking the DGA generation mechanism to perform rule filtering on the multiple DNS packets, the filtered multiple DNS packets can be accurately determined as abnormal target DNS packets.

[0040] It should be noted that, to further ensure the privacy and security of the above DNS data packets and related information, the above DNS data packets and related information can also be stored in a node of a blockchain. The technical solution of this application can also be applied to adding other data files stored on the blockchain. The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.

[0041] Step S102: Perform anomaly detection on the domain name of each DNS data packet to obtain the first anomaly probability of each DNS data packet.

[0042] It should be noted that since there are many differences between malicious domain names and normal domain names in DNS data packets, anomaly detection can be performed on the domain names of DNS data packets to obtain the first anomaly probability of the DNS data packets.

[0043] In one embodiment, as Figure 2 shown, step S102 includes: sub-step S1021 to sub-step S1023.

[0044] Sub-step S1021: Based on a preset domain name anomaly detection model, perform anomaly detection on the domain name of each DNS data packet to obtain the first probability that each DNS data packet has an anomaly.

[0045] Among them, the domain name anomaly detection model includes bigram model, HMM model, deep learning model based on word-hashing technology, malicious domain name anomaly detection model based on LSTM, and detection models for DGA domain names such as random forest, CNN, and LSTM.

[0046] It should be noted that the domain names of DNS data packets are divided into normal domain names and abnormal domain names. Abnormal domain names are generally malicious domain names generated by the Domain Generation Algorithm (DGA). For the currently publicly available DGA open dataset provided by 360netlab, the recognition rates of different families may vary, and the recognition rates of domain names in a small number of DGA families are not ideal, such as virut, suppobox, bigviktor, conficker, vawtrak, mydoom, etc. Among them, both suppobox and mydoom in the DGA family are generated through the word list mechanism, and vawtrak is generated through the hash mechanism.

[0047] In one embodiment, the domain names of multiple DNS packets are input into a bigram model for anomaly detection to obtain the first probability of anomaly for each DNS packet. Specifically, through the processing layer in the bigram model, the domain names of DNS packets are processed to obtain domain name feature values; through the screening layer in the bigram model, the domain name feature values are screened according to a preset threshold to retain the domain name feature values less than or equal to the preset threshold; through the detection layer in the bigram model, a statistical detection of the character distribution of the DGA domain names corresponding to the domain name feature values is performed to obtain the first probability of anomaly for the DNS packet. It should be noted that by performing a statistical detection of the character distribution of normal domain names and DGA domain names through the bigram model, the first probability of anomaly for the DNS packet can be accurately calculated.

[0048] It can be understood that the embodiments of the present application can also use an HMM model to perform featureless real-time detection on the domain names of DNS packets; or use a deep learning model based on the word-hashing technology to perform DGA detection on the domain names of DNS packets; or use a malicious domain name anomaly detection technology based on LSTM to realize real-time mining of the network features of domain names and classification of DGA families; or identify abnormal domain names by statistically analyzing and mining the character distribution of domain names, and calculating the edit distance, Jaccard coefficient, etc. between domain names after feature processing, so as to determine the first probability of anomaly for the DNS packet, etc. The embodiments of the present application do not make specific limitations on this.

[0049] Sub-step S1022: Based on a preset domain name statistical analysis algorithm, perform anomaly analysis on the domain names of each DNS packet to obtain the second probability of anomaly for each DNS packet.

[0050] Among them, the domain name statistical analysis algorithm includes using the proportion of vowel letters in the domain name, Gibberish detection algorithm, Shannon entropy of the characters in the domain name string, HMM coefficient of the characters in the domain name string, TF-IDF mean and variance calculated after domain name n-gram, etc. for domain name statistical analysis algorithm. Through the domain name statistical analysis algorithm, the difference between the domain name of the DNS packet and the abnormal domain name can be analyzed, so as to accurately obtain the second probability of anomaly for the DNS packet.

[0051] In one embodiment, character conversion is performed on the domain name of a DNS data packet to obtain a first domain name; the first domain name is subjected to TF-IDF conversion processing to obtain a second domain name, and the mean and variance values between the second domain name and multiple DGA domain names are calculated; according to the mean and variance values between the second domain name and multiple DGA domain names, a second probability that the DNS data packet corresponding to the second domain name is abnormal is determined. Among them, the character conversion includes uni-gram, bi-gram, and tri-gram conversions. TF-IDF conversion (term frequency–inverse document frequency) is a weighted calculation method used for information retrieval and data mining. TF is the term frequency, and IDF is the inverse document frequency index. Through the mean and variance values between the second domain name and multiple DGA domain names, the association and difference information between the domain name of the DNS data packet and the abnormal domain name can be determined, and the ability to distinguish DGA domain names is relatively good.

[0052] It should be noted that in the field of machine learning, the N-gram language model has been widely used in language processing tasks. With the development of deep learning, the neural network model may have a slight improvement in effect, but its efficiency is relatively low. By first performing uni-gram, bi-gram, and tri-gram conversions on the domain name and then performing TF-IDF conversion processing, the corresponding mean and variance are obtained respectively, and the distributions of the mean and variance of the domain name of the DNS data packet and the DGA domain name are compared. From the perspective of the mean and variance distributions, the TF-IDF model after n-gram extracts the association and difference information between domain names, which can greatly improve the calculation effect of the second abnormal probability and is more efficient than the neural network model in terms of recognition efficiency.

[0053] In one embodiment, the domain name statistical analysis algorithm using the proportion of vowel letters refers to obtaining the proportion of vowel letters in the domain name of the DNS data packet, and determining the second probability that the DNS data packet is abnormal according to the proportion of the vowel letters. It should be noted that the preference for naming letters in normal domain names generally results in a relatively large number of vowel letters because people tend to choose combinations of letters that are easy to read, and the advantage of vowel letters thus emerges. When generating DGA domain names, due to randomness and time factors, there is no preference in the selection of vowel letters during generation, so the proportion of vowel letters generated in DGA domain names is relatively low. Therefore, through the proportion of vowel letters for domain name statistical analysis, the difference between the domain name of the DNS data packet and the abnormal domain name can be analyzed, and thus the second probability that the DNS data packet is abnormal can be accurately obtained.

[0054] In one embodiment, the domain name statistical analysis algorithm using Gibberish detection refers to using a two-character-level Markov model. Through training on a preset corpus, the repetition frequency of character pairs is statistically obtained, and the second probability of the DNS data packet being abnormal is determined according to the repetition frequency of character pairs. It should be noted that whether a designed domain name is easy to read and pronounce can usually be judged by the method of Gibberish detection. After Gibberish completes the transformation of the corpus data, the probability distribution of different character pairs for each character after a given initial value has been obtained. Then, the probability of the middle character can be obtained based on the product of the probabilities of adjacent character pairs, so as to accurately obtain the second probability of the DNS data packet being abnormal.

[0055] It can be understood that the embodiments of the present application can also use the Shannon entropy of domain name string characters, the HMM coefficient of domain name string characters, etc. for domain name statistical analysis, and this embodiment does not make specific limitations in this regard. Among them, the Shannon entropy is mainly used for the quantitative measurement of information volume. Due to the randomness of domain name character combinations, the difference in information volume between domain names can be explored through the Shannon entropy. It should be noted that a normal domain name design tends to be some common character combinations. Usually, this character difference needs to be abstracted into a recognizable language, and this difference can be recognized through the HMM coefficient. By training the HMM model with English words, since DGA domain names have the characteristic of being generated irregularly, the HMM coefficient of DGA domain names is relatively low compared to normal domain names.

[0056] Sub-step S1023: Determine the first abnormal probability of each DNS data packet according to the first probability and the second probability of each DNS data packet being abnormal.

[0057] In one embodiment, the first abnormal probability of each DNS data packet is obtained by calculating the average value or weighted average value between the first probability and the second probability of each DNS data packet being abnormal.

[0058] In one embodiment, the first abnormal probability of each DNS data packet is obtained by calculating the sum value between the first probability and the second probability of each DNS data packet being abnormal.

[0059] It can be understood that other formulas can also be used to calculate the first probability and the second probability of each DNS data packet being abnormal to obtain the first abnormal probability of each DNS data packet, and this embodiment of the present application does not make specific limitations in this regard.

[0060] Step S103: Perform abnormal detection on each domain name according to the access volume of each domain name, and obtain the second abnormal probability corresponding to each domain name.

[0061] It should be noted that the access volume of abnormal DNS data packets and that of normal DNS data packets also have different characteristics. Therefore, abnormal detection of domain names can be performed based on the access volume of domain names to obtain the second abnormal probability corresponding to the domain names.

[0062] In one embodiment, as Figure 3 shown, step S101 includes: sub-step S1031 to sub-step S1032.

[0063] Sub-step S1031: Determine the traffic level corresponding to each domain name according to the access volume of each domain name.

[0064] Among them, the traffic level can be divided according to the number of DNSs within a unit time, that is, determined according to the access volume of domain names within a unit time. It can be understood that the specific setting of the traffic level can be set according to the actual situation. For example, a two-level traffic level, a three-level traffic level, or a traffic level with more than three levels can be set. The embodiments of the present application do not make specific limitations on this.

[0065] Exemplarily, the traffic level can include a first traffic level, a second traffic level, and a third traffic level. The access volume corresponding to the first traffic level is greater than the access volume corresponding to the second traffic level, and the access volume corresponding to the second traffic level is greater than the access volume corresponding to the third traffic level. Or rather, the traffic level includes a high-frequency traffic level, a medium-frequency traffic level, and a low-frequency traffic level.

[0066] Sub-step S1032: Perform traffic anomaly detection on each domain name according to the traffic level corresponding to each domain name to obtain the second abnormal probability corresponding to each domain name.

[0067] It should be noted that by differentiating according to the traffic level corresponding to each domain name, different anomaly detection and recognition strategies can be used to perform anomaly detection on different domain names, which can greatly improve the recognition accuracy of traffic anomaly detection, thereby greatly improving the recognition efficiency and accuracy of DNS data anomaly detection.

[0068] In one embodiment, the traffic level includes a first traffic level, a second traffic level, and a third traffic level. The access volume corresponding to the first traffic level is greater than the access volume corresponding to the second traffic level, and the access volume corresponding to the second traffic level is greater than the access volume corresponding to the third traffic level. Specifically, based on a preset traffic anomaly detection model, anomaly detection is performed on the domain names corresponding to the first traffic level to obtain a first anomaly detection result; based on a preset time window algorithm, anomaly detection is performed on the domain names corresponding to the second traffic level to obtain a second anomaly detection result; based on a preset correlation analysis algorithm, anomaly analysis is performed on the domain names corresponding to the third traffic level to obtain a third anomaly detection result; according to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result, the second abnormal probability corresponding to each domain name is determined.

[0069] It should be noted that, the traffic anomaly detection model is preferably used to detect anomalies in the domain names corresponding to the first traffic level (with the largest traffic volume), and then anomalies in the domain names corresponding to the second traffic level (with slightly smaller traffic volume) and the third traffic level (with the smallest traffic volume) are detected, which can reduce the server load and improve the recognition efficiency of DNS data anomaly detection.

[0070] In one embodiment, based on a preset traffic anomaly detection model, anomalies in the domain names corresponding to the first traffic level are detected to obtain a first anomaly detection result, including: invoking an XGBoost detection model corresponding to the first traffic level; using the XGBoost detection model to detect anomalies in the domain names corresponding to the first traffic level to obtain the anomaly probabilities of multiple domain names corresponding to the first traffic level; and taking the anomaly probabilities of multiple domain names corresponding to the first traffic level as the first anomaly detection result.

[0071] Among them, the traffic anomaly detection model may include an XGBoost detection model, and the XGBoost detection model can be iteratively trained through training samples constructed by parameters such as source IP and port, destination IP and port, sending and receiving rate, packet size, data length, and various response times. Compared with complex CNN, LSTM, and Bert deep learning models, the XGBoost detection model is basically unaffected by the model effect in the case of certain sample imbalance, can achieve large-scale parallel computing, has fast model detection efficiency, good effect, and strong stability. Therefore, using the XGBoost detection model can improve the accuracy and recognition efficiency of traffic anomaly detection for domain names corresponding to the first traffic level with higher traffic volume.

[0072] In one embodiment, based on a preset time window algorithm, anomalies in the domain names corresponding to the second traffic level are detected to obtain a second anomaly detection result, including: invoking a time window algorithm corresponding to the second traffic level; using the time window algorithm to detect anomalies in the domain names corresponding to the second traffic level to obtain the anomaly probabilities of multiple domain names corresponding to the second traffic level; and taking the anomaly probabilities of multiple domain names corresponding to the second traffic level as the second anomaly detection result.

[0073] It should be noted that, the time window algorithm is used to detect anomalies in the domain names corresponding to the second traffic level, so as to count the access frequencies of domain names within the time window, and when the time window accumulates to a certain number, the entropy value (the negative logarithm of the amount of probability) is calculated according to the access frequencies of a certain number of time windows. The smaller the entropy value, the higher the anomaly probability of the domain name, and the larger the entropy value, the lower the anomaly probability of the domain name. The inventor found in multiple experiments that the time window algorithm can relatively accurately calculate the anomaly probabilities of multiple domain names with medium traffic.

[0074] In one embodiment, based on a preset correlation analysis algorithm, perform anomaly analysis on the domain names corresponding to the third traffic level to obtain a third anomaly detection result, including: invoking the correlation analysis algorithm corresponding to the third traffic level; using the correlation analysis algorithm to perform anomaly detection on the domain names corresponding to the third traffic level to obtain the anomaly probabilities of multiple domain names corresponding to the third traffic level; and taking the anomaly probabilities of multiple domain names corresponding to the third traffic level as the third anomaly detection result.

[0075] It should be noted that the correlation analysis algorithm is used to perform anomaly detection on the domain names corresponding to the third traffic level, so as to construct a rule engine to determine the anomaly probabilities of multiple domain names. The rule engine can be time series observation. For example, observe the trend and volatility of traffic at the hourly or minute level in units of days (or 10 days). After several days, it is found that the access trends within the same time period every day are highly consistent. Combine IP regionality, activity range, etc. for monitoring and analysis. The monitoring and analysis means can be a series of detection indicators such as anomaly clustering analysis and label propagation mining, so as to accurately determine the similarity between the domain names corresponding to the third traffic level and the abnormal domain names, and obtain the anomaly probabilities of multiple domain names as the third anomaly detection result.

[0076] In one embodiment, according to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result, the second anomaly probability corresponding to each domain name can be determined. Exemplarily, combine the anomaly probabilities of multiple domain names in the first anomaly detection result, the anomaly probabilities of multiple domain names in the second anomaly detection result, and the anomaly probabilities of multiple domain names in the third anomaly detection result to obtain the second anomaly probability corresponding to each domain name.

[0077] In one embodiment, the traffic levels include a high-frequency traffic level, a medium-frequency traffic level, and a low-frequency traffic level; based on the traffic anomaly detection strategy corresponding to the high-frequency traffic level, perform anomaly detection on the domain names corresponding to the high-frequency traffic level to obtain a first anomaly detection result; based on the traffic anomaly detection strategy corresponding to the medium-frequency traffic level, perform anomaly detection on the domain names corresponding to the medium-frequency traffic level to obtain a second anomaly detection result; based on the traffic anomaly detection strategy corresponding to the low-frequency traffic level, perform anomaly detection on the domain names corresponding to the low-frequency traffic level to obtain a third anomaly detection result; according to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result, determine the second anomaly probability corresponding to each domain name. It should be noted that traffic anomaly detection is performed on different domain names in the order of high frequency, medium frequency, and low frequency, so as to give priority to processing the domain names corresponding to the high-frequency traffic level with more entries, which can reduce the server load.

[0078] Step S104: Determine the target anomaly evaluation information of each DNS packet based on the first anomaly probability of each DNS packet and the second anomaly probability corresponding to each domain name.

[0079] Among them, the target anomaly evaluation information includes a target anomaly probability, a target anomaly score, and / or a target anomaly rating. It should be noted that for large-scale DNS data, the embodiments of the present application propose a dual mechanism combining DNS domain name anomaly detection and DNS traffic anomaly detection, that is, based on the first anomaly probability of each DNS packet and the second anomaly probability corresponding to each domain name, to determine the target anomaly evaluation information of each DNS packet, which can improve the recognition efficiency and accuracy of DNS packets for anomaly detection.

[0080] In one embodiment, according to the second anomaly probability corresponding to each domain name, determine the third anomaly probability of multiple DNS packets; the third anomaly probability of the DNS packet matches the second anomaly probability corresponding to the domain name of the DNS packet; according to the first anomaly probability and the third anomaly probability of each DNS packet, determine the target anomaly evaluation information of each DNS packet.

[0081] Exemplarily, the target anomaly evaluation information is the target anomaly probability. By calculating the average value or weighted average value of the first anomaly probability and the third anomaly probability of each DNS packet, the target anomaly probability of the DNS packet can be calculated.

[0082] Exemplarily, the target anomaly evaluation information is the target anomaly score. By calculating the product of the first anomaly probability of the DNS packet and the first preset score, a first integral value is obtained; by calculating the product of the third anomaly probability and the second preset score, a second integral value is obtained, and by calculating the sum value of the first integral value and the second integral value, the target anomaly score of the DNS packet is obtained. Among them, the first preset score and the second preset score can be set according to the actual situation. The first preset score is, for example, 60 points, and the second preset score is, for example, 40 points.

[0083] In one embodiment, the target anomaly evaluation information includes a target anomaly score and a target anomaly rating. Among them, by calculating the product of the first anomaly probability of the DNS packet and the first preset score, a first integral value is obtained; by calculating the product of the third anomaly probability and the second preset score, a second integral value is obtained, and by calculating the sum value of the first integral value and the second integral value, the target anomaly score is obtained; according to the target anomaly score, determine the target anomaly rating of the DNS packet. Among them, different ratings are divided by different score intervals, so the first score can be determined according to the interval where the calculated first score is located. For example, if the first score is 40 points, it corresponds to the first type of anomaly level of the first rating; if the first score is 60 points, it corresponds to the second type of anomaly level of the first rating.

[0084] Step S105: According to the target anomaly evaluation information of each DNS packet, determine the target DNS packets with anomalies from multiple DNS packets.

[0085] It should be noted that the embodiment of the present application uses DNS static domain names and dynamic traffic data to mine abnormal behaviors and filter them in a targeted manner. The target anomaly evaluation information includes target anomaly probability, target anomaly score or target anomaly rating. The target anomaly evaluation information can be used to accurately determine the target DNS data packets with abnormalities from multiple DNS data packets, thereby greatly reducing the domain name anomaly misjudgment rate and greatly improving the recognition efficiency and accuracy of DNS data anomaly detection.

[0086] In one embodiment, the target anomaly evaluation information includes a target anomaly probability, and an abnormal target DNS data packet is determined from multiple DNS data packets, and the target anomaly probability of the target DNS data packet is greater than a preset anomaly probability threshold.

[0087] In one embodiment, the target anomaly evaluation information includes a target anomaly score, and an abnormal target DNS data packet is determined from multiple DNS data packets, and the target anomaly score of the target DNS data packet is greater than a preset anomaly score threshold.

[0088] In one embodiment, the target anomaly evaluation information includes a target anomaly rating, and an abnormal target DNS data packet is determined from multiple DNS data packets, and the target anomaly rating of the target DNS data packet is greater than a preset anomaly score threshold.

[0089] In one embodiment, after determining that there is an abnormal target DNS data packet, it can be determined that the target DNS data packet involves a high-risk IP, so its access frequency per unit time can be limited or directly banned; of course, it is also possible to first limit it, further observe its abnormal activities, and then directly ban it.

[0090] The DNS anomaly detection method provided in the above embodiment obtains the domain names of multiple DNS data packets, and determines the access volume of each domain name according to the domain names of multiple DNS data packets; performs anomaly detection on the domain name of each DNS data packet to obtain the first anomaly probability of each DNS data packet; performs anomaly detection on each domain name according to the access volume of each domain name to obtain the second anomaly probability corresponding to each domain name; determines the target anomaly evaluation information of each DNS data packet based on the first anomaly probability of each DNS data packet and the second anomaly probability corresponding to each domain name; determines the target DNS data packet with anomalies from multiple DNS data packets according to the target anomaly evaluation information of each DNS data packet. By combining the domain name anomaly detection of DNS data packets and the anomaly detection of the domain name access volume, the target DNS data packet with anomalies can be accurately identified, thereby greatly reducing the misjudgment rate of domain name anomalies and greatly improving the recognition efficiency and accuracy of DNS data anomaly detection.

[0091] At present, there is a lack of systematic detection methods in the field of DNS anomaly detection. Most of the anomaly detections for DNS data only detect and identify from a single dimension and a single type of abnormal behavior. With the rapid development of network technology, attack methods emerge in an endless stream, and network robots are becoming increasingly rampant. Abnormal behaviors are likely to cause network hijacking, service paralysis, high server load costs, etc. for enterprises. A detection method or system for high- and low-frequency multi-level DNS anomalies is proposed. By comprehensively detecting DNS static domain name anomalies and dynamic traffic anomalies, it can achieve overall defense against DNS data.

[0092] For domain name anomaly detection, the blacklist mechanism can directly block abnormal domain names, and the reverse DGA technology realizes the cracking of the DGA generation algorithm, and can achieve 100% defense against abnormal domain names that conform to the generation rules; the differential detection of the detection model and multi-type statistical indicators can perform anomaly scoring on domain names from many dimensions. This comprehensive mechanism can greatly reduce the misjudgment rate of domain name anomalies and apply the anomaly score to the DNS traffic anomaly detection in different frequency bands and different situations.

[0093] The XGBoost detection model proposed for large-scale DNS real-time high-frequency traffic anomaly detection can achieve efficient detection of real-time traffic with advantages such as optimized feature dimensions, domain name scoring, and parallel computing. The detection effect and real-time performance are excellent, and it can detect anomalies in real time; in the medium-frequency traffic detection, the entropy value within the sliding time window is introduced to detect whether the traffic fluctuation under the IP portrait is abnormal, and it can detect and identify multi-type traffic robots; for low-frequency DNS data, a rule engine is constructed based on statistical analysis and expert experience for matching and filtering, realizing the tracking detection of low-frequency suspicious traffic data. Through correlation analysis based on traffic behavior data, anomalies can be fully mined, so as to achieve multi-band all-round attack defense and anomaly detection for DNS data.

[0094] Through the embodiments of the present application, the workload of network security personnel is greatly reduced, and automated and process-based defense, detection, and filtering are realized, which can save a large amount of manpower and server operation costs, and comprehensively maintain the network security of enterprises. Reduce the misjudgment rate of domain name anomalies, improve the accuracy and efficiency of DNS anomaly behavior detection and identification, save costs, and improve network security.

[0095] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a DNS anomaly detection device provided by the embodiments of the present application.

[0096] As Figure 4 shown, the DNS anomaly detection device 200 includes: a domain name acquisition module 201, a first anomaly detection module 202, a second anomaly detection module 203, a target anomaly evaluation module 204, and an anomaly data determination module 205.

[0097] The domain name acquisition module 201 is configured to acquire the domain names of multiple DNS data packets, and determine the access volume of each of the domain names according to the domain names of the multiple DNS data packets;

[0098] The first anomaly detection module 202 is configured to perform anomaly detection on the domain name of each of the DNS data packets to obtain a first anomaly probability of each of the DNS data packets;

[0099] The second anomaly detection module 203 is configured to perform anomaly detection on each of the domain names according to the access volume of each of the domain names to obtain a second anomaly probability corresponding to each of the domain names;

[0100] The target anomaly evaluation module 204 is configured to determine target anomaly evaluation information of each of the DNS data packets based on the first anomaly probability of each of the DNS data packets and the second anomaly probability corresponding to each of the domain names;

[0101] The anomaly data determination module 205 is configured to determine target DNS data packets with anomalies from the multiple DNS data packets according to the target anomaly evaluation information of each of the DNS data packets.

[0102] In one embodiment, as Figure 5 shown, the first anomaly detection module 202 includes:

[0103] The anomaly detection sub-module 2021 is configured to perform anomaly detection on the domain name of each of the DNS data packets based on a preset domain name anomaly detection model to obtain a first probability that each of the DNS data packets has an anomaly;

[0104] The anomaly analysis sub-module 2022 is configured to perform anomaly analysis on the domain name of each of the DNS data packets based on a preset domain name statistical analysis algorithm to obtain a second probability that each of the DNS data packets has an anomaly;

[0105] The anomaly determination sub-module 2023 is configured to determine the first anomaly probability of each of the DNS data packets according to the first probability and the second probability that each of the DNS data packets has an anomaly.

[0106] In one embodiment, the first anomaly detection module 202 is further configured to:

[0107] Perform character conversion on the domain name of the DNS data packet to obtain a first domain name;

[0108] Perform TF-IDF conversion processing on the first domain name to obtain a second domain name, and calculate the mean and variance values between the second domain name and multiple DGA domain names;

[0109] Determine the second probability that the DNS data packet corresponding to the second domain name is abnormal according to the mean and variance values between the second domain name and multiple DGA domain names.

[0110] In one embodiment, as Figure 6 shown, the second anomaly detection module 203 includes:

[0111] A traffic level determination sub-module 2031, configured to determine the traffic level corresponding to each domain name according to the access volume of each domain name;

[0112] A traffic anomaly detection sub-module 2032, configured to perform traffic anomaly detection on each domain name according to the traffic level corresponding to each domain name, and obtain a second anomaly probability corresponding to each domain name.

[0113] In one embodiment, the traffic levels include a first traffic level, a second traffic level, and a third traffic level. The access volume corresponding to the first traffic level is greater than the access volume corresponding to the second traffic level, and the access volume corresponding to the second traffic level is greater than the access volume corresponding to the third traffic level; the second anomaly detection module 203 is further configured to:

[0114] The performing traffic anomaly detection on each domain name according to the traffic level corresponding to each domain name, and obtaining a second anomaly probability corresponding to each domain name includes:

[0115] Based on a preset traffic anomaly detection model, perform anomaly detection on the domain names corresponding to the first traffic level to obtain a first anomaly detection result;

[0116] Based on a preset time window algorithm, perform anomaly detection on the domain names corresponding to the second traffic level to obtain a second anomaly detection result;

[0117] Based on a preset association analysis algorithm, perform anomaly analysis on the domain names corresponding to the third traffic level to obtain a third anomaly detection result;

[0118] Determine the second anomaly probability corresponding to each domain name according to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result.

[0119] In one embodiment, the second anomaly detection module 203 is further configured to:

[0120] Call the XGBoost detection model corresponding to the first traffic level;

[0121] Use the XGBoost detection model to perform anomaly detection on the domain names corresponding to the first traffic level, and obtain the anomaly probabilities of multiple domain names corresponding to the first traffic level;

[0122] Use the abnormal probabilities of multiple domain names corresponding to the first traffic level as the first abnormal detection result.

[0123] In one embodiment, the target abnormal evaluation module 204 is further configured to:

[0124] Determine the third abnormal probability of multiple DNS data packets according to the second abnormal probability corresponding to each domain name; the third abnormal probability of the DNS data packet matches the second abnormal probability corresponding to the domain name of the DNS data packet;

[0125] Determine the target abnormal evaluation information of each DNS data packet according to the first abnormal probability and the third abnormal probability of each DNS data packet;

[0126] Wherein, the target abnormal evaluation information includes a target abnormal probability, a target abnormal score, and / or a target abnormal rating.

[0127] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing embodiments of the DNS abnormal detection method, and will not be elaborated herein.

[0128] The device provided in the above embodiment can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 7 shown.

[0129] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device can be a server or a terminal device.

[0130] As shown in Figure 7 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a storage medium and an internal memory, and the storage medium can be non-volatile or volatile.

[0131] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any DNS abnormal detection method.

[0132] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0133] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any DNS abnormal detection method.

[0134] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 7 the structure shown in

[0135] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0136] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0137] Obtain the domain names of multiple DNS data packets, and determine the access volume of each domain name according to the domain names of the multiple DNS data packets;

[0138] Perform anomaly detection on the domain name of each DNS data packet to obtain the first anomaly probability of each DNS data packet;

[0139] Perform anomaly detection on each domain name according to the access volume of each domain name to obtain the second anomaly probability corresponding to each domain name;

[0140] Based on the first anomaly probability of each DNS data packet and the second anomaly probability corresponding to each domain name, determine the target anomaly evaluation information of each DNS data packet;

[0141] According to the target anomaly evaluation information of each DNS data packet, determine the target DNS data packet with anomalies from the multiple DNS data packets.

[0142] In one embodiment, when the processor implements the anomaly detection on the domain name of each DNS data packet to obtain the first anomaly probability of each DNS data packet, it is used to implement:

[0143] Based on a preset domain name anomaly detection model, perform anomaly detection on the domain name of each of the DNS data packets to obtain a first probability of anomaly for each of the DNS data packets.

[0144] Based on a preset domain name statistical analysis algorithm, perform anomaly analysis on the domain name of each of the DNS data packets to obtain a second probability of anomaly for each of the DNS data packets.

[0145] Determine a first anomaly probability for each of the DNS data packets according to the first probability and the second probability of anomaly for each of the DNS data packets.

[0146] In one embodiment, when the processor implements performing anomaly analysis on the domain name of each of the DNS data packets based on the preset domain name statistical analysis algorithm to obtain a second probability of anomaly for each of the DNS data packets, it is used to implement:

[0147] Perform character conversion on the domain name of the DNS data packet to obtain a first domain name.

[0148] Perform TF-IDF conversion processing on the first domain name to obtain a second domain name, and calculate the mean and variance values between the second domain name and multiple DGA domain names.

[0149] Determine a second probability of anomaly for the DNS data packet corresponding to the second domain name according to the mean and variance values between the second domain name and multiple DGA domain names.

[0150] In one embodiment, when the processor implements performing anomaly detection on each of the domain names according to the access volume of each of the domain names to obtain a second anomaly probability corresponding to each of the domain names, it is used to implement:

[0151] Determine a traffic level corresponding to each of the domain names according to the access volume of each of the domain names.

[0152] Perform traffic anomaly detection on each of the domain names according to the traffic level corresponding to each of the domain names to obtain a second anomaly probability corresponding to each of the domain names.

[0153] In one embodiment, the traffic level includes a first traffic level, a second traffic level, and a third traffic level, the access volume corresponding to the first traffic level is greater than the access volume corresponding to the second traffic level, and the access volume corresponding to the second traffic level is greater than the access volume corresponding to the third traffic level.

[0154] When the processor implements performing traffic anomaly detection on each of the domain names according to the traffic level corresponding to each of the domain names to obtain a second anomaly probability corresponding to each of the domain names, it is used to implement:

[0155] Based on a preset traffic anomaly detection model, perform anomaly detection on the domain names corresponding to the first traffic level to obtain a first anomaly detection result;

[0156] Based on a preset time window algorithm, perform anomaly detection on the domain names corresponding to the second traffic level to obtain a second anomaly detection result;

[0157] Based on a preset correlation analysis algorithm, perform anomaly analysis on the domain names corresponding to the third traffic level to obtain a third anomaly detection result;

[0158] According to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result, determine the second anomaly probability corresponding to each of the domain names.

[0159] In one embodiment, when the processor implements performing anomaly detection on the domain names corresponding to the first traffic level based on a preset traffic anomaly detection model to obtain a first anomaly detection result, it is used to implement:

[0160] Invoke the XGBoost detection model corresponding to the first traffic level;

[0161] Use the XGBoost detection model to perform anomaly detection on the domain names corresponding to the first traffic level to obtain the anomaly probabilities of multiple domain names corresponding to the first traffic level;

[0162] Take the anomaly probabilities of multiple domain names corresponding to the first traffic level as the first anomaly detection result.

[0163] In one embodiment, when the processor implements determining the target anomaly evaluation information for each DNS packet based on the first anomaly probability of each DNS packet and the second anomaly probability corresponding to each domain name, it is used to implement:

[0164] Determine the third anomaly probability of multiple DNS packets according to the second anomaly probability corresponding to each domain name; the third anomaly probability of the DNS packet matches the second anomaly probability of the domain name corresponding to the DNS packet;

[0165] Determine the target anomaly evaluation information for each DNS packet according to the first anomaly probability and the third anomaly probability of each DNS packet;

[0166] Wherein, the target anomaly evaluation information includes a target anomaly probability, a target anomaly score, and / or a target anomaly rating.

[0167] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described computer device can refer to the corresponding process in the foregoing embodiments of the DNS anomaly detection method, and will not be elaborated herein.

[0168] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0169] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program includes program instructions, and the method implemented when the program instructions are executed can refer to the respective embodiments of the DNS anomaly detection method of the present application.

[0170] Among them, the computer-readable storage medium can be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0171] Furthermore, the computer-usable storage medium may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of blockchain nodes, etc. The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0172] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0173] It should also be understood that the term "and / or" used in this application specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0174] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A DNS anomaly detection method, characterized in that, Including: Obtain the domain names of multiple DNS packets, and determine the access volume of each domain name according to the domain names of the multiple DNS packets; Based on a preset domain name anomaly detection model, perform anomaly detection on the domain names of each DNS packet to obtain the first probability of anomaly for each DNS packet; Based on a preset domain name statistical analysis algorithm, perform anomaly analysis on the domain names of each DNS packet to obtain the second probability of anomaly for each DNS packet; Determine the first anomaly probability of each DNS packet according to the first probability and the second probability of anomaly for each DNS packet; Determine the traffic level corresponding to each domain name according to the access volume of each domain name; the traffic levels include a first traffic level, a second traffic level, and a third traffic level, the access volume corresponding to the first traffic level is greater than the access volume corresponding to the second traffic level, and the access volume corresponding to the second traffic level is greater than the access volume corresponding to the third traffic level; Based on a preset traffic anomaly detection model, perform anomaly detection on the domain names corresponding to the first traffic level to obtain a first anomaly detection result; Based on a preset time window algorithm, perform anomaly detection on the domain names corresponding to the second traffic level to obtain a second anomaly detection result; Based on a preset association analysis algorithm, perform anomaly analysis on the domain names corresponding to the third traffic level to obtain a third anomaly detection result; Determine the second anomaly probability corresponding to each domain name according to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result; Based on the first anomaly probability of each DNS packet and the second anomaly probability corresponding to each domain name, determine the target anomaly evaluation information of each DNS packet; Determine the target DNS packets with anomalies from the multiple DNS packets according to the target anomaly evaluation information of each DNS packet.

2. The DNS anomaly detection method according to claim 1, characterized in that, The step of performing anomaly analysis on the domain names of each DNS packet based on a preset domain name statistical analysis algorithm to obtain the second probability of anomaly for each DNS packet includes: Perform character conversion on the domain names of the DNS packets to obtain a first domain name; Perform TF-IDF conversion processing on the first domain name to obtain a second domain name, and calculate the mean and variance values between the second domain name and multiple DGA domain names; Determine the second probability of anomaly for the DNS packet corresponding to the second domain name according to the mean and variance values between the second domain name and multiple DGA domain names.

3. The DNS anomaly detection method according to claim 1, wherein The step of performing anomaly detection on the domain names corresponding to the first traffic level based on a preset traffic anomaly detection model to obtain a first anomaly detection result includes: Invoke the XGBoost detection model corresponding to the first traffic level; Use the XGBoost detection model to perform anomaly detection on the domain names corresponding to the first traffic level to obtain the anomaly probabilities of multiple domain names corresponding to the first traffic level; Use the anomaly probabilities of multiple domain names corresponding to the first traffic level as the first anomaly detection result.

4. The DNS anomaly detection method according to claim 1, wherein, Determining the target anomaly evaluation information for each DNS packet based on the first anomaly probability of each DNS packet and the second anomaly probability corresponding to each domain name includes: Determining a third anomaly probability for a plurality of the DNS packets according to the second anomaly probability corresponding to each domain name; the third anomaly probability of the DNS packet matches the second anomaly probability corresponding to the domain name of the DNS packet; Determining the target anomaly evaluation information for each DNS packet according to the first anomaly probability and the third anomaly probability of each DNS packet; Wherein, the target anomaly evaluation information includes a target anomaly probability, a target anomaly score, and / or a target anomaly rating.

5. A DNS anomaly detection device, characterized in that, The DNS anomaly detection device includes: A domain name acquisition module, configured to acquire the domain names of a plurality of DNS packets, and determine the access volume of each domain name according to the domain names of the plurality of DNS packets; A first anomaly detection module, configured to perform anomaly detection on the domain name of each DNS packet based on a preset domain name anomaly detection model to obtain a first probability that each DNS packet has an anomaly; perform anomaly analysis on the domain name of each DNS packet based on a preset domain name statistical analysis algorithm to obtain a second probability that each DNS packet has an anomaly; determine the first anomaly probability of each DNS packet according to the first probability and the second probability that each DNS packet has an anomaly; A second anomaly detection module, configured to determine the traffic level corresponding to each domain name according to the access volume of each domain name, where the traffic levels include a first traffic level, a second traffic level, and a third traffic level, the access volume corresponding to the first traffic level is greater than the access volume corresponding to the second traffic level, and the access volume corresponding to the second traffic level is greater than the access volume corresponding to the third traffic level; perform anomaly detection on the domain names corresponding to the first traffic level based on a preset traffic anomaly detection model to obtain a first anomaly detection result; perform anomaly detection on the domain names corresponding to the second traffic level based on a preset time window algorithm to obtain a second anomaly detection result; perform anomaly analysis on the domain names corresponding to the third traffic level based on a preset association analysis algorithm to obtain a third anomaly detection result; determine the second anomaly probability corresponding to each domain name according to the first anomaly detection result, the second anomaly detection result, and the third anomaly detection result; A target anomaly evaluation module, configured to determine the target anomaly evaluation information for each DNS packet based on the first anomaly probability of each DNS packet and the second anomaly probability corresponding to each domain name; An anomaly data determination module, configured to determine target DNS packets with anomalies from the plurality of DNS packets according to the target anomaly evaluation information of each DNS packet.

6. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the DNS anomaly detection method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the DNS anomaly detection method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Network service abnormal data detection method and device, equipment and medium

    CN111400126A

  • DNS anomaly detection method, device and equipment and storage medium

    CN111565187A