A DGA Domain Name Detection Method for IoT Botnets

Through the fusion neural network model of Small BERT and CNN, the problem of low word splicing detection performance and imbalance of multi-classification tasks in DGA domain name detection is solved, and efficient DGA domain name detection is achieved in the Internet of Things environment.

CN116633623BActive Publication Date: 2025-07-04BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310597905.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-07-04
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

In the prior art, the DGA domain name detection method has low performance on word splicing DGA domain name detection, unbalanced performance of multi-classification tasks, and is not suitable for Internet of Things environment.

Method used

The fusion neural network model of Small BERT and CNN is adopted to extract domain name features through word embedding layer and character embedding layer, and combine the semantics of Small BERT and the character randomness characteristics of CNN to perform binary classification and multi-classification of DGA domain names.

Benefits of technology

It improves the detection performance of random word type DGA domain names, improves the imbalance problem of multi-classification tasks, and is suitable for IoT environments with limited computing resources. It has the advantages of high detection performance, easy deployment and wide application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116633623B_ABST
    Figure CN116633623B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of Internet of Things security technology, and discloses a DGA domain name detection method for the Internet of Things botnet, including: obtaining the domain name string in the DNS domain name resolution request sent by the Internet of Things device; inputting the domain name string into the trained domain name detection model to obtain the binary classification and multi-classification results of the domain name; wherein, the domain name detection model is obtained by training based on the fusion network of the Small BERT pre-trained model and CNN. By combining the two feature extraction capabilities of the SmallBERT model at the sub-word granularity for domain names and CNN at the character granularity for domain names, the present invention improves the learning ability of the algorithm for the comprehensive features of domain name semantics, morphology, pronunciation, character randomness, etc., enhances the detection performance for random word type DGA domain names, improves the problem of performance imbalance in the domain name multi-classification task, meets the requirements of limited computing resources and high inference speed in the Internet of Things environment, and has the advantages of high detection performance, simple process, easy deployment, and wide available range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things security detection, and particularly to a DGA domain name detection method and system. Background Art

[0002] A botnet is a network formed by a large number of computers infected with malware. This network can be used by the controller to perform various malicious network activities, such as launching DDoS attacks, spreading spam, stealing user information, etc. The Domain Generation Algorithm (DGA) is one of the technologies for constructing the command and control channels of botnets. DGA domain names refer to malicious domain names generated by the DGA algorithm. The botnet controller integrates the DGA algorithm into the malware, dynamically generates a large number of DGA domain names, and randomly selects a part of them as the addresses of the control servers, thereby realizing the control of the zombie machines. Antivirus software and network security devices usually rely on known malicious domain names or IP addresses to identify and block malicious traffic. Therefore, detecting and blocking DGA domain names requires more advanced security defense measures and technologies.

[0003] The initial DGA algorithm adopted relatively simple random number generation methods, such as generating random numbers based on timestamps, hardware information, fixed seed values, etc., and then combining the random numbers with a preset string to generate a domain name list. With the development of network attack and defense technologies, many new algorithms and improvement methods have been proposed. For example, consonants and vowels are combined to generate syllables, and then multiple syllables are randomly concatenated to generate pronounceable domain names. There is also the method of randomly concatenating words in a dictionary to generate DGA domain names with a high similarity to legitimate domain names. The existing technology has low detection performance for word-concatenated DGA domain names, which is one of the challenges faced by DGA domain name detection technologies.

[0004] In the era of artificial intelligence, DGA domain name detection algorithms are mainly divided into two categories: machine learning methods based on artificial feature extraction and deep learning methods based on non-feature extraction. Deep learning methods can reduce the steps of artificial feature extraction and have achieved great success in the field of natural language processing. In recent years, they have been widely applied to DGA domain name detection methods. In the DGA domain name detection methods based on deep learning, in the early method research, recurrent neural networks were mostly used as the basic framework, which had good detection performance for DGA domain names with high randomness, but had poor detection performance for DGA domain names generated by word concatenation. Later, attention mechanisms were gradually fused with recurrent neural networks, convolutional neural networks, and methods of mixing multiple networks such as Transformer were adopted to improve the effect of DGA domain name multi-classification tasks. Although the overall detection performance has been improved, there is always a problem of unbalanced detection performance for the domain name multi-classification tasks of more than 70 DGA families.

[0005] In addition, in the current environment where the number of Internet of Things (IoT) devices has exploded, IoT devices are vulnerable to botnets due to various software vulnerabilities and other factors. IoT botnets have also developed into a major threat in the current field of IoT security, having a direct negative impact on people's living and production activities. In particular, in IoT environments with a large amount of device assets and a high degree of data sensitivity, such as industrial IoT, smart city IoT, and medical IoT, the security of their devices and data has received more attention. However, high-performance DGA domain name detection methods applicable to the actual IoT application environment still need to be studied and proposed. Summary of the Invention

[0006] The present invention provides a DGA domain name detection method for IoT botnets to solve the problems of low detection performance for word-concatenated DGA domain names and unbalanced performance in multi-classification tasks of DGA domain names in the prior art, as well as the defect of being inapplicable to the IoT environment.

[0007] The DGA domain name detection method for IoT botnets includes:

[0008] Obtain the domain name string in the DNS domain name resolution request sent by the IoT device;

[0009] Input the domain name string into the trained domain name detection model to obtain the binary classification and multi-classification results of the domain name; wherein, the domain name detection model is obtained by training based on a fusion neural network of Small BERT and CNN.

[0010] Further, the domain name detection model is obtained through the following steps:

[0011] Obtain an open-source DGA domain name dataset and a legitimate domain name dataset;

[0012] Divide the DGA families in the DGA domain name dataset into three categories: random character type, random syllable type, and random word type;

[0013] Perform binary classification annotation of DGA domain names and legitimate domain names on the domain name dataset respectively, as well as multi-classification annotation of the above three types of DGA domain names and legitimate domain names;

[0014] Construct a domain name classifier with a fusion neural network structure of Small BERT and CNN based on a deep learning framework;

[0015] Train the domain name classifier based on the processed domain name dataset to obtain the domain name detection model.

[0016] Furthermore, the fusion neural network structure of Small BERT and CNN specifically includes: a word embedding layer, a character embedding layer, a Small BERT sub-network, a CNN sub-network, a pooling layer, and a fully connected layer. The domain name text is input into the word embedding layer and the character embedding layer simultaneously. The word embedding layer splits the domain name text into multiple sub-words based on the wordpiece algorithm and encodes them into vectors, and the character embedding layer splits the domain name into single characters and encodes them into vectors based on the one-hot algorithm. The output of the word embedding layer serves as the input of the Small BERT sub-network, the output of the character embedding layer serves as the input of the CNN sub-network, and the outputs of the Small BERT and CNN sub-networks are merged and used as the input of the fully connected layer, finally obtaining the binary classification and multi-classification results of the domain name.

[0017] The DGA domain name detection method for the Internet of Things zombie network provided by the present invention introduces the Small BERT pre-trained language model, enhances the algorithm's ability to extract features such as domain name semantics and grammar, can effectively improve the detection performance of random word type DGA domain names, combines the CNN network's ability to extract features such as domain name morphology and character randomness, can improve the problem of unbalanced performance in the domain name multi-classification task, and at the same time meets the requirements of limited computing resources and high real-time inference speed in the Internet of Things application environment, having the advantages of high detection performance, simple process, easy deployment, and wide available range. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0019] Figure 1 is a schematic flowchart of the DGA domain name detection method shown according to an exemplary embodiment.

[0020] Figure 2 is a flowchart of the algorithm training of the DGA domain name detection method shown according to an exemplary embodiment.

[0021] Figure 3 is a flowchart of obtaining and processing the training sample data set shown according to an exemplary embodiment.

[0022] Figure 4 is a schematic diagram of the structure of the training sample data set shown according to an exemplary embodiment.

[0023] Figure 5 is a schematic diagram of the training sample data label shown according to an exemplary embodiment.

[0024] Figure 6 is a schematic diagram of the DGA domain name classification neural network structure provided by the present invention.

[0025] Figure 7 The process diagram shows the WordPiece tokenization of the word embedding layer according to an exemplary embodiment, taking "equal-future223.top" as an example. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the exemplary embodiments will be described in detail below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] In view of various problems existing in the prior art, the present invention combines a domain name text representation method based on character embedding and a domain name text representation method based on word embedding, introduces the excellent Small BERT pre-training model in the field of natural language processing, and fuses it with the CNN neural network structure to design a DGA domain name detection algorithm that can be used in the Internet of Things environment, has adjustable size, and improves detection performance.

[0029] Regarding the real-time application problems of the Internet of Things, compared with personal computers and servers, Internet of Things devices usually have fewer computing, storage, and network resources. The problem of limited computing performance of Internet of Things devices makes it impossible for deep learning models with high computing requirements to run on Internet of Things devices. Therefore, the size of the computing resources required for model inference affects whether the model can be applied to the Internet of Things environment.

[0030] For the multi-classification problem of DGA domain names, the multi-classification task of domain names has not achieved very satisfactory results in previous research work. For example, in the model that uses cost-sensitive LSTM to improve the imbalance of multi-classification results, the experimental result data shows that there is a large gap in the detection effect of the model for multiple DGA domain name types. The detection effect of this algorithm for random character-based DGA domain names is significantly better than that of random syllable-based and random word-based ones, and the detection effect for random word-based DGA domain names is the worst. According to the respective feature analyses of DGA domain names and legitimate domain names, the possible reasons for the poor performance of the DGA domain name multi-classification task lie in many factors such as the large number of feature dimensions of legitimate domain names, the continuous upgrade of DGA algorithms, the large number of DGA families, and the significant difference in the number of samples available for training among different DGA families.

[0031] For the detection problem of word-concatenated DGA domain names, the DGA algorithm with randomly combined characters was the first to appear and was used in the command and control channels of botnets. The word-concatenated DGA algorithm emerged in recent years. As the detection accuracy of traditional random character-based DGA domain names by researchers continues to improve, in order to improve the concealment and anti-strike ability of botnets, botmasters will increasingly consider using word-concatenated DGA to construct the command and control channels of botnets. Therefore, improving the detection performance for word-concatenated DGA domain names is a problem that needs to be solved.

[0032] To improve the above problems, the Small BERT pre-trained language model introduced in the present invention has a strong ability to capture features such as semantics and grammar of domain name texts, and can effectively improve the detection performance for random word-based DGA domain names. The CNN neural network structure has a good ability to capture features such as the randomness of domain name characters and morphology. The network structure integrating Small BERT and CNN combines the feature capture abilities of the two, can improve the problem of unbalanced domain name multi-classification tasks, and at the same time meet the requirements of limited computing resources and high real-time inference speed in the Internet of Things application environment.

[0033] Correspondingly, Figure 1 is a flowchart of a DGA domain name detection method for IoT botnets provided by the present invention, as Figure 1 shown, including the following steps:

[0034] S1, obtain the domain name string in the DNS domain name resolution request sent by the IoT device;

[0035] S2, input the domain name string into the pre-trained DGA domain name detection model to obtain the binary classification and multi-classification results of the domain name.

[0036] As can be seen from the above embodiments, the present application can effectively utilize the feature capture capabilities of the Small BERT pre-trained model and the CNN neural network, improve the detection ability of word-concatenated DGA domain names, and enhance the overall performance of domain name multi-classification. At the same time, it balances model performance and computational efficiency and is applicable to the Internet of Things application environment.

[0037] In the above step S2, the DGA domain name detection model is obtained through machine learning training by a fusion neural network of Small BERT and CNN. Refer to Figure 2 , the training process of the DGA domain name detection algorithm for the Internet of Things may include the following sub-steps:

[0038] S21, obtain and process the training sample data set;

[0039] S22, construct the DGA domain name detection algorithm based on the machine learning framework;

[0040] S23, perform N rounds of training on the DGA domain name detection algorithm based on the processed training data set.

[0041] In the above step S21, for the specific process of obtaining and processing the training sample data set, refer to Figure 3 , it may include the following steps:

[0042] S211, obtain the open-source data sets of DGA domain names: UMUDGA data set and UTL_DGA22 data set;

[0043] Specifically, the UMUDGA data set selects 50 important ones from all DGA families, and the data collection time is 2019; UTL_DGA22 deletes duplicate DGA families and adds new botnet families, and a total of 76 DGA family domain name data are included, and the data collection time is 2022.

[0044] S212, divide the DGA families in the data set into three categories: random character type, random syllable type, and random word type.

[0045] For dividing the DGA families into three categories, previously in the multi-classification task of DGA domain names, the classification was performed with a single DGA family as the classification granularity, that is, to determine which DGA family a certain DGA domain name belongs to. However, in the actual application scenario of DGA domain name detection for the Internet of Things botnet, it is not suitable to perform multi-classification detection of domain name text with the DGA family as the granularity. The main reasons are as follows:

[0046] First, there are no less than 76 known DGA families and their variants worldwide. According to the current research status at home and abroad, the actual performance of the DGA domain name multi-classification method based on deep learning is obviously not enough to reach the level of production system. Most DGA domain name multi-classification methods have obvious shortcomings in the detection performance of some DGA families. There is no method that can achieve an F1 value of more than 90% for the classification of all DGA families, especially for word-concatenated DGA domain names.

[0047] Second, it is known that some DGA families have derivative relationships, and many DGA families have high similarities in terms of algorithm logic. For example, the two DGA families proslikefan and pykspa both use a random combination of English letters az to generate domain names, but the domain name length and the top-level domain name used are different. Classifying DGA families for these domain names with similar features not only brings high detection difficulty, but also its classification results do not have very great practical application significance at the current stage.

[0048] Therefore, in order to make the DGA domain name multi-classification detection results based on the deep learning algorithm more credible and convenient for use in the production environment, based on the research on the text structure characteristics of the DGA domain name and the comprehensive multi-dimensional characteristics of the DGA domain name, this embodiment divides all DGA family domain names into three types: random character type, random syllable type, and random word type.

[0049] Specifically, the characteristics of random character DGA family domain names are that they are a random combination of English letters or English letters plus numbers, and the domain names are not semantic and pronounceable, such as the domain names f0fe6744.top and gfedo.info. The number of random character DGA families is the largest, including bedep, qakbot, ramdo, murofet, tempedreve, chinad, kraken, nymaim, ccleaner, ramnit, etc.

[0050] The characteristics of random syllable DGA family domain names are that English vowels and consonants are combined into a readable syllable, and then multiple readable syllables are combined into a DGA domain name. This type of DGA domain name is pronounceable and close to the characteristics of real domain names, but it does not have semantics, such as the domain name sefigwirud.com. Random syllable DGA families account for the smallest proportion of all DGA families, including pitou, pushdo, simda, symmi, vawtrak_v2, vawtrak_v3, pykspa, pykspa_noise, etc.

[0051] The characteristics of the random-word DGA family domain names are that they are composed of multiple words combined from multiple word lists, and are closest to the characteristics of real domain names. Due to the randomness of the constituent words, the semantic characteristics of this type of domain names are weaker than those of legitimate domain names, and their detection difficulty is the highest. For example, the domain name missionmore.com. The number of random-word DGA families is small, including DGA families such as bigviktor, matsnu, ngioweb, rovnix, suppobox_1, suppobox_2, suppobox_3, gozi_gpl, gozi_luther, gozi_nasa, etc.

[0052] In step S212, all DGA families in the DGA domain name dataset obtained in step S211 are classified according to the characteristics of the above three types of DGA family domain names.

[0053] S213, extract DGA family domain name data from 2 datasets according to the ratio of 2:1:1;

[0054] Specifically, refer to Figure 4 , extract all random-word DGA families and random-syllable DGA families from the UTL_DGA22 dataset and the UMUDGA dataset respectively, and extract part of the domain name data of the random-character DGA families from the UTL_DGA22 dataset. In this embodiment, 48 DGA families are extracted based on two open-source datasets, generating 480,000 DGA domain name data, and the ratio of the number of samples of random-character, random-syllable, and random-word DGA domain names is 2:1:1.

[0055] S214, obtain the legal domain name open-source dataset Tranco;

[0056] For the legal domain name dataset, in previous studies on DGA domain name detection, most experiments used the ranking list provided by the Alexa website. However, Victor Le Pochat et al. pointed out in their research that the authenticity of the list data ranked based on popularity on websites such as Alexa is questionable. They found that these rankings can be easily changed by constructing HTTP requests, so there may be malicious behaviors of large-scale manipulation of website rankings to make the results of research work based on these ranking data conform to the wishes of hackers. To enable the research community to use reliable ranking data, Victor Le Pochat et al. provided the Tranco dataset, which contains the data of the top 1 million websites in the improved rankings. In this embodiment, 240,000 domain name data in the Tranco dataset are extracted as legal domain name data.

[0057] S215, fuse the DGA domain name dataset and the legal domain name dataset;

[0058] S216, Label the domain name sample data with binary classification and multi-classification labels.

[0059] Specifically, referring to Figure 5 , labeling the sample with labels specifically includes 2 sub-steps:

[0060] First, label the sample data with binary classification labels. The binary classification label for legitimate domain names is 0, and the binary classification label for all DGA domain names is 1;

[0061] Then, label the sample data with multi-classification labels. The multi-classification label for legitimate domain names is 0, the multi-classification label for random character-based DGA domain names is 1, the multi-classification label for random syllable-based DGA domain names is 2, and the multi-classification label for random word-based DGA domain names is 3.

[0062] In this embodiment, the neural network structure of the DGA domain name detection algorithm for the IoT botnet includes a word embedding layer, a character embedding layer, a Small BERT layer, a one-dimensional CNN layer, a pooling layer, a fully connected layer, and an output layer; the Small BERT layer is used to extract features such as semantics and grammar in the domain name text; the one-dimensional CNN layer is used to extract features such as character randomness and morphology in the domain name text; the output layer is used to output the binary classification and multi-classification probability values of the domain name.

[0063] Figure 6 It is a schematic diagram of the neural network structure for classifying DGA domain names for the IoT botnet provided by the present invention. Referring to Figure 6 , the connection method of the neural network for classifying DGA domain names for the IoT botnet is as follows: The domain name string is input into the word embedding layer and the character embedding layer at the same time. The output of the word embedding layer is used as the input of the Small BERT layer, and the output of the character embedding layer is used as the input of the CNN sub-network; in the CNN sub-network, the output of the one-dimensional CNN layer is used as the input of both the max pooling layer and the average pooling layer at the same time. The outputs of the max pooling layer and the average pooling layer are used as the input of the fully connected layer; then, the output of the Small BERT layer is merged with the output of the CNN sub-network and used as the input of another fully connected layer. The output of this fully connected layer finally passes through a fully connected layer to output the classification result of the domain name.

[0064] Specifically, the word embedding layer converts the domain name text into a data format available for the Small BERT layer, and this process includes two sub-steps:

[0065] First, use the WordPiece tokenization algorithm to tokenize the domain name text.

[0066] Specifically, the WordPiece tokenization algorithm can split the input text into a set of words or sub-words and assign a unique identifier to each word or sub-word. The algorithm has two implementation goals. One is to tokenize the text data into as few chunks as possible. The other is to split a word into the sub-words with the largest count in the training data when it has to be chunked. For example, for the domain name "equal-future223.top", the tokenization result is as Figure 7 shown.

[0067] Secondly, after the domain name string is tokenized by the WordPiece algorithm, each word or sub-word is mapped to a fixed-length vector. In the BERT model, the vocabulary size of the WordPiece tokenization algorithm is usually set to 30,000 words or more.

[0068] For the implementation of the word embedding layer, the TensorFlow framework provides a preprocessing layer that adapts the text embedding for each BERT model.

[0069] The character embedding layer uses the one-hot representation method for the character-level representation of the domain name text.

[0070] The domain name consists of characters such as lowercase letters, numbers, dashes, underscores, and English full stops, a total of 38 characters. A character mapping table is established for these 38 domain name characters plus the padding character, and each character corresponds to a unique integer identifier. Each integer identifier is converted into a one-hot vector, where the length of the vector is equal to the length of the character mapping table, and only the position corresponding to the integer identifier in the vector is 1, and the other positions are 0. The vectors of each character of the input domain name are concatenated to form a matrix as the input of the model.

[0071] In this embodiment, the output dimension of the character embedding layer is 128, and the output dimension of the word embedding layer is 256.

[0072] The Small BERT layer is a lightweight BERT model proposed by Google in 2019. It has been tested on multiple natural language processing tasks and has been proven to have good results. The BERT model is based on the deep neural network architecture of Transformer and obtains rich language knowledge through unsupervised learning on large-scale unlabeled text corpora such as the entire Wikipedia and a large number of books, and has achieved excellent performance in various natural language processing tasks.

[0073] Compared with the original BERT model, Small BERT uses some optimization techniques to improve its structure, achieving a balance between performance and the number of parameters. The authors of Small BERT publicly provided 24 pre-trained Small BERT models that can be fine-tuned for downstream tasks. Among them, the smallest model has 2 Transformer layers, 128 hidden units per layer, and 2 attention heads, while the largest model has 6 Transformer layers, 768 hidden units per layer, and 8 attention heads. The base model of the BERT series contains 110M parameters, while the largest model in the Small BERT series contains 82M parameters, and the smallest model only contains 14.5M parameters.

[0074] Therefore, under the same computing resources, the Small BERT series models require less video memory and computing resources than the BERT series models, have a faster inference speed, and maintain high model performance. They are suitable for resource-constrained usage scenarios and can be applied in the Internet of Things environment.

[0075] In this embodiment, the Small BERT layer uses a Small BERT model provided by the TensorFlow framework with 2 layers, 256 hidden units, and 4 attention heads, and its output is a 256-dimensional pooled vector.

[0076] In the CNN sub-network of the DGA domain name detection and classification neural network, multiple one-dimensional convolutional layers, convolutional kernels of different sizes, and strides are used to extract n-gram features of different lengths in the domain name text sequence.

[0077] In this embodiment, the CNN layer uses 4 one-dimensional convolutional layers with convolutional kernel sizes of 2, 3, 2, and 3 respectively, strides of 1, 1, 2, and 2 respectively, and 64 convolutional kernels per layer.

[0078] Then, max-pooling and average-pooling are comprehensively used to extract features in the convolutional layer, and these features are concatenated into a long vector and input into the fully connected layer.

[0079] The role of the max-pooling layer is to retain the most significant features in each feature map while suppressing secondary features, which can improve the invariance of the model to transformations such as translation and rotation, thereby improving the generalization ability of the model. The role of the average-pooling layer is to smooth the feature map, reduce noise and redundant information in the feature map, shrink the size of the feature map, and reduce model parameters, thereby preventing overfitting.

[0080] After the output of the Small BERT layer and the output of the CNN sub-network are concatenated, they are connected to a 64-dimensional dense layer, and then respectively connected to two classification output layers.

[0081] In this embodiment, according to the above content, the TensorFlow machine learning framework is used to complete the construction of the DGA domain name detection algorithm for the Internet of Things botnet.

[0082] In step S23 above, based on the dataset processed in step S21, the algorithm constructed in step S22 is trained.

[0083] Specific algorithm training parameters include: the dataset is divided into a training set and a test set according to 9:1. Before each training round, the training set is shuffled, and then the training set is divided into the training set for this round and the validation set for this round according to 9:1; the SmallBERT network uses the AdamW optimizer with an initial learning rate of 2e-5; the convolutional neural network uses the Adam optimizer with an initial learning rate of 1e-3; the training batch size is 128, and the number of iteration rounds is 8; early termination is adopted when the training metric does not improve.

[0084] The evaluation metrics for the binary classification results of the model are recall, precision, F1, and loss value, and the recall and precision are used as the concatenated metrics for the multi-classification results.

[0085] Table 1 Binary Classification Experimental Results of Domain Names

[0086]

[0087] Table 1 shows the binary classification results of domain names. From the result data in Table 1, it can be seen that the DGA domain name detection algorithm for the Internet of Things botnet, which combines Small BERT and CNN networks, proposed in the present invention has the best performance in terms of recall, precision, and F1 value in the domain name binary classification task compared with other algorithms in the experiment.

[0088] Table 2 Multi-Class Classification Experimental Results of Domain Names

[0089]

[0090] Table 2 presents the experimental results of the DGA domain name multi-classification task. Among them, the neural network with a single Small BERT structure has the best performance in the classification tasks of random word type DGA domain names and random syllable type DGA domain names, far exceeding the model metric results of the LSTM network structure and the CNN-LSTM fusion structure; the detection performance of the Small BERT-CNN fusion model for these two types of DGA domain names is the second best, but it is also relatively close. It shows that using pre-trained language models such as Small BERT is very helpful for identifying random word type DGA domain names, and its semantic and morphological feature extraction capabilities precipitated after a large amount of text corpus training are much stronger than those of CNN and RNN models.

[0091] In addition, the Small BERT-CNN fusion model performs best in the recognition of random character-based DGA domain names and legitimate domain names, indicating that the CNN part in the fusion model can extract certain character randomness features, helping to improve the multi-classification ability of the model.

[0092] Overall, the model with the Small BERT-CNN fusion structure has achieved the best performance in the macro-average and micro-average of recall rate and precision rate. At the same time, the Small BERT-CNN fusion model also has leading performance in the binary classification task of domain names, proving the effectiveness of the Small BERT-CNN neural network model algorithm combined with word feature learning in the DGA domain name detection problem.

Claims

1. A method for detecting DGA domain names for IoT botnets, characterized in that Including: Step S1: Obtain the open-source DGA domain name datasets: UMUDGA dataset and UTL_DGA22 dataset. Classify the DGA families in the datasets into three categories: random character type, random syllable type, and random word type. Step S2: According to the above three types of DGA family domain names, extract the DGA family domain name data from the above two datasets according to the ratio of 2:1:

1. Then obtain the open-source legal domain name dataset, that is, the Tranco dataset. Among them, extract 240,000 domain name data from the Tranco dataset as legal domain name data. Integrate the DGA domain name dataset and the legal domain name dataset, and perform binary classification annotation of DGA domain names and legal domain names, as well as multi-classification annotation of the above three types of DGA domain names and legal domain names on the integrated domain name dataset. Step S3: Based on the deep learning framework, construct a domain name classifier that fuses the SmallBERT pre-training model and the CNN. The neural network connection method of this classifier is as follows: Input the domain name string into the word embedding layer and the character embedding layer at the same time. The output of the word embedding layer is used as the input of the SmallBERT layer, and the output of the character embedding layer is used as the input of the one-dimensional CNN sub-network. The output of the one-dimensional CNN layer is used as the input of the max pooling layer and the average pooling layer at the same time. The output features after the max pooling layer and the average pooling layer are concatenated into a long vector as the input of the fully connected layer. After the output of the SmallBERT layer and the output of the CNN sub-network are concatenated, they are connected to a 64-dimensional dense layer, and then respectively connected to two classification output layers. Step S4: Train the DGA domain name classifier in Step S3 based on the datasets obtained in Steps S1 and S2. Divide the training set and the test set according to 9:

1. Before each training round, shuffle the training set, and then divide the training set into the training set for this round and the validation set for this round according to 9:

1. The output of the classifier includes the binary classification result and the multi-classification result of the domain name.

2. The DGA domain name detection method for the IoT botnet according to claim 1, characterized in that The said Step S1 includes: The UMUDGA dataset selected 50 important ones from all DGA families, and the data collection time was 2019. The UTL_DGA22 deleted duplicate DGA families and added new botnet families, and a total of 76 DGA family domain name data were included, and the data collection time was 2022. The characteristics of the random character type DGA family domain name are that it is randomly combined by English letters or randomly combined by English letters and numbers, and the domain name has no semantics and pronounceability. The characteristics of the random syllable type DGA family domain name are that it is composed of English vowel letters and consonant letters to form a pronounceable syllable, and then multiple pronounceable syllables are combined to form a DGA domain name. This type of DGA domain name has pronounceability and is relatively close to the characteristics of real domain names, but has no semantics. The characteristics of the random word type DGA family domain name are that it is composed of multiple words in multiple word lists. It is the closest to the characteristics of real domain names. Due to the randomness of the composed words, the semantic characteristics of this type of domain name are weaker than those of legal domain names, and its detection difficulty is the highest.

3. The DGA domain name detection method for the IoT botnet according to claim 1, characterized in that, The said Step S2 includes: All random-word DGA families and random-syllable DGA families are extracted from the UTL_DGA22 dataset and the UMUDGA dataset respectively. Domain name data of some random-character DGA families are extracted from the UTL_DGA22 dataset. Based on two open-source datasets, 48 DGA families are extracted, generating 480,000 DGA domain name data. The ratio of the sample numbers of random-character, random-syllable, and random-word DGA domain names is 2:1:1; Binary classification labels are assigned to the sample data of the fused dataset. The binary classification label of legitimate domain names is 0, and the binary classification label of all DGA domain names is 1. Multi-classification labels are assigned to the sample data. The multi-classification label of legitimate domain names is 0, the multi-classification label of random-character DGA domain names is 1, the multi-classification label of random-syllable DGA domain names is 2, and the multi-classification label of random-word DGA domain names is 3.

4. The DGA domain name detection method for the IoT botnet according to claim 1, characterized in that, The step S3 includes: The word embedding layer uses the WordPiece tokenization algorithm to tokenize the domain name text, mapping each word or sub-word to a fixed-length vector, while the character embedding layer uses the one-hot representation method for character-level representation of the domain name text; The Small BERT has 2 Transformer layers, 256 hidden units, and 4 attention heads, and its output is a 256-dimensional pooled vector; In the CNN sub-network, multiple one-dimensional convolutional layers, convolutional kernels of different sizes, and strides are used to extract n-gram features of different lengths in the domain name text sequence. The CNN layer uses 4 one-dimensional convolutional layers, with convolutional kernel sizes of 2, 3, 2, 3 respectively, strides of 1, 1, 2, 2 respectively, and 64 convolutional kernels in each layer.

5. The DGA domain name detection method for the IoT botnet according to claim 1, wherein The step S4 includes: The Small BERT network uses the AdamW optimizer with an initial learning rate of 2e-5, while the CNN network uses the Adam optimizer with an initial learning rate of 1e-3; The batch size for training is 128, and the number of epochs is 8. The training adopts an early termination scheme when the metrics do not improve; The binary classification results include DGA domain names or legitimate domain names, and the multi-classification results include random-character DGA domain names, random-syllable DGA domain names, random-word DGA domain names, or legitimate domain names.