Key feature extraction-based DGA domain name active detection method and system

Through the expansion convolution in the two-branch model and the optimization of SE attention module, the problem of inaccurate feature extraction in DGA domain name detection is solved, and more efficient feature extraction and detection accuracy is achieved.

CN120415908AActive Publication Date: 2025-08-01CHINA CRIMINAL POLICE UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510905513.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

In the prior art, when extracting key features of DGA domain names based on deep learning, the data extraction ability of different subdomains character length features during convolution is insufficient, resulting in inaccurate feature extraction and affecting the detection effect.

Method used

Using a dual-branch active detection model, dynamically adjusting the expansion rate and graph theory optimization in the SE attention module through expansion convolution, quantifying information loss during channel compression, and extracting key features.

Benefits of technology

It improves the ability to extract character length features of different subdomains, reduces channel information loss, and improves the accuracy of DGA domain name detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120415908A_ABST
    Figure CN120415908A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security protocols, in particular to a DGA domain name active detection method and system for key feature extraction, and the method comprises the steps: screening out a plurality of sub-domains from domain name data, and converting the sub-domains into a domain name digital sequence; a double-branch active detection model is constructed, and the expansion rate of expansion convolution is dynamically adjusted in a branch 1; in the branch 2, an SE attention module is optimized, information loss caused by channel dimension reduction in the SE attention module is reduced, and key features of sample data are extracted; fusing output results of the double branches to obtain a feature map, and sending the feature map to a maximum pooling layer to extract key features so as to train an active detection model; and obtaining a detection result of the domain name to be detected by using the trained active detection model. According to the method and the device, through two branch extraction results in the active detection model, the extraction capability of the active detection model on different sub-domain character length features is improved, and the possibility that channel information is lost due to channel compression reduction is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of network security protocols, and particularly relates to a method and system for actively detecting DGA domains by extracting key features. Background Art

[0002] A DGA domain name refers to a domain name generated through a Domain Generation Algorithm, based on random characters, time, dictionaries, etc., with features as similar as possible to legitimate domain names to evade detection and blocking. DGA domain names are often used for the connection between botnets with a central structure and C2 servers to establish communication between attackers and infected hosts.

[0003] Since the survival time of DGA domain names is generally short, it poses higher requirements for the detection by the defense side. The defense side needs to detect DGA domain names within the shortest possible time and take corresponding disposal measures to effectively reduce risks. The detection methods of DGA domain names include various detection methods such as detection based on supervised learning, detection based on registration status, detection based on threat intelligence, and detection based on deep learning.

[0004] Among the above detection methods, actively detecting by extracting key features in DGA domain names based on deep learning is one of the mainstream means of current DGA domain name detection. When extracting deep features of DGA domain name data through the SE (Squeeze and Excitation) attention module, the data extraction ability of different sub-domain character length features during convolution is not considered. Additionally, when using the SE attention module to compress and recover channels through two fully connected layers, the reduction in channel compression will lead to the loss of channel information, resulting in the features extracted after recovery not being able to accurately represent the data feature information of the sample data and affecting the accuracy of feature extraction. Summary of the Invention

[0005] To solve the above technical problems, this application provides a method and system for actively detecting DGA domains by extracting key features, and the specific technical solutions adopted are as follows: In the first aspect, an embodiment of this application provides a method for actively detecting DGA domains by extracting key features, and the method includes the following steps: Screen out several sub-domains from the domain name data and convert them into a domain name digital sequence; Construct a dual-branch active detection model; In Branch 1, for the feature that the lengths of the characters forming the sub-domains in different domain name digital sequences are different, convert it into a domain name vector and dynamically adjust the dilation rate of the dilated convolution at each position in the domain name vector to obtain a corresponding feature map through the dilated convolution; In branch 2, the feature map after dilated convolution is subjected to dimensional transformation to obtain a multi-channel feature map for input to the SE attention module; when reducing the number of channels in the first fully-connected layer of the SE attention module, an undirected graph is constructed for the feature map of each channel, and the dense vector of each node in the undirected graph is used to divide the discriminant nodes of each node, and the pooling influence weight of each node is determined by combining the number of paths with an average path length between the nodes in the undirected graph; according to the weight of each feature map output by the second fully-connected layer in the SE attention module, the pooling influence weight of the corresponding node of each element in the corresponding feature map is weighted to obtain the final output of the SE attention module; The feature maps of the same channel extracted by dilated convolution of the feature maps of each channel in the final output of the SE attention module are fused by addition to obtain a fused feature map, which is fed into a max-pooling layer to extract key features for training an active detection model; the detection result of the domain name to be detected is obtained by using the trained active detection model.

[0006] Preferably, the step of screening out several subdomains from the domain name data and converting them into a domain name digital sequence includes: Each complete domain name is split into multiple subdomains according to the hierarchy; Using a dynamic threshold screening mechanism based on the Shannon entropy of position encoding, the subdomain part most likely to be a DGA domain name in each complete domain name is screened out; A mapping from characters to numbers is constructed, and the screened subdomains are converted into corresponding array lists to obtain a computer-readable domain name digital sequence; In the way of encoding and filling with 0, the missing elements in each digital sequence are filled with 0 until the length reaches the preset sequence rated length.

[0007] Preferably, the step of dynamically adjusting the dilation rate of dilated convolution at each position in the domain name word vector includes: The domain name digital sequence is one-hot encoded to generate a binary vector, and the binary vector is converted into a distributed vector by using a word embedding matrix as the domain name word vector; The embedding vector matrix composed of multiple domain name word vectors is input into a recurrent layer to generate a hidden state matrix; Taking the hidden state matrix as the input, the attention weight of each position in the domain name word vector is obtained by using an additive attention module; The dilation rate at each position is dynamically adjusted according to the attention weight at each position.

[0008] Preferably, the dynamic adjustment method of the dilation rate is: In the formula, is the dilation rate at the p-th position, , They are respectively the minimum and maximum values preset for the expansion rate. is the attention weight at the p-th position.

[0009] Preferably, the method for constructing an undirected graph for each channel includes: Regarding each element in the feature map of each channel as a node in the undirected graph, determining whether two nodes in the undirected graph are connected according to whether two elements in the feature map are adjacent elements, and taking the ratio of the frequency of two elements appearing as adjacent nodes to the sum of the frequencies of the two elements as the weight value of the connected edge.

[0010] Preferably, the step of using the dense vector of each node in the undirected graph to divide the discriminant nodes of each node includes: Dividing the Euclidean distance between the dense vector of any node and the dense vectors of the remaining nodes into two categories by means of binary classification, and taking the node corresponding to the dense vector within the category with the largest average Euclidean distance as the discriminant node of the any node.

[0011] Preferably, the calculation method of the pooling influence weight of each node is: Construct the adjacency matrix of each undirected graph, and statistically calculate the average value of the path lengths in all paths from each node to the reachable remaining nodes, denoted as K; In the matrix after calculating the K-th power of the adjacency matrix, each element is the number of paths of length K between the nodes corresponding to its row and column; Calculate the pooling influence weight of node i, denoted as : ; where m is the number of discriminant nodes of node i, is the number of discriminant nodes of node p, 、 are respectively the number of paths of length K from node i to discriminant node a and from node p to discriminant node b, is the length and width of the feature map of each channel.

[0012] Preferably, the weighting method is: ; where 、 are respectively the weighted value and the original value of the b-th element on the feature map of the first channel, is the normalized result of the pooling influence weight of the node corresponding to the b-th element, is the weight value of the feature map of the first channel output by the second fully connected layer in the SE attention module.

[0013] Preferably, before obtaining the detection result of the domain name to be detected by using the trained active detection model, it is also necessary to use the key features output by the active detection model to train a classifier to detect the probabilities of the output key features being legitimate domain names and DGA domain names.

[0014] In a second aspect, another embodiment of the present application further provides a DGA domain name active detection system for key feature extraction, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the DGA domain name active detection method for key feature extraction described in any one of the above.

[0015] The present application has at least the following beneficial effects: The present application extracts key features from domain name samples in a parallel manner. In the dilated convolution module of branch 1, a higher fine-grained dilatation rate adjustment method is generated through significant attention weights to achieve the adjustment of the dilatation rate of a single element, so that the extracted convolutional features can better adapt to the data features with different sub-domain lengths. In the SE attention module of branch 2, the influence of the information lost during the channel compression process on the data information at each element is quantified through graph theory, and the channel weights generated by SE are adjusted, solving the problem of inaccurate weighting results of the channel feature map caused by information loss during the compression of the channel feature map in the SE attention module. Finally, the key features are obtained through the additive fusion of the extraction results of the two branches in the active detection model, improving the extraction ability of the active detection model for the feature of different sub-domain character lengths and reducing the possibility of information loss caused by channel compression reduction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of a DGA domain name active detection method for key feature extraction provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the structure composition of an active detection model for DGA domain names provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the SE attention module provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Embodiment 1 An active detection method for DGA domains with key feature extraction provided by an embodiment of the present application is specifically referred to Figure 1 , and the method includes the following steps: Step 1: Screen out several subdomains from the domain name data and convert them into a domain name digital sequence.

[0019] A domain name is a string bound to the server IP address, consisting of letters, numbers, and special symbols. In this application, the domain name data comes from public data sets, including but not limited to the Tranco data set and the Netlab360 data set. Among them, the Tranco data set is a benign domain name data set, and the Netlab360 data set is DGA domain names collected in real network attacks.

[0020] Here, each complete domain name is split into multiple subdomains according to the hierarchy. For example, a.b.ele.com is split into a, b, ele, com. Then, using a dynamic threshold screening mechanism based on the Shannon entropy of positional encoding, the subdomain part most likely to become a DGA domain name in each complete domain name is screened out. Secondly, a mapping from characters to numbers is constructed, and the mapping relationship is shown in Table 1. The characters include 26 lowercase English letters, 10 numbers, and 2 special symbols ".", "-". The screened subdomains are converted into corresponding array lists to obtain a computer-readable domain name digital sequence.

[0021] Since the lengths of the screened subdomains are unstable, the lengths of the corresponding domain name digital sequences are not exactly the same. Here, the method of encoding and filling with 0 is used to fill the missing elements in each digital sequence with 0 until the length reaches the rated length of the sequence. In this embodiment, the preset rated length of the sequence takes the value of 35.

[0022] As an example, domain names with lengths of 10 and 9 are baicai.com and sdaq.info respectively, and the corresponding domain name digital sequences with length 15 are {2,1,9,3,1,9,28,3,15,13,0,0,0,0,0} and {19,4,1,17,28,9,14,6,15,0,0,0,0,0,0} respectively.

[0023] Table 1 Mapping relationship table Step 2: Construct a dual-branch active detection model. In branch 1, the dilation rate of the dilated convolution is dynamically adjusted; in branch 2, the SE attention module is optimized to reduce the information loss caused by the reduction of the channel dimension in the SE attention module and extract the key features of the sample data.

[0024] In this application, the active detection model structure of DGA domain names is a dual-branch structure, and the structural composition is as Figure 2As shown in the figure, the dilated convolution in Branch 1 consists of a dilated convolution layer 24 and a batch normalization layer 22; Branch 2 contains 3 processing units. Unit 1 and Unit 2 have the same structure, both consisting of a convolution layer 21, a batch normalization layer 22, and an activation function 23. The activation function uses the ReLU function. Unit 3 consists of a convolution layer 21, a batch normalization layer 22, and an SE attention module 27. 24 represents the addition fusion operation, 25 represents the probability distribution of the output of the max pooling layer, and 26 represents the generated distributed vector.

[0025] Among them, in Branch 1, dilated convolution is used to perform convolution operations with different receptive fields on the domain name digital sequence to adapt to the feature differences at different positions caused by different-length sub-domains padded with zeros in the domain name digital sequence of the same length; in Branch 2, graph theory is used to quantify the impact of the information lost during channel compression on the data information at each element, and to adjust the channel weights generated by SE.

[0026] 201. In Branch 1, for the characteristics of different lengths of sub-domain characters in different domain name digital sequences, they are converted into domain name vectors, and the dilation rate of the dilated convolution at each position in the domain name vector is dynamically adjusted to obtain the corresponding feature map through the dilated convolution.

[0027] First, the domain name digital sequence is one-hot encoded to generate a binary vector, and the binary vector is converted into a distributed vector 26 using a word embedding matrix. The dimension of the distributed vector is That is, it serves as the corresponding domain name vector. Among them, is the number of characters in the character set that makes up the domain name, is the preset dimension of the word vector, including but not limited to 128 and 256. In this embodiment, takes the value of 128. The reason for this is that if the binary vector after directly binary encoding the characters has only one valid bit, resulting in strong data sparsity and difficulty in effectively expressing data features, it is necessary to convert it into a word vector to obtain the data information in the sub-domain. The acquisition of the word embedding matrix is a common technique in the field of natural language processing, and the specific process will not be elaborated here.

[0028] Secondly, when using convolution kernels of different scales to perform multi-scale convolution on the domain name vector, on the premise that the sliding step size remains unchanged, for each element on the domain name vector, the smaller the size of the convolution kernel, the less context space information involved in the convolution calculation, but the more accurate the local feature expression of each element in the convolution result.

[0029] Furthermore, since the convolution matrices obtained by convolution kernels of different sizes are expressions of domain name vectors at different scales, and normal domain names usually have strong semanticity and are usually created to represent specific companies, actual products or specific organizations, including words, abbreviations or phrases with specific actual meanings; while DGA domain names usually lack semanticity and are automatically generated by algorithms. For the purpose of being difficult to predict and track, characters, letters, etc. are usually randomly combined and have strong randomness. This results in the positions that can reflect key feature content in the domain name digital sequences mapped by sub-domains of different lengths being not fixed. When using dilated convolution for convolution operations, it is necessary to dynamically adjust the dilation rate. The specific process is as follows: First of all, normal domain names usually have strong semanticity and are usually created to represent specific companies, actual products or specific organizations, including words, abbreviations or phrases with specific actual meanings; while DGA domain names usually lack semanticity and are automatically generated by algorithms. For the purpose of being difficult to predict and track, characters, letters, etc. are usually randomly combined and have strong randomness. Therefore, in the domain name vectors corresponding to the domain name data, there are also certain differences in the receptive field sizes required to extract data features at different positions. That is, the significant regions in the input data usually correspond to the key structures or high-frequency details of the domain name data, while the low-significance regions are more likely to correspond to the conventional features in the domain name data, such as the data features at the zero-padding positions. When performing dilated convolution, the low-significance regions require a larger receptive field to capture the context.

[0030] Secondly, the embedding vector matrix composed of domain name vectors (a matrix formed by stacking multiple domain name vectors in sequence) is input into the recurrent layer to generate a hidden state matrix. The recurrent layer includes but is not limited to Bi-LSTM, Bi-GRU, etc. Taking the hidden state matrix as the input, an additive attention module is used to obtain the attention weights of each position in the domain name vector. The larger the attention weight, the stronger the significance of the position in the entire input data and the more important the data feature; the smaller the attention weight, the weaker the significance of the position in the entire input data and the relatively lower importance of the data feature.

[0031] Here, the dynamic adjustment method of the dilation rate at the p-th position is: In the formula, is the dilation rate at the p-th position, and are respectively the preset minimum and maximum values of the dilation rate, and take values 1 and 8 respectively in this embodiment, is the attention weight at the p-th position.

[0032] Among them, the greater the attention weight at the p-th position indicates that the data features at the p-th position are more prominent. To extract the key structure or high-frequency details of the data at the p-th position, a smaller receptive field should be used for convolution operations.

[0033] Furthermore, for each position in the domain name vector, when the convolutional kernel performs convolution operations with the elements at each position, the receptive field during convolution is adjusted according to the dilation rate of each position to obtain the feature map corresponding to the domain name vector, with the dimension of , and the calculation of the dimension depends on the dimension of the domain name vector, the scale of the convolutional kernel, and the sliding step size. The calculation of the convolutional dimension is a well-known technology in the field of deep learning, and the specific process will not be elaborated here.

[0034] So far, the above method of generating a higher fine-grained dilation rate adjustment through the significant attention weight realizes the dilation rate adjustment of a single element, making the extracted convolutional features better adapt to the data features with different sub-domain lengths.

[0035] 202. Perform dimension conversion on the feature map after dilated convolution to obtain a multi-channel feature map for input to the SE attention module.

[0036] The SE attention module usually needs to perform dimension conversion on the convolutional feature map to obtain a multi-channel feature map U with a specific dimension required, with the size of , 、 are the length and width of each channel feature map, and is the number of channels after conversion. Among them, the dimension conversion of the feature map is completed by performing convolution operations between the convolutional feature map and a set of filter kernels with a preset size. The dimension conversion of the feature map is an existing technology, and the specific process will not be elaborated here.

[0037] 203. When reducing the number of channels in the first fully connected layer of the SE attention module, construct an undirected graph for each channel feature map, use the dense vectors of each node in the undirected graph to divide the discriminant nodes of each node, and determine the pooling influence weight of each node in combination with the number of paths with an average path length between the nodes in the undirected graph.

[0038] In the first fully connected layer of the SE attention module, use r (scaling channel scale parameter) to reduce the input feature channels, directly compress the feature map into a feature vector, and each element in the feature vector is a value obtained by compressing the channel features of the feature map. The specific compression method is: use average pooling to take the mean value of all elements in each feature map as the compressed value.

[0039] In the process of reducing feature compression in this channel, it will lead to the loss of feature information in the domain name data expressed by elements at different positions on the feature map. This application considers quantifying the impact of the lost information entering the first fully connected layer on the local information of each element in the feature map based on graph theory, and jointly applying the quantization result and the feature vector to the feature map of each channel to obtain the final output of the SE attention module. The specific process is as follows: First, construct an undirected graph for each channel in the multi-channel feature map of a specific dimension in the channel dimension. When performing channel compression through global average pooling, the sizes of elements at different positions, which reflect different feature information of the domain name data, are ignored. Here, each element on the feature map within each channel is regarded as a node, and the relevance between different nodes is inferred using undirected graph technology. This is because global average pooling compresses the elements of the entire feature map into a real number, and the loss of channel information has a greater impact on the local information of the elements.

[0040] Secondly, each element in the feature map of each channel is regarded as a node in the undirected graph. Whether two nodes in the undirected graph are connected is determined by whether the two elements in the feature map are adjacent elements. If two elements are adjacent elements, the corresponding nodes of the two elements are connected, and the ratio of the frequency of the two elements appearing as adjacent nodes to the sum of the frequencies of the two elements is used as the weight of the connected edge, obtaining the undirected graph corresponding to the feature map of each channel.

[0041] Subsequently, use the constructed undirected graph as the input, and use the Node2Vec algorithm to output the dense vector of each node. In this embodiment, the number of steps for each random walk in the algorithm is set to 30, the number of random walks for each node is set to 200, the embedding dimension is set to 64, the return parameters p and in-out q are set to 0.5 and 2 respectively. The Node2Vec method is a common technique in the field of undirected graphs, and the specific process will not be elaborated here.

[0042] The dense vector of each node is an expression of the local information at the node. Since the global average pooling of the first fully connected layer compresses the elements of the entire feature map into a real number, the loss of channel information is more an expression of local information. Use the dense vector of the node to evaluate the impact degree of channel compression on each node, and then reconstruct the lost information through interpolation.

[0043] Specifically, the Euclidean distances between the dense vector of node i and the dense vectors of the remaining nodes are divided into two categories by means of binary classification. The node corresponding to the dense vector within the category with the largest average Euclidean distance is used as the discriminant node of node i. If the information propagation between nodes with large differences in local structural features is easier, then the data features of the element corresponding to node i are more similar to the data features expressed by the entire feature map. Among them, the binary classification can be to divide all the Euclidean distances into two categories by using a clustering algorithm, and the clustering algorithm includes but is not limited to k-means clustering, DBSCAN clustering, etc. In other embodiments, statistical classification can also be used. For example, the average value of all the Euclidean distances is used as the binary classification threshold to achieve binary classification.

[0044] After that, the adjacency matrix of each undirected graph is constructed, and the average value of the path lengths in all paths from each node to the remaining reachable nodes is calculated and denoted as K. Calculate the K-th power of the adjacency matrix. Taking the element in the i-th row and j-th column of the K-th power of the adjacency matrix as an example, this element represents the number of paths from node i to node j with a length of K.

[0045] Among them, the fewer the number of paths between node i and node j, the more difficult the data information propagation between node i and node j; on the contrary, the more the number of paths, the easier the information propagation between node i and node j. At the same time, the more nodes node i can communicate with through shorter paths, the greater the influence of the data information of node i on the data information of the remaining nodes, and the greater the influence of global average pooling on the domain name information expressed by the element corresponding to node i.

[0046] Here, calculate the pooling influence weight of node i, denoted as : In the formula, m is the number of discriminant nodes of node i, is the number of discriminant nodes of node p, , are respectively the number of paths from node i to discriminant node a with a length of K and the number of paths from node p to discriminant node b with a length of K, is the length and width of the feature map of each channel.

[0047] 204. According to the weights of each feature map output by the second fully connected layer in the SE attention module, the pooling influence weights of the nodes corresponding to each element in the corresponding feature map are weighted to obtain the final output of the SE attention module.

[0048] First, use the multi-channel feature map U converted in step 202 as the input of the SE attention module. First, perform global average pooling to compress into The eigenvector; then, using two fully connected layers, the first fully connected layer compresses the channels into channels to reduce the computational amount, and then through a RELU non-linear activation layer, the second fully connected layer restores the number of channels back to channels, and then through the Sigmoid activation to obtain the weight vector , where , , are the weights of the first, second, and th channel feature maps output by the second fully connected layer in the SE attention module respectively, r is the compression ratio, which takes the value of 16 in this embodiment. The SE attention module is a commonly used technology in the field of deep learning, and the specific process will not be elaborated.

[0049] Secondly, for the feature map of each channel, taking the feature map F1 of the first channel as an example, according to step 203, calculate the pooling influence weight of each element corresponding node in the feature map F1, and combine the weight of the feature map F1 to weight each element's pooling influence weight to obtain the weighted result of the feature map F1. The weighted results of

[0050] channels feature maps are arranged in channel order to obtain the final output X of the SE attention module. Figure 3 In Figure 3 shown in, 31 represents Figure 2 the output of unit 2 in branch 2 in , 32 represents the multi-channel feature map U, 33 represents the feature vector compressed by the first fully connected layer, 34 represents the obtained weight vector S, 35 represents the undirected graph constructed by the single-channel feature map, 36 represents the adjacency matrix of the undirected graph, 37 represents the final output X of the SE attention module,

[0051] Here, calculate the weighted result of the b-th element on the feature map F1: In the formula, , are the weighted value and the original value of the b-th element on the feature map of the first channel respectively, is the normalized result of the pooling influence weight of the b-th element corresponding node , is the weight of the first channel feature map output by the second fully connected layer in the SE attention module. Among them, the normalization is based on the pooling influence weights of all elements corresponding nodes on the feature map F1 for normalization.

[0052] Among them, the greater the pooling influence weight of the b-th element corresponding to the node on the feature map F1, the more local information of the b-th element is lost during the global average pooling process. To reduce the impact of information loss in channel compression, the weighted value should be amplified during weighted reconstruction; the less local information of the b-th element is lost during the global average pooling process, a relatively small weighted value is used during weighted reconstruction to ensure the accuracy of the weighted result corresponding to the feature map F1.

[0053] Step 3: Additively fuse the feature maps of the same channels extracted by dilated convolution for each channel in the final output of the SE attention module to obtain a fused feature map, and send it to the max pooling layer to extract key features for training the active detection model; use the trained active detection model to obtain the detection result of the domain name to be detected.

[0054] First, additively fuse the feature maps of each channel in the final output X of the SE attention module with the feature maps of the same channels extracted by dilated convolution to obtain a fused feature map, and send the fused feature map to the max pooling layer to output a probability distribution. The max pooling operation compresses the dimension of the vector after additively computing feature fusion, and filters and outputs the most discriminative key features by extracting the maximum response of the features.

[0055] Furthermore, use the legitimate domain names in the legitimate domain name set and DGA domain names to train the active detection model. The ratio range of the training set to the test set is 7:3 - 9:1. Preferably, in this embodiment, the ratio of the training set to the test set is 8:2. Extract the key features of the samples in the training set through the above process, train a classifier based on the key features, the optimizer is the Adam (Adaptive Moment Estimation) optimizer, and the loss function is the cross-entropy loss function. The training of the neural network is a well-known technology, and the specific process will not be elaborated here.

[0056] After that, for the domain name to be detected, perform digital mapping encoding on it and use it as the input to the trained active detection model to output the detection result of the domain name to be detected. The detection result includes the probability of the legitimate domain name and the probability of the DGA domain name.

[0057] Embodiment 2 Another embodiment of the present application also provides a DGA domain name active detection system for key feature extraction, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the DGA domain name active detection method for key feature extraction described in any one of the above.

[0058] Other embodiments of the present application will be readily contemplated by those skilled in the art in view of the specification and practice of the invention herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not invented by the present application.

[0059] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. An active detection method for DGA domain names with key feature extraction, characterized in that, The method includes the following steps: Filter out several subdomains from the domain name data and convert them into a domain name digital sequence; Construct a two-branch active detection model; In Branch 1, for the feature that the lengths of the characters forming the subdomains in different domain name digital sequences are different, convert them into domain name word vectors and dynamically adjust the dilation rate of the dilated convolution at each position in the domain name word vector, so as to obtain the corresponding feature map through the dilated convolution; In Branch 2, perform dimensional conversion on the feature map after dilated convolution to obtain a multi-channel feature map for input to the SE attention module; when reducing the number of channels in the first fully connected layer of the SE attention module, construct an undirected graph for the feature map of each channel, use the dense vector of each node in the undirected graph to divide the discriminant nodes of each node, and determine the pooling influence weight of each node in combination with the number of paths with an average path length existing between the nodes in the undirected graph; according to the weight of each feature map output by the second fully connected layer in the SE attention module, weight the pooling influence weight of the corresponding node of each element in the corresponding feature map to obtain the final output of the SE attention module; Additively fuse the feature maps of the same channel extracted by the dilated convolution of the feature maps of each channel in the final output of the SE attention module to obtain a fused feature map, and send it to the max pooling layer to extract key features for training the active detection model; use the trained active detection model to obtain the detection result of the domain name to be detected.

2. The active detection method for DGA domain names with key feature extraction as described in claim 1, characterized in that, The step of filtering out several subdomains from the domain name data and converting them into a domain name digital sequence includes: Split each complete domain name into multiple subdomains according to the hierarchy; Use a dynamic threshold screening mechanism based on position-encoded Shannon entropy to screen out the subdomain part in each complete domain name that is most likely to be a DGA domain name; Construct a mapping from characters to numbers, convert the screened subdomains into corresponding array lists to obtain a computer-readable domain name digital sequence; Use the method of encoding and padding with 0 to pad the missing elements in each digital sequence with 0 until the length reaches the preset sequence rated length.

3. The active detection method for DGA domain names with key feature extraction according to claim 2, characterized in that, The step of dynamically adjusting the dilation rate of the dilated convolution at each position in the domain name word vector includes: Perform one-hot encoding on the domain name digital sequence to generate a binary vector, and use the word embedding matrix to convert the binary vector into a distributed vector as the domain name word vector; Input the embedding vector matrix composed of multiple domain name word vectors into the recurrent layer to generate a hidden state matrix; Use the hidden state matrix as the input, and use the additive attention module to obtain the attention weight at each position in the domain name word vector; Dynamically adjust the dilation rate at this position according to the attention weight at each position.

4. The active detection method for DGA domain names with key feature extraction according to claim 3, characterized in that, The dynamic adjustment method of the dilation rate is: In the formula, is the expansion rate at the p-th position, , are respectively the preset minimum and maximum values of the expansion rate, is the attention weight at the p-th position.

5. The active detection method for DGA domain names with key feature extraction according to claim 1, characterized in that, The method of constructing an undirected graph for each channel includes: Take each element in the feature map of each channel as a node in the undirected graph, determine whether two nodes in the undirected graph are connected according to whether the two elements in the feature map are adjacent elements, and use the ratio of the frequency of the two elements appearing as adjacent nodes to the sum of the frequencies of the two elements as the weight of the connected edge.

6. The active detection method for DGA domain names with key feature extraction according to claim 1, characterized in that The step of using the dense vector of each node in the undirected graph to divide the discriminant nodes of each node includes: Using a binary classification method, the Euclidean distance between any node and the dense vectors of the remaining nodes is divided into two categories, and the node corresponding to the dense vectors within the category with the largest average Euclidean distance is used as the discriminant node of the said any node.

7. The active detection method for DGA domain names with key feature extraction according to claim 6, wherein, The calculation method of the pooling influence weight of each node is as follows: Construct the adjacency matrix of each undirected graph, and calculate the average value of the path lengths in all paths from each node to the remaining reachable nodes, denoted as K; In the matrix after calculating the K-th power of the adjacency matrix, each element is the number of paths of length K between the nodes corresponding to its row and column; The pooling influence weight of computing node i, denoted as : ; where m is the number of discriminant nodes of node i, is the number of discriminant nodes of node p, , are respectively the number of paths from node i to discriminant node a with length K and from node p to discriminant node b with length K, is the length and width of the feature map of each channel.

8. The active detection method for DGA domain names with key feature extraction according to claim 7, characterized in that, The weighting method is as follows: ; In the formula, , are the weighting value and the original value of the b-th element on the feature map of the first channel respectively, is the pooling influence weight of the node corresponding to the b-th element is the normalization result of is the weight value of the feature map of the first channel output by the second fully connected layer in the SE attention module.

9. The active detection method for DGA domain names with key feature extraction as described in claim 1, characterized in that, Before using the trained active detection model to obtain the detection result of the domain name to be detected, it is also necessary to use the key features output by the active detection model to train a classifier to detect the probability that the output key features are legitimate domain names and the probability of DGA domain names.

10. An active detection system for DGA domain names with key feature extraction, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the DGA domain name active detection method for key feature extraction as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Training method, system, application method and system of DGA domain name detection model

    CN115758263A

  • DGA domain name detection method under condition of unbalanced positive and negative sample proportion

    CN116318845A

  • Malicious DGA domain name detection method and system

    CN117834292A

  • Hypergraph-based domain name test method and apparatus

    WO2024183339A1