DGA malicious domain name detection method based on federated learning and deep learning

By combining federated learning with deep learning framework, the high false positive rate and data privacy issues in DGA malicious domain name detection are solved, efficient and accurate malicious domain name detection is achieved, data privacy is protected, and the generalization ability and robustness of the model are improved.

CN120263512BActive Publication Date: 2025-09-19JIANGXI LINLU TECHNOLOGY SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510536069.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-09-19
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Existing technologies have high false positive rates, high operating costs, data privacy and security issues in DGA malicious domain name detection, and insufficient model generalization and robustness, making it difficult to cope with diverse and dynamic network attacks.

Method used

A method based on federated learning and deep learning is adopted. The model is trained on a local dataset through a federated learning framework. Combined with a character-level dynamic encoding algorithm, a multi-scale text convolutional neural network, and a bidirectional long short-term memory attention enhancement network, the domain name detection model is optimized and feature extracted, avoiding direct sharing of raw data, protecting data privacy, and improving the model's generalization ability and recognition accuracy.

Benefits of technology

The accuracy and generalization ability of DGA malicious domain name detection are improved, the robustness of the model and data privacy protection are enhanced, it can effectively respond to diverse and dynamic network attacks, realize the protection of data privacy, and improve the generalization ability and recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263512B_ABST
    Figure CN120263512B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of malicious domain name detection and proposes a method for detecting DGA malicious domain names based on federated learning and deep learning. By designing a domain name detection model based on a federated learning framework, participants are allowed to independently train the model on a local data set, thereby protecting data privacy. By regularly sending model parameters or gradient information to a central server for aggregation and updating, global model optimization is achieved, thereby improving the performance of detecting DGA malicious domain names. Furthermore, by utilizing a portion of the DGA malicious domain name data set held by each participant and through the collaborative training mechanism of the federated learning framework, each participant can share the optimization results of the model, thereby improving the generalization ability of the model. By combining a character-level dynamic encoding algorithm, a multi-scale text convolutional neural network, and a bidirectional long short-term memory attention enhancement network, the accuracy and stability of identifying DGA malicious domain names are greatly improved. The present invention improves the accuracy and generalization ability of DGA malicious domain name detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of malicious domain name detection, and in particular to a DGA malicious domain name detection method based on federated learning and deep learning. Background Art

[0002] With the rapid development of Internet technology, the types of network attacks are becoming increasingly diverse, and security incidents such as denial of service attacks and phishing attacks occur frequently. Network attacks usually use domain generation algorithms (DGAs) to generate random domain names to hide the real IP addresses of malware command and control (C&C) servers, thereby building botnets and carrying out attacks.

[0003] In the existing technology, traditional detection methods based on rule bases or static blacklists are often used. However, the randomness and dynamic characteristics of DGA domain name generation have a great impact on the accuracy of detection. First, the existing detection methods have high false alarm rates and operating costs. Traditional detection methods are usually based on manually defined rules (such as domain name length, character entropy, and N-gram frequency) for detection. These rules are prone to misjudgment when facing legal but "abnormal" domain names. Moreover, the blacklist relies on historical data and cannot distinguish between "maliciously registered but unused domain names" and real malicious domain names, which leads to the easy blocking of legitimate domain names. Manual review is required, increasing the operation and maintenance burden. Second, the existing detection methods have data privacy and security issues. Due to industry regulatory requirements and data privacy protection restrictions, it is difficult to achieve effective circulation of sample data between different security agencies, forming a "data island" phenomenon. The threat intelligence stored by each agency is often limited to specific family variants within its own monitoring range. This data fragmentation pattern directly restricts the generalization ability of the detection model. When a new DGA family or upgraded variant appears, the model trained by a single agency is easy to There are blind spots in identification. Third, existing detection methods have insufficient generalization and robustness. Since DGA malicious domain names are highly diverse and dynamic, detection models need to have good generalization and robustness to cope with various unknown attack scenarios. However, in practical applications, deep learning models face the problems of overfitting or underfitting, resulting in poor performance on unknown data. The detection rate of the generated grammatically compliant malicious domain name model is less than 70%, and the misjudgment rate of the domain name IDN homograph attack model rises to 32.1%. In addition, when the model is sensitive to noise or outliers, the robustness of the model will also be affected. Fourth, existing detection methods have difficulties in training and evaluation. Since data on DGA malicious domain names is often difficult to obtain and the labeling cost is high, training deep learning models will face the problem of insufficient data. In addition, evaluating model performance also requires a large amount of test data and appropriate evaluation indicators. Existing studies use different indicators and test sets, resulting in poor comparability of model performance. Therefore, in practical applications, it is difficult to obtain sufficient test data and design reasonable evaluation indicators.

[0004] Therefore, how to design a DGA malicious domain name detection method to improve the accuracy, generalization ability and robustness of detection has become an urgent problem to be solved. Summary of the Invention

[0005] Based on this, the present invention proposes a DGA malicious domain name detection method based on federated learning and deep learning. By designing a domain name detection model based on a federated learning framework, participants are allowed to independently train models on local data sets without directly sharing original data, thereby protecting data privacy. By regularly sending model parameters or gradient information to a central server for aggregation and updating, global model optimization is achieved, improving the performance of detecting DGA malicious domain names, and fully utilizing the partial DGA malicious domain name data sets held by each participant. Through the collaborative training mechanism of the federated learning framework, each participant can share the optimization results of the model, accelerate the training speed of the model, and improve the generalization ability of the model. The combination of a character-level dynamic encoding algorithm, a multi-scale text convolutional neural network, and a bidirectional long short-term memory attention enhancement network greatly improves the recognition accuracy and stability of DGA malicious domain names. The present invention improves the accuracy and generalization ability of DGA malicious domain name detection.

[0006] The present invention proposes a method for detecting DGA malicious domain names based on federated learning and deep learning, including:

[0007] Obtain the target original domain name and input it into a domain name detection model. The domain name detection model is based on a federated learning framework and includes a data preprocessing module, a feature extraction module, and a feature fusion module.

[0008] Normalizing the target original domain name according to a character-level dynamic encoding algorithm to obtain a semantic representation of the target domain name. The character-level dynamic encoding algorithm extracts effective features of the domain name sequence in the target original domain name based on character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by a domain name detection model. The semantic representation of the target domain name is a vector representation of fixed dimension.

[0009] Performing feature extraction processing on the target domain name semantic representation according to a multi-scale text convolutional neural network to obtain multi-scale local spatial features, wherein the multi-scale text convolutional neural network includes an input layer, a convolution layer, a pooling layer, and a fully connected layer;

[0010] Performing feature extraction processing on the multi-scale local spatial features according to a bidirectional long short-term memory attention enhancement network to obtain global temporal features, wherein the bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is used to capture contextual information from the sequence of the target domain name semantic representation forward to backward, and the attention enhancement mechanism is used to focus attention on key characters in the target domain name semantic representation;

[0011] The multi-scale local spatial features and the global temporal features are subjected to feature fusion to obtain domain name fusion features, and classification and recognition are performed according to the domain name fusion features.

[0012] In summary, according to the above-mentioned DGA malicious domain name detection method based on federated learning and deep learning, by designing a domain name detection model based on a federated learning framework, participants are allowed to independently train models on local data sets without directly sharing original data, thereby protecting data privacy. By regularly sending model parameters or gradient information to a central server for aggregation and updating, global model optimization is achieved, improving the performance of detecting DGA malicious domain names, and being able to fully utilize part of the DGA malicious domain name data set held by each participant. Through the collaborative training mechanism of the federated learning framework, each participant can share the optimization results of the model, accelerate the training speed of the model, and improve the generalization ability of the model. The combination of a character-level dynamic encoding algorithm, a multi-scale text convolutional neural network, and a bidirectional long short-term memory attention enhancement network greatly improves the recognition accuracy and stability of DGA malicious domain names. The present invention improves the accuracy and generalization ability of DGA malicious domain name detection. Specifically, the target original domain name is obtained and input into the domain name detection model. The domain name detection model is based on the federated learning framework. The domain name detection model includes a data preprocessing module, a feature extraction module and a feature fusion module. It allows participants to independently train the model on the local data set without directly sharing the original data, thereby protecting data privacy and being able to make full use of part of the DGA malicious domain name data set held by each participant. Through the collaborative training mechanism of the federated learning framework, the generalization ability of the model is improved. The target original domain name is normalized according to the character-level dynamic encoding algorithm to obtain the target domain name semantic representation. The character-level dynamic encoding algorithm extracts the effective features of the domain name sequence in the target original domain name according to the character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by the domain name detection model. The target domain name semantic representation is a vector representation of fixed dimension, and the character-level convolutional network layer is used to dynamically generate the embedding representation of the whole word, so that the domain name detection model can It has more advantages when dealing with special text types, especially in terms of spelling errors, term variants and unregistered words, showing stronger robustness and generalization ability. Relying on character-level dynamic encoding, it can completely retain the original sequence information, avoid information loss caused by excessive segmentation, and further improve adaptability and generalization ability. The target domain name semantic representation is feature extracted according to the multi-scale text convolutional neural network to obtain multi-scale local spatial features. The multi-scale text convolutional neural network includes an input layer, a convolution layer, a pooling layer and a fully connected layer. Malicious domain names usually contain a large number of random characters, variant spellings or meaningless fragments injected by humans. Traditional methods based on feature rules or manual analysis are difficult to effectively handle these variations and new threats. The present invention is based on character-level convolution operations, which can automatically learn pattern features in domain name character sequences, and then identify potential abnormal or malicious feature combinations. It can effectively deal with the problems of spelling variants or domain name noise, especially when faced with a large amount of short and irregular domain name data.It shows higher robustness and generalization ability, and can automatically learn malicious feature patterns in domain names without manually defining features, so that it is more adaptable to rapidly changing malicious domain name detection scenarios, effectively improving the accuracy and real-time performance of malicious domain name recognition. The multi-scale local spatial features are feature extracted and processed according to the bidirectional long short-term memory attention enhancement network to obtain global temporal features. The bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is used to capture contextual information from the sequence of the target domain name semantic representation forward to backward, and the attention enhancement mechanism is used to focus on the key points in the semantic representation of the target domain name. The model focuses on key characters and simultaneously traverses the text sequence in both the forward and backward directions to more comprehensively capture the context. The bidirectional structure can better perceive the semantic association between the preceding and following characters in the domain name string, and further highlight the characteristics of abnormal characters or key character sequences through the attention mechanism. The dynamic focusing capability enables the model to show higher recognition ability for variant domain names, randomly generated domain names, and human interference features, further improving recognition accuracy. The multi-scale local spatial features and the global temporal features are fused to obtain domain name fusion features, and classification and recognition are performed based on the domain name fusion features. This invention improves the accuracy and generalization ability of DGA malicious domain name detection.

[0013] Furthermore, the federated learning framework is specifically as follows:

[0014] The domain name detection model is based on a federated learning framework, wherein the federated aggregation algorithm of the federated learning framework is based on a federated performance averaging algorithm, and the federated performance averaging algorithm includes a dynamic weighting mechanism;

[0015] The specific formula of the federated performance averaging algorithm is as follows:

[0016] ,

[0017] ,

[0018] ,

[0019] ,

[0020] in, represents the global model parameters after aggregation, represents a time sequence number, Indicates the total number of clients, Indicates the client ordinal number, represents the comprehensive weight, Indicates the total amount of client data. Represents a single client data, represents the performance adjustment factor, Represents the client performance score, represents the scoring weight, Indicates the current accuracy of the client. Indicates the current F1 value of the client.

[0021] Furthermore, the federated learning framework also includes:

[0022] Privacy protection is performed based on a differential privacy algorithm. The differential privacy algorithm adds pre-calibrated random noise to sensitive information during training to limit sensitivity to changes in a single training sample. The random noise is added based on Gaussian noise.

[0023] Furthermore, the step of normalizing the target original domain name according to the character-level dynamic encoding algorithm to obtain the semantic representation of the target domain name is as follows:

[0024] Normalizing the target original domain name according to a character-level dynamic encoding algorithm, splitting the target original domain name string into a character sequence character by character, and embedding each character in the character sequence to obtain a character embedding vector for the corresponding character;

[0025] Performing convolution processing on the character embedding vector according to a convolutional network layer to obtain character features;

[0026] Performing feature screening processing on the character features according to the maximum pooling layer to obtain significant character features;

[0027] Smoothing the significant character features according to the high-speed network layer to obtain smoothed character features;

[0028] Converting the smoothed character features into a vector representation of a fixed dimension to obtain a semantic representation of the target domain name;

[0029] The specific formula of the character-level dynamic encoding algorithm is as follows:

[0030] ,

[0031] ,

[0032] ,

[0033] in, Represents character features, represents the convolution kernel number, represents the activation function, represents the weight matrix of the convolution kernel, represents consecutive character embedded subsequences, represents the bias term, Indicates significant character features, represents the input feature of the maximum pooling layer, represents smooth character features, represents the gate control unit, represents the nonlinear transformation of the high-speed network layer, Represents the input features of the highway network layer.

[0034] Furthermore, the step of performing feature extraction processing on the semantic representation of the target domain name according to the multi-scale text convolutional neural network to obtain multi-scale local spatial features is specifically as follows:

[0035] Performing feature extraction processing on the semantic representation of the target domain name according to a multi-scale text convolutional neural network, wherein the multi-scale text convolutional neural network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer;

[0036] Extract local features of different scales based on multi-size convolution kernels;

[0037] Performing local maximum pooling processing on the local features of different scales to reduce the feature dimensions of the local features and obtain significant local features of different scales;

[0038] Fuse the significant local features at different scales to obtain multi-scale local spatial features;

[0039] The specific algorithm of the multi-scale text convolutional neural network is as follows:

[0040] ,

[0041] ,

[0042]

[0043] in, Represents local features at different scales, represents the activation function, represents the convolution operation, Represents the input text sequence, represents the convolution kernel size, Indicates the output channel, represents the size of a single convolution kernel, Represents significant local features at different scales, Indicates the batch size, Indicates the length of the input text sequence, Represents multi-scale local spatial features, Represents the characteristic scale.

[0044] Furthermore, the step of performing feature extraction processing on the multi-scale local spatial features according to the bidirectional long short-term memory attention enhancement network to obtain global temporal features is specifically as follows:

[0045] Context information is extracted from multi-scale local spatial features according to a bidirectional long short-term memory mechanism to obtain temporal features, wherein the context information extraction performs a bidirectional traversal of the text sequence in forward and backward order;

[0046] The specific formula of the bidirectional long short-term memory mechanism is as follows:

[0047] ,

[0048] ,

[0049] in, Represents the time series characteristics, represents a bidirectional long short-term memory mechanism, Represents multi-scale local spatial features, Indicates the batch size, represents the characteristic scale, and Represent the forward and backward time series features respectively, and represent forward and backward traversal respectively, Indicates time sequence;

[0050] Attention enhancement processing is performed on the temporal features according to an attention enhancement mechanism to obtain global temporal features.

[0051] Furthermore, the step of performing attention enhancement processing on the time series features according to the attention enhancement mechanism to obtain global time series features specifically includes:

[0052] The attention enhancement mechanism is used to enhance the temporal features. The temporal features of each time step are calculated based on a single-layer linear transformation and a nonlinear activation function to obtain the importance score.

[0053] Then calculate the attention weight according to the Softmax function;

[0054] Performing attention weighting processing on the time series feature according to the attention weight to obtain a global time series feature;

[0055] The specific formula of the attention enhancement mechanism is as follows:

[0056] ,

[0057] in, Represents the global timing characteristics, represents a time sequence number, represents the attention weight, represents the temporal characteristics of a single time step, Indicates the batch size.

[0058] The present invention proposes a DGA malicious domain name detection system based on federated learning and deep learning, including:

[0059] A preprocessing module is used to obtain the target original domain name and input it into a domain name detection model. The domain name detection model is based on a federated learning framework and includes a data preprocessing module, a feature extraction module, and a feature fusion module.

[0060] Normalizing the target original domain name according to a character-level dynamic encoding algorithm to obtain a semantic representation of the target domain name. The character-level dynamic encoding algorithm extracts effective features of the domain name sequence in the target original domain name based on character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by a domain name detection model. The semantic representation of the target domain name is a vector representation of fixed dimension.

[0061] A local feature extraction module is used to perform feature extraction processing on the semantic representation of the target domain name according to a multi-scale text convolutional neural network to obtain multi-scale local spatial features. The multi-scale text convolutional neural network includes an input layer, a convolution layer, a pooling layer, and a fully connected layer;

[0062] a global feature enhancement module, configured to perform feature extraction processing on the multi-scale local spatial features based on a bidirectional long short-term memory attention enhancement network to obtain global temporal features, wherein the bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is configured to capture contextual information from the sequence of the target domain name semantic representation forward to backward, and the attention enhancement mechanism is configured to focus attention on key characters in the target domain name semantic representation;

[0063] The feature fusion module is used to fuse the multi-scale local spatial features and the global temporal features to obtain domain name fusion features, and perform classification and recognition based on the domain name fusion features.

[0064] The present invention also provides a storage medium, which stores one or more programs. When the programs are executed by a processor, they implement the above-mentioned DGA malicious domain name detection method based on federated learning and deep learning.

[0065] The present invention further provides a computer device, comprising a memory and a processor, wherein:

[0066] The memory is used to store computer programs;

[0067] When the processor is used to execute the computer program stored in the memory, it implements the above-mentioned DGA malicious domain name detection method based on federated learning and deep learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a flowchart of the DGA malicious domain name detection method based on federated learning and deep learning proposed in the first embodiment of the present invention;

[0069] Figure 2 This is a schematic diagram of the structure of a DGA malicious domain name detection system based on federated learning and deep learning proposed in the second embodiment of the present invention;

[0070] Figure 3 A flowchart of the domain name detection model of the present invention;

[0071] Figure 4 Schematic diagram of the structure of the federated learning framework of the present invention.

[0072] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0073] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.

[0074] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0076] See also Figure 1, which is a flow chart of a DGA malicious domain name detection method based on federated learning and deep learning proposed in the first embodiment of the present invention, the DGA malicious domain name detection method based on federated learning and deep learning includes steps S01 to S05, wherein:

[0077] Step S01: Obtain the target original domain name and input it into the domain name detection model;

[0078] It should be noted that in this embodiment, the domain name detection model is based on the federated learning framework. For the specific process structure of the domain name detection model, please refer to Figure 3 , the specific structure of the federated learning framework can be found in Figure 4 The domain name detection model includes a data preprocessing module, a feature extraction module, and a feature fusion module. To address the special needs of malicious domain name detection, this paper proposes FedPerAvg. This algorithm retains the data volume weight of FedAvg and innovatively incorporates a performance-based dynamic weighting mechanism. This design retains the basic influence of data volume in the original FedAvg and adjusts the final weight through a performance adjustment factor, giving better-performing clients greater influence. The federated learning framework is as follows:

[0079] The domain name detection model is based on a federated learning framework, wherein the federated aggregation algorithm of the federated learning framework is based on a federated performance averaging algorithm, and the federated performance averaging algorithm includes a dynamic weighting mechanism;

[0080] The specific formula of the federated performance averaging algorithm is as follows:

[0081] ,

[0082] ,

[0083] ,

[0084] ,

[0085] in, represents the global model parameters after aggregation, represents a time sequence number, Indicates the total number of clients, Indicates the client ordinal number, represents the comprehensive weight, Indicates the total amount of client data. Represents a single client data, represents the performance adjustment factor, Represents the client performance score, represents the scoring weight, Indicates the current accuracy of the client. Indicates the current F1 value of the client;

[0086] The federated learning framework also includes:

[0087] Privacy protection for DGA malicious domain name detection is achieved under the federated learning framework, mainly relying on differential privacy technology. This framework provides formal privacy protection for DGA malicious domain name detection under federated learning, enabling each client to collaboratively train efficient models while protecting sensitive domain name data. Differential privacy adds carefully calibrated random noise to sensitive information during training to ensure that the model output is sensitive to changes in a single training sample, thereby preventing privacy leakage. Privacy protection processing is performed based on the differential privacy algorithm. The differential privacy algorithm adds pre-calibrated random noise to sensitive information during training to limit sensitivity to changes in a single training sample. The random noise is added based on Gaussian noise.

[0088] Step S02: normalizing the target original domain name according to a character-level dynamic encoding algorithm to obtain a semantic representation of the target domain name;

[0089] It should be noted that in this embodiment, the character-level dynamic encoding algorithm extracts effective features of the domain name sequence in the target original domain name based on character-level embedding, the normalization process is used to convert the target original domain name into a normalized input that can be processed by the domain name detection model, the target domain name semantic representation is a vector representation of fixed dimension, the target original domain name is normalized according to the character-level dynamic encoding algorithm, the character string of the target original domain name is split into a character sequence character by character, and each character in the character sequence is embedded to obtain a character embedding vector for the corresponding character;

[0090] Performing convolution processing on the character embedding vector according to a convolutional network layer to obtain character features;

[0091] Performing feature screening processing on the character features according to the maximum pooling layer to obtain significant character features;

[0092] Smoothing the significant character features according to the high-speed network layer to obtain smoothed character features;

[0093] Converting the smoothed character features into a vector representation of a fixed dimension to obtain a semantic representation of the target domain name;

[0094] The specific formula of the character-level dynamic encoding algorithm is as follows:

[0095] ,

[0096] ,

[0097] ,

[0098] in, Represents character features, represents the convolution kernel number, represents the activation function, represents the weight matrix of the convolution kernel, represents consecutive character embedded subsequences, represents the bias term, Indicates significant character features, represents the input feature of the maximum pooling layer, represents smooth character features, represents the gate control unit, represents the nonlinear transformation of the high-speed network layer, Represents the input features of the highway network layer.

[0099] Step S03: performing feature extraction processing on the semantic representation of the target domain name according to the multi-scale text convolutional neural network to obtain multi-scale local spatial features;

[0100] It should be noted that in this embodiment, the multi-scale text convolutional neural network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer. Feature extraction processing is performed on the semantic representation of the target domain name according to the multi-scale text convolutional neural network. The multi-scale text convolutional neural network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer.

[0101] Extract local features of different scales based on multi-size convolution kernels;

[0102] Performing local maximum pooling processing on the local features of different scales to reduce the feature dimensions of the local features and obtain significant local features of different scales;

[0103] Fuse the significant local features at different scales to obtain multi-scale local spatial features;

[0104] The specific algorithm of the multi-scale text convolutional neural network is as follows:

[0105] ,

[0106] ,

[0107]

[0108] in, Represents local features at different scales, represents the activation function, represents the convolution operation, Represents the input text sequence, represents the convolution kernel size, Indicates the output channel, represents the size of a single convolution kernel, Represents significant local features at different scales, Indicates the batch size, Indicates the length of the input text sequence, Represents multi-scale local spatial features, Represents the characteristic scale.

[0109] Step S04: performing feature extraction processing on the multi-scale local spatial features according to the bidirectional long short-term memory attention enhancement network to obtain global temporal features;

[0110] It should be noted that in this embodiment, the bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is used to capture context information from the sequence of the target domain name semantic representation forward to backward. The attention enhancement mechanism is used to focus on the key characters in the target domain name semantic representation. The step of performing feature extraction processing on the multi-scale local spatial features according to the bidirectional long short-term memory attention enhancement network to obtain global temporal features is as follows:

[0111] Context information is extracted from multi-scale local spatial features according to a bidirectional long short-term memory mechanism to obtain temporal features, wherein the context information extraction performs a bidirectional traversal of the text sequence in forward and backward order;

[0112] The specific formula of the bidirectional long short-term memory mechanism is as follows:

[0113] ,

[0114] ,

[0115] in, Represents the time series characteristics, represents a bidirectional long short-term memory mechanism, Represents multi-scale local spatial features, Indicates the batch size, represents the characteristic scale, and Represent the forward and backward time series features respectively, and represent forward and backward traversal respectively, Indicates time sequence;

[0116] Performing attention enhancement processing on the temporal features according to an attention enhancement mechanism to obtain global temporal features;

[0117] The attention enhancement mechanism is used to enhance the temporal features. The temporal features of each time step are calculated based on a single-layer linear transformation and a nonlinear activation function to obtain the importance score.

[0118] Then calculate the attention weight according to the Softmax function;

[0119] Performing attention weighting processing on the time series feature according to the attention weight to obtain a global time series feature;

[0120] The specific formula of the attention enhancement mechanism is as follows:

[0121] ,

[0122] in, Represents the global timing characteristics, represents a time sequence number, represents the attention weight, represents the temporal characteristics of a single time step, Indicates the batch size.

[0123] Step S05: performing feature fusion on the multi-scale local spatial features and the global temporal features to obtain domain name fusion features, and performing classification and recognition based on the domain name fusion features;

[0124] It should be noted that the malicious domain name detection algorithm based on the domain name detection model (CharacterBert-TextCNN-BiLSTM-ATT) in this embodiment is a comprehensive detection method that integrates the advantages of multiple deep neural network structures. First, the CharacterBert module is used to perform character-level encoding and semantic representation on the input domain name string, and the original domain name text is converted into a character-level vector matrix rich in contextual semantics. Subsequently, these encoded feature vector matrices are input into the TextCNN module for convolution operation to further capture the local key character patterns or short fragment features implicit in the domain name, and effectively identify abnormal character sequences or suspicious patterns through the convolution kernel. Next, the BiLSTM-ATT module is used to further capture the semantic relationship before and after the domain name sequence and the dependency between the features. The attention mechanism is used to highlight important features to obtain more accurate feature expression.

[0125] In summary, according to the above-mentioned DGA malicious domain name detection method based on federated learning and deep learning, by designing a domain name detection model based on a federated learning framework, participants are allowed to independently train models on local data sets without directly sharing original data, thereby protecting data privacy. By regularly sending model parameters or gradient information to a central server for aggregation and updating, global model optimization is achieved, improving the performance of detecting DGA malicious domain names, and being able to fully utilize part of the DGA malicious domain name data set held by each participant. Through the collaborative training mechanism of the federated learning framework, each participant can share the optimization results of the model, accelerate the training speed of the model, and improve the generalization ability of the model. The combination of a character-level dynamic encoding algorithm, a multi-scale text convolutional neural network, and a bidirectional long short-term memory attention enhancement network greatly improves the recognition accuracy and stability of DGA malicious domain names. The present invention improves the accuracy and generalization ability of DGA malicious domain name detection. Specifically, the target original domain name is obtained and input into the domain name detection model. The domain name detection model is based on the federated learning framework. The domain name detection model includes a data preprocessing module, a feature extraction module and a feature fusion module. It allows participants to independently train the model on the local data set without directly sharing the original data, thereby protecting data privacy and being able to make full use of part of the DGA malicious domain name data set held by each participant. Through the collaborative training mechanism of the federated learning framework, the generalization ability of the model is improved. The target original domain name is normalized according to the character-level dynamic encoding algorithm to obtain the target domain name semantic representation. The character-level dynamic encoding algorithm extracts the effective features of the domain name sequence in the target original domain name according to the character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by the domain name detection model. The target domain name semantic representation is a vector representation of fixed dimension, and the character-level convolutional network layer is used to dynamically generate the embedding representation of the whole word, so that the domain name detection model can It has more advantages when dealing with special text types, especially in terms of spelling errors, term variants and unregistered words, showing stronger robustness and generalization ability. Relying on character-level dynamic encoding, it can completely retain the original sequence information, avoid information loss caused by excessive segmentation, and further improve adaptability and generalization ability. The target domain name semantic representation is feature extracted according to the multi-scale text convolutional neural network to obtain multi-scale local spatial features. The multi-scale text convolutional neural network includes an input layer, a convolution layer, a pooling layer and a fully connected layer. Malicious domain names usually contain a large number of random characters, variant spellings or meaningless fragments injected by humans. Traditional methods based on feature rules or manual analysis are difficult to effectively handle these variations and new threats. The present invention is based on character-level convolution operations, which can automatically learn pattern features in domain name character sequences, and then identify potential abnormal or malicious feature combinations. It can effectively deal with the problems of spelling variants or domain name noise, especially when faced with a large amount of short and irregular domain name data.It shows higher robustness and generalization ability, and can automatically learn malicious feature patterns in domain names without manually defining features, so that it is more adaptable to rapidly changing malicious domain name detection scenarios, effectively improving the accuracy and real-time performance of malicious domain name recognition. The multi-scale local spatial features are feature extracted and processed according to the bidirectional long short-term memory attention enhancement network to obtain global temporal features. The bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is used to capture contextual information from the sequence of the target domain name semantic representation forward to backward, and the attention enhancement mechanism is used to focus on the key points in the semantic representation of the target domain name. The model focuses on key characters and simultaneously traverses the text sequence in both the forward and backward directions to more comprehensively capture the context. The bidirectional structure can better perceive the semantic association between the preceding and following characters in the domain name string, and further highlight the characteristics of abnormal characters or key character sequences through the attention mechanism. The dynamic focusing capability enables the model to show higher recognition ability for variant domain names, randomly generated domain names, and human interference features, further improving recognition accuracy. The multi-scale local spatial features and the global temporal features are fused to obtain domain name fusion features, and classification and recognition are performed based on the domain name fusion features. This invention improves the accuracy and generalization ability of DGA malicious domain name detection.

[0126] See also Figure 2 , which is a schematic diagram of the structure of a DGA malicious domain name detection system based on federated learning and deep learning proposed in the second embodiment of the present invention, and includes:

[0127] A preprocessing module 10 is used to obtain the target original domain name and input it into a domain name detection model. The domain name detection model is based on a federated learning framework and includes a data preprocessing module, a feature extraction module, and a feature fusion module.

[0128] Normalizing the target original domain name according to a character-level dynamic encoding algorithm to obtain a semantic representation of the target domain name. The character-level dynamic encoding algorithm extracts effective features of the domain name sequence in the target original domain name based on character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by a domain name detection model. The semantic representation of the target domain name is a vector representation of fixed dimension.

[0129] A local feature extraction module 20 is configured to perform feature extraction processing on the semantic representation of the target domain name according to a multi-scale text convolutional neural network to obtain multi-scale local spatial features, wherein the multi-scale text convolutional neural network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer;

[0130] A global feature enhancement module 30 is configured to perform feature extraction processing on the multi-scale local spatial features based on a bidirectional long short-term memory attention enhancement network to obtain global temporal features. The bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is configured to capture contextual information from the sequence of the target domain name semantic representation forward and backward, and the attention enhancement mechanism is configured to focus attention on key characters in the target domain name semantic representation.

[0131] The feature fusion module 40 is configured to fuse the multi-scale local spatial features and the global temporal features to obtain domain name fusion features, and perform classification and recognition based on the domain name fusion features.

[0132] The present invention also proposes a computer storage medium on which one or more programs are stored, which, when executed by a processor, implement the above-mentioned DGA malicious domain name detection method based on federated learning and deep learning.

[0133] The present invention also proposes a computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the above-mentioned DGA malicious domain name detection method based on federated learning and deep learning.

[0134] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device.

[0135] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.

[0136] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0137] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0138] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A DGA malicious domain name detection method based on federated learning and deep learning, characterized by: include: Obtain the target original domain name and input it into a domain name detection model. The domain name detection model is based on a federated learning framework and includes a data preprocessing module, a feature extraction module, and a feature fusion module. Normalizing the target original domain name according to a character-level dynamic encoding algorithm to obtain a semantic representation of the target domain name. The character-level dynamic encoding algorithm extracts effective features of the domain name sequence in the target original domain name based on character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by a domain name detection model. The semantic representation of the target domain name is a vector representation of fixed dimension. Performing feature extraction processing on the target domain name semantic representation according to a multi-scale text convolutional neural network to obtain multi-scale local spatial features, wherein the multi-scale text convolutional neural network includes an input layer, a convolution layer, a pooling layer, and a fully connected layer; Performing feature extraction processing on the multi-scale local spatial features according to a bidirectional long short-term memory attention enhancement network to obtain global temporal features, wherein the bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is used to capture contextual information from the sequence of the target domain name semantic representation forward to backward, and the attention enhancement mechanism is used to focus attention on key characters in the target domain name semantic representation; The multi-scale local spatial features and the global temporal features are subjected to feature fusion to obtain domain name fusion features, and classification and recognition are performed according to the domain name fusion features.

2. The DGA malicious domain name detection method based on federated learning and deep learning according to claim 1 is characterized in that: The federated learning framework is as follows: The domain name detection model is based on a federated learning framework, wherein the federated aggregation algorithm of the federated learning framework is based on a federated performance averaging algorithm, and the federated performance averaging algorithm includes a dynamic weighting mechanism; The specific formula of the federated performance averaging algorithm is as follows: , , , , in, represents the global model parameters after aggregation, represents a time sequence number, Indicates the total number of clients, Indicates the client ordinal number, represents the comprehensive weight, Indicates the total amount of client data. Represents a single client data, represents the performance adjustment factor, Represents the client performance score, represents the scoring weight, Indicates the current accuracy of the client. Indicates the current F1 value of the client.

3. The DGA malicious domain name detection method based on federated learning and deep learning according to claim 2 is characterized in that: The federated learning framework also includes: Privacy protection is performed based on a differential privacy algorithm. The differential privacy algorithm adds pre-calibrated random noise to sensitive information during training to limit sensitivity to changes in a single training sample. The random noise is added based on Gaussian noise.

4. The DGA malicious domain name detection method based on federated learning and deep learning according to claim 1 is characterized in that: The step of normalizing the target original domain name according to the character-level dynamic encoding algorithm to obtain the semantic representation of the target domain name is as follows: Normalizing the target original domain name according to a character-level dynamic encoding algorithm, splitting the target original domain name string into a character sequence character by character, and embedding each character in the character sequence to obtain a character embedding vector for the corresponding character; Performing convolution processing on the character embedding vector according to a convolutional network layer to obtain character features; Performing feature screening processing on the character features according to the maximum pooling layer to obtain significant character features; Smoothing the significant character features according to the high-speed network layer to obtain smoothed character features; Converting the smoothed character features into a vector representation of a fixed dimension to obtain a semantic representation of the target domain name; The specific formula of the character-level dynamic encoding algorithm is as follows: , , , in, Represents character features, represents the convolution kernel number, represents the activation function, represents the weight matrix of the convolution kernel, represents consecutive character embedded subsequences, represents the bias term, Indicates significant character features, represents the input feature of the maximum pooling layer, represents smooth character features, represents the gate control unit, represents the nonlinear transformation of the high-speed network layer, Represents the input features of the highway network layer.

5. The DGA malicious domain name detection method based on federated learning and deep learning according to claim 1 is characterized in that: The step of performing feature extraction processing on the semantic representation of the target domain name according to the multi-scale text convolutional neural network to obtain multi-scale local spatial features is specifically as follows: Performing feature extraction processing on the semantic representation of the target domain name according to a multi-scale text convolutional neural network, wherein the multi-scale text convolutional neural network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer; Extract local features of different scales based on multi-size convolution kernels; Performing local maximum pooling processing on the local features of different scales to reduce the feature dimensions of the local features and obtain significant local features of different scales; Fuse the significant local features at different scales to obtain multi-scale local spatial features; The specific algorithm of the multi-scale text convolutional neural network is as follows: , , in, Represents local features at different scales, represents the activation function, represents the convolution operation, Represents the input text sequence, represents the convolution kernel size, Indicates the output channel, represents the size of a single convolution kernel, Represents significant local features at different scales, Indicates the batch size, Indicates the length of the input text sequence, Represents multi-scale local spatial features, Represents the characteristic scale.

6. The DGA malicious domain name detection method based on federated learning and deep learning according to claim 1 is characterized in that: The step of performing feature extraction processing on the multi-scale local spatial features according to the bidirectional long short-term memory attention enhancement network to obtain global temporal features is specifically as follows: Context information is extracted from multi-scale local spatial features according to a bidirectional long short-term memory mechanism to obtain temporal features, wherein the context information extraction performs a bidirectional traversal of the text sequence in forward and backward order; The specific formula of the bidirectional long short-term memory mechanism is as follows: , , in, Represents the time series characteristics, represents a bidirectional long short-term memory mechanism, Represents multi-scale local spatial features, Indicates the batch size, represents the characteristic scale, and Represent the forward and backward time series features respectively, and represent forward and backward traversal respectively, Indicates time sequence; Attention enhancement processing is performed on the temporal features according to an attention enhancement mechanism to obtain global temporal features.

7. The DGA malicious domain name detection method based on federated learning and deep learning according to claim 6 is characterized in that: The step of performing attention enhancement processing on the time series features according to the attention enhancement mechanism to obtain global time series features specifically includes: The attention enhancement mechanism is used to enhance the temporal features. The temporal features of each time step are calculated based on a single-layer linear transformation and a nonlinear activation function to obtain the importance score. Then calculate the attention weight according to the Softmax function; Performing attention weighting processing on the time series feature according to the attention weight to obtain a global time series feature; The specific formula of the attention enhancement mechanism is as follows: , in, Represents the global timing characteristics, represents a time sequence number, represents the attention weight, represents the temporal characteristics of a single time step, Indicates the batch size.

8. A DGA malicious domain name detection system based on federated learning and deep learning, characterized by: include: A preprocessing module is used to obtain the target original domain name and input it into a domain name detection model. The domain name detection model is based on a federated learning framework and includes a data preprocessing module, a feature extraction module, and a feature fusion module. Normalizing the target original domain name according to a character-level dynamic encoding algorithm to obtain a semantic representation of the target domain name. The character-level dynamic encoding algorithm extracts effective features of the domain name sequence in the target original domain name based on character-level embedding. The normalization process is used to convert the target original domain name into a normalized input that can be processed by a domain name detection model. The semantic representation of the target domain name is a vector representation of fixed dimension. A local feature extraction module is used to perform feature extraction processing on the semantic representation of the target domain name according to a multi-scale text convolutional neural network to obtain multi-scale local spatial features. The multi-scale text convolutional neural network includes an input layer, a convolution layer, a pooling layer, and a fully connected layer; a global feature enhancement module, configured to perform feature extraction processing on the multi-scale local spatial features based on a bidirectional long short-term memory attention enhancement network to obtain global temporal features, wherein the bidirectional long short-term memory attention enhancement network includes a bidirectional long short-term memory mechanism and an attention enhancement mechanism. The bidirectional long short-term memory mechanism is configured to capture contextual information from the sequence of the target domain name semantic representation forward to backward, and the attention enhancement mechanism is configured to focus attention on key characters in the target domain name semantic representation; The feature fusion module is used to fuse the multi-scale local spatial features and the global temporal features to obtain domain name fusion features, and perform classification and recognition based on the domain name fusion features.

9. A storage medium, characterized in that: The storage medium stores one or more programs, which, when executed by the processor, implement the DGA malicious domain name detection method based on federated learning and deep learning as described in any one of claims 1 to 7.

10. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, it implements the DGA malicious domain name detection method based on federated learning and deep learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • DGA domain name detection method oriented to botnet of Internet of Things

    CN116633623A

  • Malicious DGA domain name detection method and system

    CN117834292A