An Early Malicious Traffic Detection Method for the Internet of Things Combining Domain Name and DNS Traffic Characteristics

Through deep learning methods combining domain names and DNS traffic characteristics, DNS query packet characteristics of IoT devices are extracted, which solves the problem of malicious traffic detection lag in the early stage of IoT devices in the prior art, and achieves efficient and accurate malicious traffic detection.

CN119484036BActive Publication Date: 2025-05-27SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411491458.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-05-27
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The prior art is difficult to detect malicious traffic in IoT devices in early stages, resulting in a certain lag in the detection results.

Method used

By combining domain name and DNS traffic characteristics, using deep learning methods based on attention-based long and short-term memory neural networks and convolutional neural networks, the multi-dimensional features of DNS query data packets are extracted to realize the detection of early malicious traffic in the Internet of Things.

Benefits of technology

Accurate detection of early malicious traffic of IoT devices has been achieved, with the detection accuracy rate reaching more than 98%, improving detection efficiency and shortening training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484036B_ABST
    Figure CN119484036B_ABST
Patent Text Reader

Abstract

The present invention proposes an early malicious traffic detection method for the Internet of Things that combines domain names and DNS traffic characteristics, which is divided into four parts. The first part is traffic capture and filtering; the second part is DNS traffic information extraction, the third part is the training of the detection model, and the fourth part is the detection of malicious traffic in the Internet of Things. The specific content is that after extracting traffic characteristics such as Internet of Things DNS domain names and query information, the long short-term memory model with a multi-head attention mechanism is used to learn DNS domain name characteristics, and the convolutional neural network is used to extract DNS traffic characteristics to detect in the early stage of malicious traffic occurrence. The method proposed by the present invention can quickly and efficiently detect malicious traffic in the Internet of Things in the early stage, with an identification accuracy rate reaching 98%, and the model training time and detection time are short. The method for early malicious traffic detection facilitates network managers to quickly respond and protect Internet of Things devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cyberspace security, and relates to an early malicious traffic detection method for the Internet of Things (IoT) that combines domain name and DNS traffic characteristics. Background Art

[0002] The Internet of Things refers to a system in which various physical devices, sensors, software, and other technologies are interconnected through a network for data exchange and communication. These devices generally have the capabilities of perception, computing, processing, and communication, and are capable of collecting and transmitting information in real time.

[0003] It is reported that in terms of regional distribution, the number of IoT devices in China exceeds 500 million, and the number of IoT devices in North America and Europe respectively exceeds 300 million, which proves that IoT technology has been widely adopted and implemented globally, and it is expected that the number and scale of IoT devices will continue to expand in the future. Along with its rapid development in scale, the potential security issues of devices have gradually attracted people's attention. Currently, there are serious security vulnerabilities in IoT devices. Due to the unique nature of lightweight and low-cost IoT devices, their operating systems mostly use customized systems, which are rarely maintained and updated, and cannot provide sufficient protection against modern viruses and vulnerabilities. Secondly, general users will neglect the maintenance of the account passwords of IoT devices, and attackers can easily control these IoT devices by using weak password login credentials and default account passwords. In addition, many IoT device vendors produce devices with weak authentication policies, which are easily brute-forced by weak password combinations, and thus can be exploited by attackers using known vulnerabilities to control the devices and achieve their attack purposes.

[0004] To defend against malware attacks on IoT devices, existing research on IoT devices mostly relies on the statistical characteristics of malicious traffic, and identifies potential abnormal traffic by deeply mining the behavioral characteristics of abnormal traffic. However, this method mainly depends on the acquisition of complete original network traffic and the analysis of statistical characteristics, which greatly limits its ability in early detection. It needs to wait until the attack flow ends or after a certain time window to complete the detection, resulting in a certain lag in its detection results.

[0005] Therefore, the present invention filters the DNS query traffic of early communication of malware in the network. This method focuses on DNS traffic, combines DNS domain name characteristics and DNS traffic characteristics to form a DNS query feature set, and uses deep learning to train the model, which is more conducive to real-time threat detection. At the same time, it can more lightly implement the detection of early malicious traffic in the IoT, and improve the efficiency of detecting that IoT devices are infected. Summary of the Invention

[0006] To effectively identify malicious traffic and achieve early detection of malicious traffic in the Internet of Things (IoT), the present invention proposes an early malicious traffic detection method for the IoT that combines domain names and DNS traffic characteristics. Aiming at the problem that it is difficult to achieve early detection of malicious traffic in the IoT in existing research, an attention-based early malicious traffic detection method for the IoT is proposed. This method is based on the DNS query traffic in the early stage of malware communication, and converts the malicious traffic detection problem into a binary classification problem of multi-dimensional features of DNS query packets. First, extract the text features of the query domain name and the headers and connection features of DNS query packets; then, use a long short-term memory neural network based on the multi-head attention mechanism to fully mine the distinguishable features of domain name samples, and use a convolutional neural network to extract the key information of DNS query packets; finally, splice the two features and use a fully connected layer to classify the feature vectors to achieve traffic detection.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] An early malicious traffic detection method for the IoT that combines domain names and DNS traffic characteristics, comprising the following steps:

[0009] (1) From the perspective of identifying malicious traffic in the IoT at an early stage, start with the IoT DNS query traffic, analyze the DNS query traffic characteristics, and use a traffic extraction tool to screen and filter to obtain DNS traffic;

[0010] (2) Analyze the filtered DNS traffic, select domain name features and traffic features respectively, and obtain the text features of the query domain name and the corresponding DNS query packet headers and connection features;

[0011] (3) Use deep learning methods to achieve the detection of early malicious traffic in the IoT, construct a dataset to train the model and evaluate the model;

[0012] (4) Input the IoT traffic to be detected into the trained model for detection.

[0013] Further, the step (1) specifically includes the following sub-steps:

[0014] (1.1) Use a network analysis tool to monitor network traffic and capture packets on an intermediate router;

[0015] (1.2) Use a network analysis tool command to screen and filter to obtain DNS traffic.

[0016] Further, the step (2) specifically includes the following sub-steps:

[0017] (2.1) Use a network traffic analysis tool to analyze the DNS query packets obtained in step (1.2), and extract the timing features and connection features of the DNS traffic;

[0018] (2.2) Use the network traffic analysis tool ZeeK to obtain the domain name query features of DNS queries, and combine them with the features obtained in step (2.1) to form a DNS query feature set.

[0019] Furthermore, the features in the DNS query feature set extracted by the present invention are shown in Table 1 as follows:

[0020] Table 1:

[0021]

[0022] Furthermore, step (3) specifically includes the following sub-steps:

[0023] (3.1) Process the obtained DNS query feature set and divide it into a set for model training and a test set; preprocess the DNS query feature set obtained in step (2.2).

[0024] (3.2) For the domain name feature part, first clean the domain name text and remove non-alphanumeric characters in the domain name. Then construct a dictionary for the characters, with a total of 38 characters. Replace each character with the corresponding index so that the text data is converted into a numerical format that can be processed by the model. Finally, keep a fixed length of 40 for the domain name feature sequence, and fill the shorter sequences with 0 to ensure that all input sequence lengths are the same.

[0025] (3.3) For the DNS connection feature and query feature parts, first use the interpolation technique of median filling to fill in the missing values, and use the one-hot encoding technique for non-numerical fields to convert the values that cannot be recognized by the model into the corresponding numerical form. Finally, perform normalization processing on the numerical values to map the data values to [0,1] to eliminate the scale differences caused by different dimensions between different features.

[0026] (3.4) Use a long short-term memory model with an attention mechanism to fully mine the distinguishable features of domain name samples, and use the multi-head attention mechanism to process the input text sequence. Each attention head calculates the attention score from a different subspace, assigns higher weights to the important parts of the sequence, and captures the key information distributed at different positions in the sequence.

[0027] (3.5) Use a one-dimensional convolutional neural network to extract DNS traffic features, and capture the key information and abnormal behaviors in the connection process through multi-layer convolution and pooling operations; use two convolutional layers, two pooling layers, 1 Flatten layer, 1 Dropout layer and a fully connected layer to extract the selected DNS traffic features.

[0028] (3.6) Concatenate the domain name text features extracted by the multi-head attention LSTM and the DNS traffic features extracted by the CNN, and use a classifier combined with a stacked fully-connected layer structure composed of three fully-connected layers stacked in sequence and a softmax classification layer to perform the prediction task;

[0029] (3.7) Use the training set for model training, and use the test set to update the model parameters; and use other deep learning algorithms: LSTM and CNN to compare the used method from multiple metrics.

[0030] Further, the step (4) specifically includes the following sub-steps:

[0031] (4.1) Obtain features for the unrecognized early-stage malicious traffic of the Internet of Things by referring to the same processing process as in (2.2) to obtain the DNS query feature set of the unrecognized traffic;

[0032] (4.2) Input the obtained DNS query feature set into the model for detection, and judge whether it is malicious traffic of the Internet of Things according to the output result, so as to detect malicious traffic of the Internet of Things in the early infection stage.

[0033] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for detecting early-stage malicious traffic of the Internet of Things by combining domain names and DNS traffic features.

[0034] A computer-readable storage medium stores computer instructions thereon. When the computer instructions are executed by a processor, they implement the method for detecting early-stage malicious traffic of the Internet of Things by combining domain names and DNS traffic features.

[0035] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0036] (1) The present invention can effectively identify malicious software traffic in the early stage of malicious software infringement, and the detection accuracy rate reaches more than 98%. It can detect malicious traffic of the Internet of Things in the early infection stage.

[0037] (2) The present invention is more lightweight compared to the conventional use of complete malicious traffic of the Internet of Things, effectively improving the efficiency of detecting malicious traffic of the Internet of Things.

[0038] (3) The present invention proposes a new feature selection metric method, innovatively combines the DNS query domain names and DNS query traffic features of the Internet of Things to identify malicious traffic of the Internet of Things, and the proposed feature extraction and selection method improves the model efficiency.

[0039] (4) When the method of the present invention is compared with the current mainstream models, under the condition that the detection accuracy rate, precision rate and other indicators are on average the same, the training time is shortened by nearly half, greatly saving the training time and the time used for detection; while compared with the models with similar training time to that of the present invention, the detection accuracy rate, precision rate and other indicators of this method are improved by 1%. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 Schematic diagram of the recognition framework;

[0041] Figure 2 Comparison of the accuracy rates of each model in the ablation experiment;

[0042] Figure 3 Confusion matrix results of an early malicious traffic detection method for the Internet of Things combining domain names and DNS traffic characteristics;

[0043] Figure 4 Comparison of the training time and accuracy rate with other mainstream models;

[0044] Figure 5 Comparison of the detection time consumption with other mainstream models. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The technical solution provided by the present invention will be described in detail below in combination with specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0046] Embodiment 1: The present invention proposes an early malicious traffic detection method for the Internet of Things combining domain names and DNS traffic characteristics. The detection framework diagram is as Figure 1 shown, which is divided into four parts. The first part is to screen DNS traffic. Specifically, starting from the perspective of early identification of malicious traffic in the Internet of Things, starting from the Internet of Things DNS query traffic and domain names, DNS traffic is screened and filtered through the use of traffic extraction tools. The second part is to extract features from DNS traffic. Specifically, the Zeek tool is used to batch parse DNS traffic packets to obtain key log files, including the detailed information of DNS queries and the details of network connections, so as to reveal abnormal DNS activities caused by malware or data leakage. The third part uses deep learning methods to achieve the detection of early malicious traffic in the Internet of Things. Specifically, a long short-term memory model with an attention mechanism is used to mine the distinguishable features of domain name samples, and a convolutional neural network is used to extract DNS traffic features to distinguish normal traffic and malicious Internet of Things DNS traffic. The fourth part is to detect early malicious traffic in the Internet of Things. Specifically, the Internet of Things traffic to be detected is input into the model for detection.

[0047] Specifically, an efficient detection method for early malicious traffic in the Internet of Things includes the following steps:

[0048] (1) Deploy traffic extraction tools at key network nodes to screen and filter out IoT DNS traffic.

[0049] The specific process of this step is as follows:

[0050] (1.1) Deploy packet capture tools at key nodes such as intermediate routers, and capture packets in network traffic in real time through tools such as Wireshark and Tshark.

[0051] (1.2) Use the Tshark tool command to screen and filter out DNS traffic, and run the command: "tshark -Y \"dns\" -T fields -e dns.qry.name -e dns.qry.type -e frame.time -r [pcap] > dns.pcap".

[0052] (2) Extract features from DNS traffic to obtain key log files, including detailed information of DNS queries and details of network connections, so as to reveal abnormal DNS activities caused by malware or data leakage.

[0053] The specific process in this step is as follows:

[0054] (2.1) Use the Zeek network traffic analysis tool to deeply analyze the traffic packets, and run the command: "zeek -r [dns.pcap] > conn.log, dns.log" to obtain the key log files dns.log and conn.log. Among them, the dns.log file records the detailed information of DNS queries, including query timestamps, domain names, query types, response codes, etc., and the conn.log file records the details of network connections, such as the start and end times of connections, source and destination IP addresses and ports, protocol types, connection status, etc.

[0055] (2.2) According to the log files obtained in (2.1), extract and construct a DNS query feature set. The specific process includes traversing each record in dns.log and conn.log, and adding the corresponding key fields to the DNS feature list after finding them. The specific fields and corresponding meanings of the extracted DNS query feature set are shown in Table 2

[0056] Table 2:

[0057]

[0058] (3) After completing and obtaining the DNS query feature set, use the feature set as the data set and divide it into a training set and a test set. Use a long short-term memory model with an attention mechanism to fully mine the distinguishable features of domain name samples, use a convolutional neural network to extract DNS traffic features, and finally, splice the two features and use a fully connected layer to classify the feature vectors to achieve traffic detection.

[0059] The specific process in this step is as follows:

[0060] (3.1) Process the obtained DNS query feature set, and the processing methods for the training set and the test set are the same.

[0061] (3.2) For the domain name features, first clean the domain name text, including removing non-alphanumeric characters in the domain name, such as underscores, hyphens, and other special symbols, to reduce the noise of the input data. Build a dictionary for the characters, with a total of 38 characters. And replace each character with the corresponding index, so that the text data is converted into a numerical format that the model can process. Considering the diverse lengths of the characters, retain a fixed length of 40 for the input feature sequences, and use 0 as the padding value in the shorter sequences to ensure that all input sequence lengths are the same.

[0062] (3.3) For the DNS connection features and query features, for the features selected in Table 2, first use the interpolation technique of median filling to fill in the missing values, and use the one-hot encoding technique for non-numerical fields to convert the values that cannot be recognized by the model into the corresponding numerical forms. Finally, perform normalization processing on the numerical values to map the data values to [0,1] to eliminate the scale differences caused by different dimensions between different features.

[0063] (3.4) For the model part, use a long short-term memory model with an attention mechanism to fully mine the distinguishable features of domain name samples. Specifically, use the multi-head attention mechanism to effectively capture the time series dependence and its key information in the domain name text data. Use a convolutional neural network to extract DNS traffic features.

[0064] Specifically, the LSTM model based on the multi-head attention mechanism can more effectively capture the time series dependence and its key information in the domain name text data. First, input the domain name character features obtained after preprocessing in step (2) into the word embedding layer to map the domain name features to a higher dimension, and obtain a high-dimensional domain name feature representation with a fixed dimension through a word embedding layer. Then, process the domain name text data in parallel through different attention heads, and each attention head focuses on capturing different aspects of the data, such as semantic depth and context connection. The attention calculation formula is as follows:

[0065]

[0066] Among them, Q represents the query item matrix, K represents the key matrix, V represents the value matrix, and d K is the dimension of the key vector, and softmax(·) is the normalized exponential function. The multi-head attention mechanism introduces representations in multiple subspaces based on self-attention. The specific calculation formula is shown in the following two formulas:

[0067] head i = Attention(QW i Q , KW i K , VW i V )

[0068] MultiHead(Q, K, V) = Concat(head 1 , …, head h )

[0069] Among them, QWi i Q , W i K and VW i V are learnable weight matrices used to transform the input into an appropriate dimension. After the query domain name text vector passes through the multi-head attention channel, it is input into the LSTM layer to further enhance the model's memory ability for text sequence data. LSTM effectively manages the information flow through its unique gating mechanism, that is, while retaining long-term dependency information, it filters out unnecessary short-term noise, which is particularly suitable for processing DNS query data that requires long-term memory and complex time-dependent parsing.

[0070] Specifically, a one-dimensional convolutional neural network is used to analyze the connection features and query features extracted from DNS traffic, capturing key information and abnormal behaviors during the connection process. Through multi-layer convolution and pooling operations, the model can extract key features from high-dimensional data. The present invention uses two convolutional layers, two pooling layers, 1 Flatten layer, one Dropout layer and a fully connected layer to extract the selected DNS traffic features.

[0071] Among them, the convolutional layer is used to perform a convolution operation on the input data through a set of learned filters to extract local features in the DNS traffic data. And max pooling is used to divide the input data area into several rectangular areas and output the maximum value of each area, which is used to capture more abstract-level feature representations and reduce the spatial dimension of the feature map, improving the computational efficiency of the model. The Flatten layer converts the multi-dimensional DNS features output by the convolutional layer into a one-dimensional vector. This process can be simply represented by the following formula:

[0072] Flatten(D, H, W) = Vector(D · H · W)

[0073] Where D, H, and W represent the sizes of different dimensions respectively. The Dropout layer randomly discards a part of the neuron data in the network to prevent the model from overfitting. Finally, the fully connected layer is used to perform weighted summation on the output DNS feature vectors of the Dropout layer to achieve the mapping from feature extraction to decision output.

[0074] Finally, there is the classification module. Two types of features are concatenated and the fully connected layer is used to classify the feature vectors to achieve traffic detection. The output feature vectors of the multi-head attention LSTM and CNN come from two different dimensions. In the present invention, the domain name text features extracted by the multi-head attention LSTM and the DNS traffic features extracted by the CNN are concatenated to form a vector set containing both types of features. A classifier composed of a stacked fully connected layer structure formed by sequentially stacking three fully connected layers and a softmax classification layer is used for the prediction task. The expression after feature fusion is as follows:

[0075] merge = Concatenate(A, B)

[0076] Where A and B are two sets of features, and merge is the feature representation after concatenating the features. The fully connected layer extracts and integrates the concatenated feature vectors and converts them to a lower dimension while ensuring the retention of key information. Then the features are handed over to the softmax function to map the feature vectors to the interval [0, 1] and perform normalization processing on the output of the entire vector to obtain the final prediction result.

[0077] The hyperparameter settings of the model used in the present invention are as follows: The hyperparameter settings of the model include the following: A total of 80 rounds were carried out during the model training process, the batch size was 32, and the learning rate was 0.0001. The model adopted the Dropout mechanism with a coefficient set to 0.3, selected Adam as the optimizer, and used sparse categorical cross-entropy as the loss function. In the LSTM layer based on the multi-head attention mechanism, the input dimension was set to 95, the word embedding dimension was 100, the number of attention heads was 10, and the hidden unit size and output dimension were both 128. For the one-dimensional convolutional neural network part, two convolutional layers were used, each with 64 filters, a kernel size of 2, and the activation function was ReLU. Pooling layers were used after the two convolutional layers respectively, with a window size and a stride of 3. The convolutional neural network layer finally adopted a fully connected layer as the feature output layer, with an output dimension of 16 and the activation function being softmax(). The Concatenate() feature fusion layer was used to fuse the outputs of the multi-head attention LSTM and the CNN. In the classification module, three fully connected layers were used as the hidden layer, and RELU was used as the activation function after the first two fully connected layers. The last fully connected layer was connected with softmax as the activation function to obtain the final output.

[0078] The specific ablation experiment results are as Figure 2 shown.

[0079] (3.5) Evaluate the accuracy, training time, and detection time of the method for early malicious traffic detection in the Internet of Things by combining domain names and DNS traffic characteristics. Compared with the current mainstream models, when the detection accuracy and other indicators are the same, the training time and the test time of the method of the present invention are greatly reduced, and the test time is reduced by nearly a hundred times. The reason is that compared with other models, this model only selects DNS traffic, which improves the efficiency of the model; similarly, compared with the model that only uses DNS traffic, when the training time is approximately the same, the accuracy of this model is increased by 1%. The specific detection results and time are as Figure 3 、 Figure 4 、 Figure 5 shown.

[0080] (4) Detect early malicious traffic in the Internet of Things

[0081] This step specifically includes the following processes:

[0082] (4.1) Process the unrecognized early malicious traffic in the Internet of Things according to the same processing process as in (2.2) to obtain the DNS query feature set of the unrecognized traffic.

[0083] (4.2) Input the obtained DNS query feature set into the model for detection, and let the model judge whether it is malicious traffic.

[0084] Embodiment 2: An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for detecting early malicious traffic in the Internet of Things by combining domain names and DNS traffic characteristics as described above.

[0085] Embodiment 3: A computer-readable storage medium stores computer instructions. When the computer instructions are executed by a processor, they implement the method for detecting early malicious traffic in the Internet of Things by combining domain names and DNS traffic characteristics as described above.

[0086] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. An IoT early malicious traffic detection method combining domain name and DNS traffic characteristics, characterized in that: The steps include: (1) Based on the perspective of early identification of malicious IoT traffic, we start with the IoT DNS query traffic and use traffic extraction tools to screen and filter the DNS traffic; (2) Analyze the filtered DNS traffic, select domain name features and traffic features respectively, and obtain the text features of the query domain name and the corresponding DNS query data packet header and connection features; Step (2) analyzes the filtered DNS traffic, which specifically includes the following sub-steps: (2.1) The network traffic analysis tool analyzes the DNS query message obtained in step (1.2) to extract the timing characteristics and connection characteristics of the DNS traffic; (2.2) Using the network traffic analysis tool ZeeK, obtain the domain name query features of the DNS query, and combine them with the features obtained in step (2.1) to form a DNS query feature set; The features extracted from the DNS query feature set in step (2.2) are shown in the following table: (3) Use deep learning methods to detect early malicious traffic in the Internet of Things, build data sets to train models, and evaluate models; Process the obtained DNS query feature set and divide it into model training and test sets; Preprocess the DNS query feature set obtained in step (2.2); A long short-term memory model based on an attention mechanism is used to fully exploit the distinguishable features of domain name samples. A multi-head attention mechanism is used to process the input text sequence. Each attention head calculates the attention score from a different subspace, assigning higher weights to important parts of the sequence and capturing key information distributed at different positions in the sequence. A one-dimensional convolutional neural network is used to extract DNS traffic features. Through multi-layer convolution and pooling operations, key information and abnormal behaviors in the connection process are captured. Two convolution layers, two pooling layers, one Flatten layer, one Dropout layer, and a fully connected layer are used to extract the selected DNS traffic features. The domain name text features extracted by the multi-head attention LSTM and the DNS traffic features extracted by the CNN are concatenated, and a classifier composed of a stacked fully connected layer structure consisting of three fully connected layers stacked in sequence and a softmax classification layer is used for the prediction task; The preprocessed training set is used for model training, and the test set is used to update the model parameters. Other deep learning algorithms are used: LSTM and CNN are used to compare the methods used from multiple indicators. (4) Input the IoT traffic to be tested into the trained model for testing.

2. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 1 is characterized in that: Step (1) uses a traffic extraction tool to screen and filter DNS traffic; specifically includes the following sub-steps: (1.1) Use network analysis tools to monitor network traffic and capture packets at intermediate routers; (1.2) Use network analysis tool commands to screen and filter DNS traffic.

3. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 1 is characterized in that: Preprocessing the DNS query feature set obtained in step (2.2) includes the following sub-steps: (1) For the domain name feature part, first clean the domain name text, remove non-alphanumeric characters in the domain name, and build a dictionary of characters, a total of 38 characters. Replace each character with the corresponding index so that the text data is converted into a numerical format that the model can process. Finally, the domain name feature sequence is kept at a fixed length of 40, and the shorter sequence is filled with 0 to ensure that the length of all input sequences is consistent; (2) For the DNS connection features and query features, we first use the median filling interpolation technique to fill in the missing values. For non-numeric fields, we use the independent hot encoding technique to convert the values ​​that cannot be recognized by the model into corresponding numerical forms. Finally, we use normalization processing for the numerical values ​​and map the data values ​​to [0, 1] to eliminate the scale differences caused by the different dimensions of different features.

4. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 3 is characterized in that: Step (3) specifically includes the following sub-steps: (3.1) Processing the obtained DNS query feature set and dividing it into a set for model training and a test set; preprocessing the DNS query feature set obtained in step (2.2); (3.2) For the domain name feature part, first clean the domain name text, remove non-alphanumeric characters in the domain name, and build a dictionary of characters, a total of 38 characters, replace each character with the corresponding index so that the text data is converted into a numerical format for model processing. Finally, the domain name feature sequence is kept at a fixed length of 40, and the shorter sequence is filled with 0 to ensure that the length of all input sequences is consistent; (3.3) For the DNS connection features and query features, we first use the median filling interpolation technique to fill the missing values. For non-numeric fields, we use the independent hot encoding technique to convert the values ​​that cannot be recognized by the model into corresponding numerical forms. Finally, we use normalization processing for the numerical values ​​and map the data values ​​to [0,1] to eliminate the scale differences caused by the different dimensions of different features.

5. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 4 is characterized in that: In step (3.4), the LSTM model based on the multi-head attention mechanism can more effectively capture the time series dependency and key information in the domain name text data. First, the domain name character features obtained after the preprocessing in step (2) are input into the word embedding layer, and the domain name features are mapped to a higher dimension to obtain a high-dimensional domain name feature representation. Then, the domain name text data is processed in parallel through different attention heads. The attention calculation formula is as follows: Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d K is the dimension of the key vector, softmax(·) is the normalized exponential function, and the multi-head attention mechanism introduces multiple subspace representations based on self-attention. The specific calculation formulas are shown in the following two formulas: MultiHead(Q,K,V)=Concat(head1,…,head h ) Among them, QW i Q , and It is a learnable weight matrix used to convert the input into an appropriate dimension. After the query domain name text vector passes through the multi-head attention channel, it is input into the LSTM layer to further enhance the model's memory ability for text sequence data.

6. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 5 is characterized in that: In step (3.5), the convolution layer is used to perform convolution operations on the input data through a set of learned filters to extract local features in the DNS traffic data, and use maximum pooling to divide the input data area into several rectangular areas and output the maximum value of each area to capture feature representations at a more abstract level and reduce the spatial dimension of the feature map, thereby improving the computational efficiency of the model. The Flatten layer converts the DNS multi-dimensional features output by the convolution layer into a one-dimensional vector. The process is simply expressed by the following formula: Flatten(D,H,W)=Vector(D·H·W) Among them, D, H, and W represent the sizes of different dimensions respectively. Then the Dropout layer randomly discards part of the neuron data in the network to prevent the model from overfitting. Finally, the fully connected layer is used to perform weighted summation on the output DNS feature vector of the dropout layer to achieve the mapping from feature extraction to decision output.

7. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 6 is characterized in that: In step (3.6), two features are concatenated and the feature vector is classified using a fully connected layer to achieve traffic detection. The output feature vectors of the multi-head attention LSTM and CNN come from two different dimensions. A classifier composed of a stacked fully connected layer structure consisting of three fully connected layers stacked in sequence and a softmax classification layer is used for the prediction task. The expression after feature fusion is as follows: merge=Concatenate(A,B) Among them, A and B are two sets of features, merge is the feature representation after splicing the features, the fully connected layer extracts and integrates the spliced ​​feature vectors, and converts them into lower dimensions while ensuring the retention of key information. The features are then handed over to the softmax function to map the feature vector to the [0,1] interval, and the output of the entire vector is normalized to obtain the final prediction result.

8. The method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics according to claim 3 is characterized in that: Step (4) specifically includes the following sub-steps: (4.1) The unidentified early malicious traffic of IoT is processed in the same way as (2.2) to obtain the features, and the DNS query feature set of the unidentified traffic is obtained; (4.2) The obtained DNS query feature set is input into the model for detection, and whether it is malicious IoT traffic is determined based on the output results, thereby detecting malicious IoT traffic in the early infection stage.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics as described in any one of claims 1 to 8 above is implemented.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, the method for early malicious traffic detection in the Internet of Things combining domain name and DNS traffic characteristics as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Malicious domain name detection method and device, equipment and storage medium

    CN114513355A

  • Trained model to detect malicious command and control traffic

    US11843624B1