DoH malicious tunnel traffic detection method and system based on feature fusion
The DoH malicious traffic detection method based on feature fusion and semi-supervised learning, combined with byte sequence and statistical features, solves the shortcomings of existing detection methods in accuracy and adaptability, and achieves efficient identification and stable detection of DoH malicious traffic.
Patent Information
- Application Number
- CN202510835203.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-05
AI Technical Summary
Existing DoH malicious tunnel traffic detection methods have shortcomings in detection accuracy, adaptability and deployment cost, and it is difficult to effectively identify malicious traffic generated by various DNS tunneling tools.
A DoH malicious traffic detection method based on feature fusion is adopted, which combines byte sequence features and statistical features, performs feature fusion through a multi-head attention mechanism, introduces a semi-supervised learning mechanism, and uses dynamic pseudo-label screening for training.
It improves the accuracy and generalization of DoH malicious traffic detection, enhances the model's ability to discriminate complex communication behaviors, and improves deployment adaptability and detection reliability in real network environments.
Smart Images

Figure CN120602179A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to cyberspace security, and specifically to a DoH malicious tunnel traffic detection method and system based on feature fusion. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of encrypted network communication technology, the DNS over HTTPS (DoH) protocol has been widely adopted by mainstream browsers and operating systems because it encapsulates traditional plaintext DNS queries within an HTTPS encrypted channel, significantly improving the privacy and anti-eavesdropping capabilities of user communications. However, the encryption and stealthiness of the DoH protocol also facilitate the construction of covert communication channels by network attackers. Attackers can use this to bypass traditional intrusion detection and content filtering mechanisms based on plaintext traffic analysis, carrying out malicious activities such as data theft and remote control, posing a significant threat to network security. Therefore, researching and implementing effective detection methods for DoH malicious tunneling traffic is of great practical significance and urgency.
[0004] Currently, there are various DoH tunnel detection methods, including rule-based methods, statistical feature-based machine learning methods, and deep learning algorithms based on automatic feature extraction. Rule-based methods rely on pre-set traffic characterization rules, resulting in high maintenance costs and difficulty adapting to new attacks, resulting in low detection accuracy. Statistical feature-based machine learning methods extract statistical features of traffic (such as packet length, flow duration, and inter-packet interval) for modeling and classification. While these statistical feature-based machine learning methods can extract traffic characteristics to a certain extent, they are highly dependent on expert experience and the single features (statistical features) used are easily circumvented by attackers, similarly suffering from low detection accuracy. Deep learning algorithms based on automatic feature extraction have automatic feature extraction capabilities, but their reliance on large amounts of high-quality annotated data makes them difficult to deploy in real-world network environments. Therefore, existing methods still have significant shortcomings in detection accuracy, adaptability, and deployment costs. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a DoH malicious tunnel traffic detection method and system based on feature fusion. This method integrates multimodal features and introduces a semi-supervised learning mechanism to effectively identify malicious DoH traffic generated by various DNS tunnel tools under a limited annotated data set, thereby improving detection accuracy and generalization ability.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: One or more embodiments provide a method for detecting DoH malicious traffic based on feature fusion, including the following steps: Preprocess the acquired DoH traffic data to be detected; The session data in the preprocessed DoH traffic data is segmented into a token sequence, and features are extracted using a byte sequence feature extractor to obtain a byte sequence feature vector. For the session data in the preprocessed DoH traffic data, statistical features are extracted and standardized, and features are extracted through the statistical feature extraction sub-network to obtain statistical feature vectors; The byte sequence feature vector and the statistical feature vector are fused based on the multi-head attention mechanism to obtain the classification results of DoH traffic; The byte sequence feature extractor and the statistical feature extraction sub-network extract features, and are trained using a semi-supervised learning framework with dynamic pseudo-label screening for unlabeled training samples.
[0007] One or more embodiments provide a DoH malicious traffic detection system based on feature fusion, including: A preprocessing module, configured to preprocess the acquired DoH traffic data to be detected; The byte sequence feature vector extraction module is configured to perform sequence segmentation processing on the session data in the preprocessed DoH traffic data to obtain a token sequence, and extract features based on the byte sequence feature extractor to obtain a byte sequence feature vector; A statistical feature vector extraction module is configured to extract and normalize statistical features from the session data in the preprocessed DoH traffic data, and extract features through a statistical feature extraction subnetwork to obtain a statistical feature vector; The fusion module is configured to fuse the byte sequence feature vector and the statistical feature vector based on the multi-head attention mechanism to obtain the DoH traffic classification result; The byte sequence feature extractor and the statistical feature extraction sub-network extract features, and are trained using a semi-supervised learning framework with dynamic pseudo-label screening for unlabeled training samples.
[0008] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps in the above-mentioned DoH malicious traffic detection method based on feature fusion are completed.
[0009] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps in the above-mentioned DoH malicious traffic detection method based on feature fusion are completed.
[0010] Compared with the prior art, the present invention has the following beneficial effects: The present invention improves the effectiveness and practicality of DoH malicious traffic detection in many aspects. On the one hand, it integrates the multi-dimensional information of byte sequences and statistical features, and improves the adequacy and discriminability of feature representation through a multi-head attention mechanism, effectively enhancing the model's ability to discriminate complex communication behaviors. On the other hand, it adopts a semi-supervised learning framework with dynamic pseudo-label screening, which does not rely on a large amount of manually labeled data. Through weak and strong enhancement strategies combined with a dynamic screening mechanism, the model can continuously self-optimize in a real network environment and improve deployment adaptability. In addition, through feature standardization and noise perturbation, the model's robustness to abnormal disturbances is improved, and the reliability and stability of detection are enhanced.
[0011] The advantages of the present invention and its additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0013] Figure 1 is an overall flow chart of the method of Example 1 of the present invention; Figure 2 Schematic diagram of the byte sequence feature extraction structure of Example 1 of the present invention; Figure 3 Schematic diagram of the statistical feature depth representation extraction structure of Example 1 of the present invention; Figure 4 is a flowchart of unsupervised training of unlabeled data according to embodiment 1 of the present invention; DETAILED DESCRIPTION The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0014] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0015] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof. It should be noted that, in the absence of conflict, the various embodiments of the present invention and the features in the embodiments can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.
[0016] Technical term explanation: DNS (Domain Name System): Translated as domain name system, it is one of the core services of the Internet. It converts human-readable domain names into machine-readable IP addresses, enabling users to access websites and services.
[0017] HTTPS (HyperText Transfer Protocol Secure): Secure Hypertext Transfer Protocol. HTTPS is a secure version of HTTP that adds TLS / SSL encryption to the HTTP protocol to ensure the confidentiality and integrity of data transmission between the client and the server. DoH (DNS over HTTPS): DNS protocol transmitted via HTTPS; DNSTT (DNS Tunnel Tool): A DNS-based tunneling tool. DNSTT is a tunneling communication tool that encapsulates arbitrary data in DNS queries, using specific encoding techniques to hide TCP / UDP data streams in DNS messages. GoDoH: Chinese is the DoH tunnel tool based on the Go language. GoDoH is an open source project that uses the DNS over HTTPS protocol channel to achieve hidden two-way communication.
[0018] Iodine: A DNS-based tunneling tool. Iodine is a classic DNS tunneling communication tool that can achieve network connectivity in restricted networks (such as networks that only allow DNS egress).
[0019] Example 1 In the technical solutions disclosed in one or more embodiments, Figures 1 to 4 As shown, a DoH malicious traffic detection method based on feature fusion includes the following steps: Step 1: Preprocess the acquired DoH traffic data to be detected; Step 2: Tokenize the session data in the preprocessed DoH traffic data to obtain a token sequence, and extract features based on the byte sequence feature extractor to obtain a byte sequence feature vector. Step 3: For the session data in the preprocessed DoH traffic data, extract statistical features and normalize them, and extract features through the statistical feature extraction sub-network to obtain statistical feature vectors; Step 4: The byte sequence feature vector and the statistical feature vector are fused based on the multi-head attention mechanism to obtain the classification result of DoH traffic; The byte sequence feature extractor and the statistical feature extraction sub-network extract features and are trained using a semi-supervised learning framework with dynamic pseudo-label screening for unlabeled training samples; This implementation improves the effectiveness and practicality of DoH malicious traffic detection in many aspects. On the one hand, it integrates the multi-dimensional information of byte sequences and statistical features, and improves the adequacy and discriminability of feature representation through a multi-head attention mechanism, effectively enhancing the model's ability to discriminate complex communication behaviors. On the other hand, it adopts a semi-supervised learning framework with dynamic pseudo-label screening, which does not rely on a large amount of manually labeled data. Through weak and strong enhancement strategies combined with a dynamic screening mechanism, the model can continuously self-optimize in a real network environment and improve deployment adaptability. In addition, through feature standardization and noise perturbation, the model's robustness to abnormal disturbances is improved, and the reliability and stability of detection are enhanced.
[0020] Furthermore, a semi-supervised learning framework with dynamic pseudo-label screening is used for training unlabeled training samples. Specifically, before the first round of training, the token sequence is randomly masked and the statistical features are weakly enhanced with Gaussian noise. Samples with a prediction probability greater than a set confidence threshold are screened out. The first round of prediction results are used as pseudo-labels, and the samples screened out in the first round of training are strongly enhanced before the second round of training. In this embodiment, a DoH malicious traffic detection framework is constructed by fusing byte-level and statistical-level features. First, the raw DoH session data is cleaned, de-redundanted, and normalized to facilitate subsequent modeling. A byte sequence feature extractor segments the session data into token sequences and extracts high-dimensional feature representations from these token sequences, capturing the semantic patterns hidden in encrypted communications. Simultaneously, a statistical feature extraction subnetwork constructs feature vectors based on traffic statistical properties (such as average packet length, flow duration, and standard deviation of inter-packet intervals) to reflect the overall behavioral characteristics of DoH traffic. Both networks fuse features using a multi-head attention mechanism, effectively integrating fine-grained and macro-grained information to improve classification accuracy. For unlabeled samples, a semi-supervised learning strategy based on dynamic pseudo-label screening is introduced to alleviate the reliance of supervised learning on labeled data. Weak augmentation is performed by masking the token sequence and adding Gaussian noise to the statistical features. After initial model training, samples with high prediction confidence are selected as pseudo-labels. Further training is performed using strong augmentation, gradually improving the model's generalization and discrimination capabilities for unlabeled samples.
[0021] In step 1, DoH traffic data comes from packet capture files (such as PCAP) and is constructed in units of complete communication sessions.
[0022] Byte sequence features extract the payload bytes of each data packet in a session and concatenate them in chronological order to capture fine-grained contextual information in encrypted traffic. Statistical features quantify the overall behavior patterns of a session from a global perspective, such as session duration, mean and variance of data packet length, directional distribution, and time interval distribution.
[0023] In step 1, DoH traffic data is preprocessed, including splitting traffic into quintuples and filtering sessions with too few packets. Specifically, DoH traffic data is pre-processed, including traffic splitting and traffic filtering: Step 11: Traffic splitting: Use the SplitCap tool to divide the DoH traffic data session into five-tuple pairs to obtain multiple DoH sessions with clear structures. The five-tuple includes source address, destination address, source port, destination port, and protocol type; Step 12: Traffic filtering: Set the minimum packet length min_window_size and delete sessions that are shorter than the minimum packet length. Given the varying lengths of sessions, to eliminate anomalous samples that may be missing TLS handshake information, we filter out sessions with fewer than min_window_size packets. Furthermore, the byte sequence feature extractor requires a fixed-length input. Short sessions require extensive zero-padding, which can introduce noise and affect feature extraction. Given that the TLS handshake typically completes within the first six packets, in this example, min_window_size is set to 6.
[0024] The preprocessing method of this embodiment can generate a traffic representation suitable for a subsequent detection model while eliminating invalid data that may affect the model performance.
[0025] In step 2, the message byte sequence of the session data packet in the DoH traffic data is segmented into a sequence, i.e., tokenized. This includes defining the BURST_Session structure, using a combined Bigram and Unigram strategy for word segmentation, and constructing an embedding representation. Step 21: Build a BURST_Session structure, extract the payloads of the first n packets in each session data packet in the DoH traffic data, arrange them in chronological order, and mark the direction of the payload of each packet to obtain the BURST_Session structure data; Specifically, the BURST_Session structure includes the original byte content of the data packet and the direction tag (direction). Each data packet in the BURST_Session is accompanied by a direction tag. The structure example is as follows: BURST_Session = [ {"packet": pkt1_bytes, "direction": "request"}, {"packet": pkt2_bytes, "direction": "response"}, ... ].
[0026] Among them, packet represents the raw byte content; pkt1_bytes and pkt2_bytes are variable names used to represent the raw byte content of the data packet; direction represents the direction variable used to indicate whether this data packet is a request or a response; The BURST_Session constructed in this embodiment is different from the traditional BURST construction method. It is constructed based on sessions, retaining the first n data packets in each session and arranging these data packets in chronological order to form a BURST_Session, which represents the initial interaction part in each session, including requests or responses; this BURST_Session structure can not only effectively capture the semantic information in network communication, but also extract the interaction features between requests and responses, providing efficient and structured input for subsequent feature extractors.
[0027] In this embodiment, it is proposed to construct a BURST_Session structure based on sessions and introduce a data packet direction attribute to enhance the modeling capability of malicious DNS over HTTPS traffic behavior.
[0028] Step 22: Use the Bigram and Unigram joint strategy to segment the BURST_Session structure data and obtain the Token sequence as follows: Step 221: Encode the original byte content in the BURST_Session structure data in hexadecimal format, padded with the set length, and converted into a character sequence to obtain a hexadecimal string. To convert the BURST_Session structure into a token sequence recognizable by the post-feature extractor, this embodiment performs hexadecimal encoding on the byte contents of each packet in the BURST_Session. During the encoding process, a maximum of 128 bytes are extracted from each packet, and any missing bytes are padded with zeros. The resulting string is then converted into a character sequence (e.g., FF00AABB). The direction field is not encoded and is retained as metadata.
[0029] Step 222: The hexadecimal string output from step 221 is encoded using a fusion of the Bigram and Unigram models, and a preliminary structural processing is performed on the Bigram to obtain a Tokenized Segment. The Unigram model selects subsequences with high learning frequency and large information content based on statistical rules as the token sequence after word segmentation, i.e., the word segmentation result; Tokenized Segment is the token segment after Bigram processing; Specifically, by performing preliminary structural processing using Bigram, the hexadecimal string can be divided into groups of two characters to reduce noise; Due to the inconsistency in the length of each field in DoH protocol packets, the entire protocol exhibits a variable-length and non-fixed byte pattern. Traditional fixed-granularity word segmentation methods (such as Bigram) struggle to accurately delineate semantic boundaries and can easily overlook important information between fields. Subword encoding methods used in natural language processing (such as BPE or WordPiece) are either too coarse-grained when processing hexadecimal data, failing to capture details, or too fine-grained, resulting in information fragmentation. Such encoding methods are also unsuitable for traffic semantic modeling. To address this issue, the word segmentation method described in this embodiment draws on the principles of Chinese word segmentation to propose an encoding method that integrates Bigram and Unigram models. Bigram is used to initially structure hexadecimal traffic, thereby reducing single-byte noise; while the Unigram model, based on statistical laws, automatically learns high-frequency, high-information subsequences as token units. Dynamically adjusting the granularity of the Unigram model preserves key information and improves the ability to express DoH traffic semantics. This method can handle variable-length, unstructured traffic segments and supports presetting the vocabulary size to control resource consumption during training.
[0030] The Unigram model selects subsequences with high learning frequency and large information content as token sequences based on statistical laws. High frequency and large information content can be achieved by sorting and selecting a set proportion of sequences, such as selecting the top 20% token sequences with high frequency and large information content. The Unigram model is obtained through training, and the initial vocabulary of the training dataset includes all tags and high-frequency substrings output by the existing pre-tagged device.
[0031] Unigram is a subword segmentation model supported by SentencePiece (SPM) and can be used to build subword vocabulary; During the training process, each time a token is deleted, the principle of minimizing the increase in loss is followed. The loss function is calculated as shown in formula (1): ; Where, Represents a word in a given data set, which in this embodiment refers to a hexadecimal encoded string of a data packet; Indicates the i-th word ( ; ; …; ); for each The set of possible word segmentation methods is represented as .
[0032] After the data set is determined, the word segmentation method set for each word It is also determined that each word segmentation method in the set has a certain probability of being selected If some tokens are deleted from the vocabulary, then for some words, the set The number of word segmentation methods will be reduced, thereby reducing the summation term in the loss function log(), thereby increasing the overall loss. Optionally, the Unigram model algorithm selects 10% to 20% of tokens from the vocabulary that contribute the least to the incremental loss in each iteration and deletes them, that is, the loss function value changes the least after deletion.
[0033] Step 223: Construct a structure: bind the token sequence after segmentation to the corresponding direction information (Direction) to obtain a new structure BURST_Session_tokens; Specifically, each data packet in the new BURST_Session_tokens structure is represented in dictionary form, containing the fields "tokens" and "direction", which correspond to semantic features and behavioral markers respectively, so that subsequent models can distinguish the contextual semantics of requests / responses.
[0034] The token sequence of each data packet is bound to the direction, forming a combination of "semantics + behavior". The binding process is as follows: BURST_Session_tokens=[{"tokens":tokenized_data[i],"direction":BURST_Session[i]["direction"]} for i in range(len(tokenized_data))]; Step 224 : Split the data in the newly obtained structure BURST_Session_tokens into two parts: a request data sequence and a response data sequence.
[0035] Step 225: Add a token to the segmented data sequence to obtain a token sequence. Specifically, during the tokenization process, special tags are added for further processing: [CLS]: indicates the beginning of the sequence and is used to capture the characteristics of the entire sequence.
[0036] [SEP]: The boundary between sub-BURSTs used to divide requests and responses.
[0037] [PAD]: Used to pad the sequence to a fixed length to ensure the consistency of the input data.
[0038] The sequence structure of the Token sequence output in step 225 can be exemplified as follows: [CLS] + REQ_BURST + [SEP] + RES_BURST + [PAD]... (to a fixed length); Step 23: Construct an embedding representation: Sum the obtained token embedding, position embedding, and segment embedding vectors of the token sequence element by element to generate the final token sequence embedding representation.
[0039] Token Embedding: Specifically, each token sequence output in step 22 is mapped into a fixed-dimensional vector through a lookup table to represent its semantic information. Position Embedding: Specifically, it encodes the position information of each token in the sequence. Since the semantics of network traffic depends heavily on the order of packets, position embedding is introduced to encode the position information of tokens in the sequence, thereby preserving the sequential characteristics and enhancing the model's ability to perceive sequential patterns. Segment Embedding: Specifically, the request token (REQ_BURST) is assigned to segment A, and the response token (RES_BURST) is assigned to segment B.
[0040] This embodiment uses a combined Bigram and Unigram word segmentation strategy to process byte sequences, combined with a multi-layer Transformer model to implement byte-level semantic modeling, improving the ability to understand the contextual features of encrypted traffic.
[0041] Furthermore, in step 2, the Transformer network is used as a byte feature extractor, and the obtained token sequence embedding representation is used as input to extract sequence features. The Transformer-based byte sequence feature extractor can be expressed as: Transformer Layer; In step 3, the statistical features of the session data are extracted using the DoHLyzer tool and then normalized. Statistical features are extracted using the DoHLyzer tool, and fields that are strongly related to the network environment are removed, retaining only 28 common behavioral statistical features, including the number of traffic bytes sent, the traffic byte rate sent, the number of traffic bytes received, the traffic byte rate received, the average packet length, the median of the packet length, the mode of the packet length, the variance of the packet length, the standard deviation of the packet length, the coefficient of variation of the packet length, the skewness relative to the median of the packet length, the skewness relative to the mode of the packet length, the average packet time interval, the median of the packet time interval, the mode of the packet time interval, the variance of the packet time interval, the standard deviation of the packet time interval, the coefficient of variation of the packet time interval, the skewness relative to the median of the packet time, the skewness relative to the mode of the packet time, the average request / response time difference, the median request / response time difference, the mode request / response time difference, the request / response time difference variance, the request / response time difference standard deviation, the coefficient of variation of the request / response time difference, the skewness of the request / response time difference relative to the median, and the skewness of the request / response time difference relative to the mode.
[0042] The raw output contains multiple environment-related fields (such as source / destination IP addresses, ports, start timestamps, and duration). These fields are significantly dependent on the network environment and can cause the model to overfit to specific environments, reducing generalization capabilities. To improve the model's robustness in unknown environments, this example removes these sensitive fields and retains only 28 statistical features that are closely related to traffic behavior and have good generalization capabilities as input.
[0043] Furthermore, since these 28 statistical features have different dimensions and numerical ranges (such as number of bytes, number of packets, time intervals, etc.), directly inputting them would interfere with the model's learning of feature importance. Therefore, the features are normalized before being input into the statistical feature extraction subnetwork.
[0044] The original statistical features of the i-th session are: ; After standardization, we get: ; in, and are the mean and standard deviation of the kth feature in the training set, respectively.
[0045] To further explore the interactions and high-level expressive power of statistical features, this example designs a statistical feature extraction subnetwork consisting of a sequentially connected one-dimensional convolutional neural network (1D-CNN) and a channel-spatial attention module (CABM) to extract deep representations of statistical features. This architecture models the relationships between features at both the local pattern and global dependency levels, thereby enhancing the discriminative power of statistical features.
[0046] In step 3, the process of extracting features based on the statistical feature extraction sub-network includes the following steps: Step 31, 1D-CNN local feature extraction: the standardized statistical features Input to the one-dimensional convolution layer to perform convolution and batch normalization; In this embodiment, 32 convolution kernels are used to extract local combination features:
[0047] Output feature map Represents the response of 32 channels (each channel corresponds to a convolution kernel) on 28 feature dimensions.
[0048] Step 32: The obtained local combined features are weighted based on the attention operation of the channel-spatial attention module (CABM); To highlight the highly discriminative channel features, a channel attention mechanism is introduced. After performing global average pooling and maximum pooling on each channel, the attention weight is generated through a shared MLP mapping: ; Among them, the weight , used to weight the importance of each channel, the output is: ⋅X; The spatial attention mechanism introduces an attention mechanism in the statistical feature dimension (i.e., spatial dimension) to identify the location of key feature dimensions. After performing average pooling and maximum pooling in the channel dimension, the spatial attention weights are generated through MLP: ; in, is the spatial attention weight, and the final output is the potential representation of the statistical features of the fused attention: ⋅ ; In this embodiment, through the CABM module, the model can simultaneously perceive the collaborative relationship between channels and the key areas of feature dimensions, effectively improving the representation ability of statistical features. It will be used for subsequent multimodal fusion modeling.
[0049] In step 4, the byte sequence feature vector obtained in step 2 is fused with the statistical feature vector obtained in step 3, which includes the following steps: The two types of features characterize traffic from different perspectives: byte sequence features can reflect the structure and protocol behavior of the original data packet, while statistical features characterize the overall statistical laws of the session. The two are highly complementary.
[0050] Step 41: uniformly map the two types of modal features to the same dimensional space; The same dimensional space can achieve effective alignment and fusion. The byte sequence feature vector is represented as: , the statistical feature depth is represented as , where L is the length of the byte sequence and d is the unified feature dimension.
[0051] Step 42: Perform multi-head attention operation on the mapped features to convert the byte sequence feature vector As a query, statistical feature vector As key (Key) and value (Value), the byte sequence feature Through the attention mechanism fusion, the fused features are obtained; This embodiment uses a multi-head attention operation to guide the model to extract contextual information related to byte features from statistical features. The calculation process is as follows: ; ; ; in, , , , , h represents the number of attention heads; the final fusion representation , as the input of the classifier, the classification result is obtained by the classifier, thereby identifying whether the DoH traffic data to be detected is DoH malicious traffic; In this example, 1D-CNN and CABM are introduced to extract statistical features, combined with a multi-head attention mechanism to achieve cross-modal feature fusion, forming a more discriminative joint representation. Multi-Head Attention (MHA) is also introduced to model fine-grained associations between different modalities. This multi-head attention mechanism enables parallel modeling of fine-grained cross-modal associations in multiple subspaces, further enhancing the expressiveness and discriminative power of the fused features.
[0052] A further technical solution is to train the byte sequence feature extractor and the statistical feature extraction sub-network, including the following steps: Step S1: Obtain DoH traffic data for preprocessing, construct an original data set, and divide it into training set 1 and training set 2; First, we selected the CIRA-CIC-DoHBrw-2020 public dataset as the basic data source and designed various network configurations, including different client locations, browser types, and DoH servers. We also constructed the EDNSv6 dataset in an IPv6 network environment. The two datasets were combined to form a complete training set.
[0053] This example selects two datasets, CIRA-CIC-DoHBrw-2020 and the collected EDNSv6, as the training data sources, covering benign DoH traffic and DoH malicious tunnel traffic generated by a variety of typical tunnel tools (including dns2tcp, dnscat2, Iodine, DNSTT, GoDoH and DNSExfiltrator), which are highly representative and diverse.
[0054] The data in the training set is preprocessed, including traffic splitting and filtering, and divided into two subsets: training set 1 is used for pre-training, and training set 2 is used for backbone training. Specifically: 80% of the total data in the original dataset are labeled and divided into two equal parts. Half of the data constitutes training set 1 (unlabeled dataset), which is used for pre-training of the byte sequence feature extractor; the other half is combined with 20% of the data with retained labels to form training set 2 (semi-supervised training set), which contains both unlabeled samples and labeled samples for backbone training.
[0055] Step S2: Tokenize the message byte sequence of the session data in the DoH traffic data to obtain a token sequence; Tokenization includes defining the BURST_Session structure, using a combined Bigram and Unigram strategy for word segmentation, and constructing an embedding representation. The implementation of this step is the same as step 2 and will not be repeated here. Step S3: Use the DoHLyzer tool to extract the statistical features of the session and perform normalization. The implementation process of this step is the same as that of step 3 and will not be repeated here. Step S4: For the Transformer-based byte sequence feature extractor, two self-supervised tasks, mask prediction and same-session prediction, are set, and pre-training is performed using training set 1 to obtain a pre-trained byte sequence feature extractor; To enhance the feature extractor's ability to model network communication semantics and behavior patterns, this embodiment designs two self-supervised tasks, focusing on byte-level context modeling and session-level interaction modeling. The self-supervised tasks include: (1) Masked BURST_Session Prediction: Perform random masking operations on some tokens to predict the original content; Specifically, a first set proportion of tokens in the input token sequence is selected and randomly masked, such as 15%; the masked content is predicted using contextual information, thereby learning the semantic dependencies between byte sequences.
[0056] Specifically, the random masking strategy can be as follows: in the selected token sequence, the second set ratio of segmented words is replaced with a mask [MASK]; the third set ratio of segmented words is replaced with random segmented words (Token), and the remaining segmented words remain unchanged; for example, the second set ratio can be set to 80%, the third set ratio can be set to 10%, and the remaining 10% remains unchanged; (2) Same-Session BURST_Session Prediction: Predict whether two sub-BURST_Session structure data fragments belong to the same session; Specifically, the same-session prediction task uses bidirectional burst-session pair prediction. This involves sampling a request sub-burst-session (REQ_BURST) and a response sub-burst-session (RES_BURST) from the same session to construct positive pairs. Sub-burst-session pairs from different sessions are then constructed as negative samples. A byte sequence feature extractor is trained based on these positive and negative pairs to determine whether they belong to the same session. This task simulates a real request-response interaction process, helping the model learn cross-directional information associations and behavioral semantics.
[0057] To obtain a byte sequence feature extractor with good generalization capabilities, we performed multiple rounds of iterative training on the aforementioned self-supervised task using token sequences from training set 1. During training, hyperparameters such as the learning rate and batch size were dynamically adjusted based on the model's progress until the model converged or reached the preset maximum number of rounds, resulting in a pre-trained byte sequence feature extractor.
[0058] Step S5: weakly enhance the samples in the second training set. The weak enhancement includes randomly masking the token sequence obtained in step S2 and adding noise to the statistical features obtained in step S3. The weakly enhanced samples are respectively input into the pre-trained byte sequence feature extractor and the statistical feature extraction sub-network for the first round of backbone training. The first round of trunk training includes: Step S5.1, supervised training: The weakly enhanced data of the labeled samples is input into the byte sequence feature extractor and the statistical feature extraction sub-network to extract the two types of modal features. The multi-head attention mechanism is used to fuse them, and the classifier outputs the prediction results, and the supervised loss is calculated; Step S5.2, unsupervised training: For unlabeled samples, use the model to generate pseudo labels for weakly enhanced samples, and combine the adaptive confidence threshold to screen high-confidence samples. After screening out samples with a predicted probability greater than the set confidence threshold, perform strong enhancement and then conduct a second round of training. Input the model and calculate the unsupervised loss.
[0059] The final backbone training loss is the sum of supervised loss and unsupervised loss; Step S6: Based on the obtained backbone training loss value, adjust the parameters of the model, and iteratively train to obtain the trained byte sequence feature extractor and statistical feature extraction sub-network.
[0060] In step S5, a weak enhancement method is used for the token sequence. Specifically, a first set ratio of tokens in the input token sequence is selected and randomly masked. For example, the first set ratio is 15%. The random masking strategy can be as follows: in the selected token sequence, the second set ratio of words is replaced with a mask [MASK]; the third set ratio of words is replaced with random words (Token), and the remaining words remain unchanged; for example, the second set ratio can be 80%, the third set ratio can be set to 10%, and the remaining 10% remains unchanged; In step S5, a weak enhancement method for statistical features is used to add a small noise that follows a normal distribution to each eigenvalue of the standardized statistical feature vector to simulate the natural fluctuations in the network and enhance the model's tolerance to input disturbances.
[0061] In step S5.1, the supervised training process is as follows: Input the token sequence in the weakly enhanced labeled sample processed in step S5 into the byte sequence feature extractor pre-trained in step S4 to obtain the byte sequence feature representation; Input the statistical features of the weakly enhanced labeled samples processed in step S5 into the statistical feature extraction sub-network (1D-CNN + CABM) structure to extract the statistical feature potential representation; Compute the supervised loss: ; in, is a labeled sample, represents the weakly enhanced labeled samples, is the one-hot vector of the true label, represents the model prediction operation, is the cross entropy loss.
[0062] In step S5.2, the unsupervised training process is as follows: Step S5.21: Dynamically generate pseudo labels for unlabeled samples and select samples for the second round of training. This involves adaptively adjusting the confidence threshold of the pseudo labels based on the results of the first round of training, selecting samples with a prediction probability greater than the set confidence threshold, and using the first round of prediction results as pseudo labels. This process includes the following: (21-1) Inputting the token sequence in the weakly enhanced unlabeled sample processed in step S5 into the byte sequence feature extractor pre-trained in step S4 to obtain a byte sequence feature representation; (21-2) Inputting the statistical features of the weakly enhanced unlabeled samples processed in step S5 into the statistical feature extraction subnetwork (1D-CNN + CABM) to extract the statistical feature potential representation; (21-3) Use the above step 4 to perform multi-head attention mechanism fusion, and send the fused features into the classifier to obtain the predicted probability distribution: ; in, represents the predicted category label, is an unlabeled sample, , Represents the model prediction operation; (21-4) For each weakly enhanced unlabeled sample , independently predict the byte sequence features and statistical features, and obtain two classification prediction results:
[0063]
[0064] (21-5) Calculate the cosine similarity between the two classification prediction results as the consistency score: ) (21-6) Adaptively adjust the confidence threshold of the pseudo-label according to the consistency score;
[0065] in, and is a hyperparameter.
[0066] The dynamic generation method of adaptively adjusting the confidence threshold of pseudo labels in this embodiment can be adaptively adjusted according to classification consistency. If the consistency is low, the threshold is increased to avoid low-confidence pseudo labels; if the consistency is high, the threshold is lowered to adopt more high-quality pseudo labels.
[0067] (21-7) Select samples greater than the confidence threshold as sample data for the second round of training, and use the results of the first round of prediction as pseudo labels; like , then set the pseudo label to , that is, assign pseudo labels to samples that meet the threshold conditions; and delete samples that do not meet the threshold conditions and conduct a second round of training; This embodiment designs a dynamically adjusted pseudo-label confidence mechanism, which can improve the learning efficiency and detection accuracy of the model in a weakly labeled environment.
[0068] Step S5.22: Perform a second round of training on the samples with generated pseudo labels; Specifically, a strong enhancement operation is performed on the token sequence of samples that meet the threshold condition in (21-7); Strong enhancement operation, specifically: delete continuous segments of a set length from the token sequence and fill them with set characters; select features with a set proportion for statistical features and set a unified set value for the selected features; Specifically, for each token sequence, a "span deletion" operation is performed: one to two consecutive spans, each one to three tokens long, are randomly selected and deleted to simulate unexpected anomalies or traffic loss in network communications. To avoid excessive information loss, the total deletion ratio is limited to no more than 30% of the original sequence length. After the deletion operation, to ensure consistency in the model input length, all sequences are padded with [PAD] tokens to restore them to the preset fixed input length.
[0069] Specifically, the statistical features of samples that meet the threshold conditions are enhanced. Specifically, 20% to 30% of the features are randomly selected from the 28 features and their values are set to 0 or the training set mean to simulate missing information or measurement anomalies. In this embodiment, the uniform setting value is 0 or the training set mean; (22-2) Input the strongly enhanced Token sequence and statistical features into the byte sequence feature extractor and the statistical feature extraction sub-network respectively to extract two types of features; (22-2) Use the multi-head attention mechanism to fuse features and feed them into the classifier. Use the pseudo-label obtained in step S5.21 as the supervision signal to calculate the unsupervised loss: ); in, is the adaptive confidence threshold, is the indicator function.
[0070] The final backbone training loss is the sum of supervised loss and unsupervised loss, and the formula is: ; in, An important hyperparameter for balancing supervised and unsupervised losses.
[0071] When the loss function After several rounds of training, the training was terminated when the performance stabilized. After training, a SemiSBF-Net model with multimodal feature perception capabilities was obtained. The SemiSBF-Net model includes a byte sequence feature extractor and a statistical feature extraction subnetwork. It can efficiently detect DoH malicious tunneling traffic. It can accurately identify and distinguish different types of DoH traffic, including but not limited to legitimate DNS over HTTPS traffic, as well as malicious tunneling traffic generated by tunneling tools such as DNSTT, GoDoH, and Iodine.
[0072] Example 2 Based on Example 1, this embodiment provides a DoH malicious traffic detection system based on feature fusion, including: A preprocessing module, configured to preprocess the acquired DoH traffic data to be detected; The byte sequence feature vector extraction module is configured to perform sequence segmentation processing on the session data in the preprocessed DoH traffic data to obtain a token sequence, and extract features based on the byte sequence feature extractor to obtain a byte sequence feature vector; A statistical feature vector extraction module is configured to extract and normalize statistical features from the session data in the preprocessed DoH traffic data, and extract features through a statistical feature extraction subnetwork to obtain a statistical feature vector; The fusion module is configured to fuse the byte sequence feature vector and the statistical feature vector based on the multi-head attention mechanism to obtain the DoH traffic classification result; The byte sequence feature extractor and the statistical feature extraction sub-network extract features, and are trained using a semi-supervised learning framework with dynamic pseudo-label screening for unlabeled training samples.
[0073] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.
[0074] Example 3 This embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and running on the processor. When the computer instructions are run by the processor, the steps in the DoH malicious traffic detection method based on feature fusion in Example 1 are completed.
[0075] Example 4 This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps in the DoH malicious traffic detection method based on feature fusion in Example 1 are completed.
[0076] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0077] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A DoH malicious traffic detection method based on feature fusion, characterized by: The steps include: Preprocess the acquired DoH traffic data to be detected; The session data in the preprocessed DoH traffic data is segmented into a token sequence, and features are extracted using a byte sequence feature extractor to obtain a byte sequence feature vector. For the session data in the preprocessed DoH traffic data, statistical features are extracted and normalized, and features are extracted through the statistical feature extraction sub-network to obtain statistical feature vectors; The byte sequence feature vector and the statistical feature vector are fused based on the multi-head attention mechanism to obtain the classification results of DoH traffic; The byte sequence feature extractor and the statistical feature extraction sub-network extract features, and are trained using a semi-supervised learning framework with dynamic pseudo-label screening for unlabeled training samples.
2. The DoH malicious traffic detection method based on feature fusion according to claim 1 is characterized in that: A semi-supervised learning framework with dynamic pseudo-label screening is used for training unlabeled training samples. Specifically, before the first round of training, the token sequence is randomly masked and the statistical features are weakly enhanced with Gaussian noise. Based on the results of the first round of training, the confidence threshold of the pseudo-label is adaptively adjusted to screen out samples with a prediction probability greater than the set confidence threshold. The samples screened out in the first round of training are strongly enhanced before the second round of training.
3. The DoH malicious traffic detection method based on feature fusion according to claim 2 is characterized in that: The random masking method is as follows: in the selected token sequence, the second set ratio of segmented words is replaced with a mask; the third set ratio of segmented words is replaced with random segmented words, and the remaining segmented words remain unchanged; Alternatively, a strong enhancement operation is performed, specifically: for the token sequence, continuous segments of a set length are deleted and padded with set characters; for the statistical features, features with a set proportion are selected, and a uniform set value is set for the selected features.
4. The DoH malicious traffic detection method based on feature fusion according to claim 2, characterized in that: The process of adaptively adjusting the confidence threshold of the pseudo-label based on the first round of training results, screening out samples with a prediction probability greater than the set confidence threshold, and using the first round of prediction results as pseudo-labels includes the following: Input the token sequence in the weakly enhanced unlabeled sample into the pre-trained byte sequence feature extractor to obtain the byte sequence feature representation; The statistical features of the weakly enhanced unlabeled samples are input into the statistical feature extraction sub-network to extract the potential representation of the statistical features; The obtained byte sequence feature representation and statistical feature potential representation are fused through a multi-head attention mechanism, and the fused features are fed into the classifier to obtain the predicted probability distribution: For each weakly enhanced unlabeled sample , independently predict the byte sequence features and statistical features, and obtain two classification prediction results; Compute the cosine similarity between two classification predictions as a consistency score: Adaptively adjust the confidence threshold of the pseudo-label based on the consistency score; Samples with a confidence level greater than the threshold are selected as sample data for the second round of training, and the results of the first round of predictions are used as pseudo labels.
5. The DoH malicious traffic detection method based on feature fusion according to claim 1 is characterized in that: Sequence segmentation is performed on the message byte sequence of the session data packet in the DoH traffic data, including: Construct a BURST_Session structure, extract the payloads of the first n packets in each session data packet in the DoH traffic data, arrange them in chronological order, and mark the direction of the payload of each packet to obtain the BURST_Session structure data; The combined strategy of Bigram and Unigram is used to perform word segmentation modeling on the BURST_Session structure data to obtain the Token sequence.
6. The DoH malicious traffic detection method based on feature fusion according to claim 5 is characterized in that: The combined strategy of Bigram and Unigram is used to segment the BURST_Session structure data and obtain the Token sequence. The process is as follows: The original byte content in the BURST_Session structure data is hexadecimal-encoded and padded to the set length, and then converted into a character sequence to obtain a hexadecimal string; The hexadecimal string is encoded using a fusion of Bigram and Unigram models, and initially structured using Bigram. The Unigram model selects subsequences with high learning frequency and large information content based on statistical rules as the token sequence after word segmentation to obtain the word segmentation result; Bind the token sequence after word segmentation with the corresponding direction information to obtain a new structure; The data in the obtained new structure is divided into two parts: request data sequence and response data sequence; Add tags to the segmented data sequence to obtain a Token sequence.
7. The DoH malicious traffic detection method based on feature fusion according to claim 1, characterized in that: The byte sequence feature vector is integrated with the statistical feature vector, including the following steps: Map the two types of modal features uniformly to the same dimensional space; Perform a multi-head attention operation on the mapped features, using the byte sequence feature vector as the query and the statistical feature vector as the key and value. The byte sequence features and the statistical feature vector are fused through the attention mechanism to obtain the fused features. Alternatively, the process of training the byte sequence feature extractor and the statistical feature extraction sub-network includes the following steps: Step S1: Obtain DoH traffic data for preprocessing, construct an original data set, and divide it into training set 1 and training set 2; Step S2: Tokenize the message byte sequence of the session data in the DoH traffic data to obtain a token sequence; Step S3: Use the DoHLyzer tool to extract the statistical features of the session and perform normalization processing; Step S4: For the Transformer-based byte sequence feature extractor, two self-supervised tasks, mask prediction and same-session prediction, are set, and pre-training is performed using training set 1 to obtain a pre-trained byte sequence feature extractor; Step S5: weakly enhance the samples in the second training set. The weak enhancement includes randomly masking the token sequence obtained in step S2 and adding noise to the statistical features obtained in step S3. The weakly enhanced samples are respectively input into the pre-trained byte sequence feature extractor and the statistical feature extraction sub-network for the first round of backbone training. The first round of trunk training includes: Supervised training: The weakly enhanced data of the labeled samples is input into the byte sequence feature extractor and the statistical feature extraction sub-network. The two types of modal features are extracted and fused, and the prediction results are output. The supervised loss is calculated. Unsupervised training: For unlabeled samples, weakly enhanced samples are used to generate pseudo labels. High-confidence samples are screened using an adaptive confidence threshold. Samples with predicted probabilities greater than the set confidence threshold are screened for strong enhancement and then a second round of training is performed. The unsupervised loss is then calculated. Based on the obtained loss value, the parameters of the model are adjusted, and the training is iteratively performed to obtain the trained byte sequence feature extractor and statistical feature extraction sub-network.
8. The DoH malicious traffic detection system based on feature fusion is characterized by: include: A preprocessing module, configured to preprocess the acquired DoH traffic data to be detected; The byte sequence feature vector extraction module is configured to perform sequence segmentation processing on the session data in the preprocessed DoH traffic data to obtain a token sequence, and extract features based on the byte sequence feature extractor to obtain a byte sequence feature vector; A statistical feature vector extraction module is configured to extract and normalize statistical features from the session data in the preprocessed DoH traffic data, and extract features through a statistical feature extraction subnetwork to obtain a statistical feature vector; The fusion module is configured to fuse the byte sequence feature vector and the statistical feature vector based on the multi-head attention mechanism to obtain the DoH traffic classification result; The byte sequence feature extractor and the statistical feature extraction sub-network extract features, and are trained using a semi-supervised learning framework with dynamic pseudo-label screening for unlabeled training samples.
9. An electronic device, characterized in that: The invention comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the DoH malicious traffic detection method based on feature fusion according to any one of claims 1 to 7 are completed.
10. A computer-readable storage medium, characterized in that Used to store computer instructions, which, when executed by a processor, complete the steps of the DoH malicious traffic detection method based on feature fusion as described in any one of claims 1-7.
Citation Information
Cited By
Malicious DoH tunnel traffic detection method and device, and storage medium
CN121125349A
Malicious doh tunnel traffic detection method, device and storage medium
CN121125349B
DoH tunnel detection method based on feature fusion and large language model
CN121864426A
Coating color formula prediction method, system and equipment, medium and product
CN121963943A