Malicious encrypted traffic detection method and system based on multi-modal feature fusion
By adopting a deep convolutional neural network model with multimodal feature fusion in encrypted traffic detection, the problem of difficulty in identifying malicious encrypted traffic in the existing technology is solved, and higher detection accuracy and network security are achieved.
Patent Information
- Application Number
- CN202510062073.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
In the face of encrypted traffic, it is difficult for the existing technology to effectively identify malicious behaviors, which makes it difficult for malicious encrypted traffic to be discovered and blocked in time, posing a huge threat to network security.
The malicious encrypted traffic detection method based on multimodal feature fusion is adopted. By constructing a deep convolutional neural network model, multi-head attention mechanism and multimodal factor analysis algorithm are used to weighted fusion and classify the multimodal features of encrypted traffic to realize the detection of malicious encrypted traffic.
Improve the accuracy and reliability of encrypted traffic detection, can more effectively identify and block malicious encrypted traffic, and enhance network security.
Smart Images

Figure CN119966688A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a malicious encrypted traffic detection method and system based on multi-modal feature fusion. Background Art
[0002] With the rapid development of network technology, encryption technology is increasingly used in network communications. More and more network applications use encryption protocols to ensure data security and privacy, such as the widespread use of HTTPS protocol in web browsing, as well as various encrypted instant messaging software and cloud storage services. This has led to an increasing proportion of encrypted traffic in network traffic. However, while encryption technology protects the security of legitimate user data, it also provides cover for malicious attackers, making it difficult for traditional traffic detection methods based on ports, protocol features, etc. to effectively identify malicious behavior in encrypted traffic. For example, hackers can use encrypted channels to transmit malware, steal sensitive information, or launch network attacks, and existing detection methods are often unable to penetrate the encryption layer to obtain effective information when facing encrypted traffic, making it difficult to detect and block malicious encrypted traffic in a timely manner, posing a huge threat to network security.
[0003] The prior art has the following technical problems:
[0004] First, the preprocessing operation is simple, resulting in a lot of data noise information and the loss of malicious encrypted traffic;
[0005] Second, the processed information is not effectively integrated, some features are lost, and the final detection accuracy is affected. Summary of the invention
[0006] In order to solve the deficiencies of the prior art, the present invention provides a malicious encrypted traffic detection method and system based on multimodal feature fusion;
[0007] On the one hand, a malicious encrypted traffic detection method based on multimodal feature fusion is provided, including:
[0008] Constructing a data set, wherein the data set is encrypted traffic of known malicious or non-malicious detection results;
[0009] Constructing a deep convolutional neural network model, using a training set to train the deep convolutional neural network model to obtain a trained model; the deep convolutional neural network model includes: a convolutional layer, a pooling layer, a feature fusion layer and a fully connected layer connected in sequence, the feature fusion layer first calculates the initial weight of the feature through a multi-head attention mechanism, and then uses a multimodal factor analysis algorithm to adjust the initial weight to obtain a fine-tuning weight, and finally uses the fine-tuning weight to perform weighted fusion on the features to obtain a fused feature, and uses a fully connected layer to classify the fused feature to obtain a classification result;
[0010] The encrypted traffic to be predicted is obtained, the encrypted traffic to be predicted is processed, and the processed encrypted traffic to be predicted is input into the trained model to obtain the detection result of the encrypted traffic.
[0011] On the other hand, a malicious encrypted traffic detection system based on multimodal feature fusion is provided, including:
[0012] A construction module is configured to: construct a data set, wherein the data set is encrypted traffic of known malicious or non-malicious detection results;
[0013] The training module is configured to: construct a deep convolutional neural network model, use a training set to train the deep convolutional neural network model, and obtain a trained model; the deep convolutional neural network model includes: a convolution layer, a pooling layer, a feature fusion layer, and a fully connected layer connected in sequence, the feature fusion layer first calculates the initial weight of the feature through a multi-head attention mechanism, and then uses a multimodal factor analysis algorithm to adjust the initial weight to obtain a fine-tuning weight, and finally uses the fine-tuning weight to perform weighted fusion on the features to obtain a fused feature, and uses a fully connected layer to classify the fused feature to obtain a classification result;
[0014] The detection module is configured to: obtain the encrypted traffic to be predicted, process the encrypted traffic to be predicted, input the processed encrypted traffic to be predicted into the trained model, and obtain the detection result of the encrypted traffic.
[0015] On the other hand, there is also provided an electronic device, comprising:
[0016] a memory for non-transitory storage of computer-readable instructions; and
[0017] a processor for executing the computer readable instructions,
[0018] When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.
[0019] On the other hand, a storage medium is provided, which non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the method described in the first aspect is executed.
[0020] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is used to implement the method described in the first aspect when running on one or more processors.
[0021] The above technical solution has the following advantages or beneficial effects:
[0022] This method is used for encrypted traffic detection. First, in the data collection phase, encrypted traffic data packets and related metadata are captured from the network communication interface, and multimodal features such as statistics, timing, and protocols are extracted and stored separately. Then data preprocessing is performed, and data cleaning uses a hybrid method of statistics and clustering algorithms to deal with outliers and missing values. Normalization selects appropriate methods based on data characteristics, encoding converts and expands non-numerical data, and data enhancement operations are also performed. After that, a deep convolutional neural network is constructed, which includes convolution, pooling, and fully connected layers, and the input layer has independent channels to receive multiple data.
[0023] In the feature fusion and model training phase, the initial weights are calculated through the multi-head attention mechanism, and then the multimodal factor analysis is used to extract factors and adjust weights. Then the model is trained with the back propagation algorithm, with cross entropy loss as the target, and the parameters are optimized with stochastic gradient descent and early stopping method to prevent overfitting. Finally, in the traffic detection and identification step, the processed encrypted traffic is input into the trained model, and the traffic type, source and potential malicious behavior probability are output, and abnormal traffic is alerted and recorded for analysis.
[0024] Data cleaning improves data purity and model training stability and accuracy, and filling missing values ensures data continuity, which is conducive to the model learning feature rules; normalization enables features to be compared and calculated at the same scale, avoiding uneven feature learning and improving model accuracy and convergence speed; encoding reduces data storage and computational complexity, and enhances the model's ability to recognize category features; data enhancement increases data diversity, improving the model's generalization performance and its ability to recognize different encrypted traffic features.
[0025] The attention mechanism calculates preliminary weights to analyze feature correlations from multiple perspectives, automatically learns feature relationships and highlights the contributions of important features, integrates information to improve recognition accuracy, and enhances system performance reliability; multimodal factor analysis adjusts weights to extract factors, and adjusts them according to their importance to make the fusion take into account both commonality and uniqueness, improve system adaptability and detection accuracy, optimize detection effects, and enhance the ability to identify malicious traffic. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0027] Figure 1 This is a flow chart of the method of embodiment 1. DETAILED DESCRIPTION
[0028] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] Embodiment 1
[0031] This embodiment provides a malicious encrypted traffic detection method based on multimodal feature fusion;
[0032] like Figure 1 As shown in the figure, the malicious encrypted traffic detection method based on multimodal feature fusion includes:
[0033] S101: Construct a data set, where the data set is encrypted traffic of known malicious or non-malicious detection results;
[0034] S102: constructing a deep convolutional neural network model, using a training set to train the deep convolutional neural network model to obtain a trained model; the deep convolutional neural network model includes: a convolutional layer, a pooling layer, a feature fusion layer and a fully connected layer connected in sequence, the feature fusion layer first calculates the initial weight of the feature through a multi-head attention mechanism, and then uses a multimodal factor analysis algorithm to adjust the initial weight to obtain a fine-tuning weight, and finally uses the fine-tuning weight to perform weighted fusion on the features to obtain a fused feature, and uses a fully connected layer to classify the fused feature to obtain a classification result;
[0035] S103: Obtain the encrypted traffic to be predicted, process the encrypted traffic to be predicted, input the processed encrypted traffic to be predicted into the trained model, and obtain the detection result of the encrypted traffic.
[0036] Furthermore, constructing the data set includes:
[0037] (1-1): Capture encrypted traffic data packets from the network communication structure, synchronously collect metadata of each data packet, perform multimodal feature extraction on the encrypted traffic, and obtain statistical features and time series features;
[0038] (1-2): Preprocessing operations are performed on the collected metadata and multimodal features, and the preprocessing operations include: data cleaning, data normalization, data encoding and data enhancement.
[0039] Exemplarily, the (1-1): capturing encrypted traffic data packets from the network communication structure, synchronously collecting metadata of each data packet, performing multimodal feature extraction on the encrypted traffic, and obtaining statistical features and time series features, including:
[0040] Encrypted traffic data packets are captured from the network communication interface, and metadata of each data packet is synchronously collected, including packet size, time interval, and traffic direction information; at the same time, multimodal feature extraction is performed on the encrypted traffic, including statistical feature extraction, that is, obtaining the size distribution of traffic packets, statistical characteristics of time intervals, and timing feature extraction, that is, building the sequence relationship of traffic packets in the time dimension, and protocol feature extraction, that is, collecting metadata information on different protocol layers (such as TCP, UDP), and storing these data and features as independent data sets.
[0041] Furthermore, the data cleaning includes: removing abnormal data by using statistical analysis methods, performing cluster analysis on the data by using clustering algorithms to remove abnormal data and invalid data; and filling in missing data.
[0042] It should be understood that a hybrid method based on statistical analysis and clustering algorithm is used to remove invalid and abnormal data points. For data points that deviate significantly from the overall data distribution, they are divided into abnormal clusters through clustering analysis and eliminated. At the same time, statistical methods are used to identify and fill in some valuable missing values, such as using the mean or median filling method based on adjacent data.
[0043] Furthermore, the use of statistical analysis methods to remove abnormal data includes:
[0044] First, we use statistical analysis methods to calculate the mean μ and standard deviation σ of the data. c The 3σ principle is used, that is, data points that exceed the range of the mean plus or minus 3 times the standard deviation are considered abnormal to preliminarily screen possible abnormal data. For the feature of traffic packet size, if its mean is μ s , with standard deviation σ s , then it is greater than μ s +3σ or less than μ s -3σ s The relevant traffic packet size data will be marked as suspected abnormal points.
[0045] Furthermore, the clustering algorithm is used to perform cluster analysis on the data to remove abnormal data and invalid data, including:
[0046] K-Means clustering is used to perform cluster analysis on the data, and suspected outliers are divided into outlier clusters and removed. The goal of the clustering algorithm is to divide data points into different clusters so that the data points in the same cluster have high similarity, while the data points between different clusters have low similarity.
[0047] In this way, abnormal data that is significantly different from the normal data distribution pattern can be identified more accurately, avoiding misjudgment, improving the purity of the data, and thus improving the stability and accuracy of model training.
[0048] Furthermore, the filling of missing data includes:
[0049] Use the mean or median filling method based on adjacent data.
[0050] For the time interval feature sequence T = [t1, t2, ..., t n ], if t i For missing values, calculate the mean of its adjacent data (when i>1 and i<n) as the filling value; if it is at both ends of the sequence, such as t1 is missing, t2 is used as the filling value; if t n If missing, use t n-1 as fill value.
[0051] This filling method can maintain the continuity and trend of the data to a certain extent, so that the subsequent feature extraction and model training will not be excessively disturbed by missing values, ensuring that the model can fully learn the inherent laws of the time interval characteristics, so as to better detect and analyze the encrypted traffic and improve the reliability and effectiveness of the entire encrypted traffic detection method.
[0052] Furthermore, the data normalization includes:
[0053] Adaptive normalization is used to automatically select the appropriate normalization method based on the dynamic range and distribution characteristics of the data. For example, Z-score normalization is used for data with a near-normal distribution, and Min-Max normalization is used for data within a set interval, mapping each eigenvalue to a unified numerical range.
[0054] The dynamic range and distribution characteristics of the data are analyzed to determine the appropriate normalization method. Skewness and kurtosis are used to measure whether the data is approximately normally distributed. Skewness is used to describe the asymmetry of the data distribution, while kurtosis reflects the concentration of the data around the mean and the steepness of the distribution. If the skewness of the data is close to 0 and the kurtosis is close to 3, it can be considered to be approximately normally distributed, and the Z-score normalization method is used.
[0055] The Z-score normalization formula is:
[0056]
[0057] Where x is the original eigenvalue, μ is the mean of the feature, and σ is the standard deviation. The effect of this formula is to convert the original data into a standard normal distribution with a mean of 0 and a standard deviation of 1, so that different features can be compared and calculated on the same scale. For example, in the packet size feature of encrypted traffic, assuming that the mean of the original packet size data is 100 and the standard deviation is 20, a certain packet size value is 140. After Z-score normalization, its value is (140-100) / 20=2. This allows the packet size feature to be processed at the same numerical level as other features that have undergone the same normalization processing (such as time interval features) in subsequent model training, avoiding excessive or insufficient learning of certain features during model training due to large differences in the range of eigenvalues, thereby improving the accuracy and convergence speed of the model.
[0058] If the data is distributed in a specific interval, for example, some characteristic values of the traffic packet are always in the interval [a, b], then Min-Max normalization is used. The formula is:
[0059]
[0060] Where min and max are the minimum and maximum values of the feature respectively. This normalization method linearly maps the data to the interval [0,1], which can preserve the original distribution of the data while unifying it to the same numerical range.
[0061] For example, for the port number feature of the traffic, if its value range is between 1000 and 65535, the value of a port number of 2000 is approximately 0015 after Min-Max normalization. This allows the model to learn the importance of each feature more evenly when processing features in different ranges, thereby improving the accuracy of encrypted traffic detection and closely integrating with the overall encrypted traffic detection method, providing a good data foundation for subsequent feature extraction, fusion, and training and application of deep convolutional neural network models, ensuring that the entire system can more accurately identify the type, source, and potential malicious behavior of encrypted traffic.
[0062] Furthermore, the data encoding includes:
[0063] For non-numeric data, a hash function-based encoding method is used to convert it into a numerical form suitable for model input. At the same time, the encoded values are expanded by one-hot encoding to enhance the model's ability to recognize category features.
[0064] Furthermore, a coding method based on a hash function is adopted. The hash function has the characteristic of mapping data of any length to a set length value. Let the non-numeric data be D, the hash function be H, and the value after hash coding be H(D). For the protocol type strings such as "TCP" and "UDP" in the network protocol, they can be converted into specific values through the hash function, such as "TCP" is encoded as 56, and "UDP" is encoded as 89.
[0065] It should be understood that the purpose of this encoding method is to quickly and uniformly convert complex non-numeric data into numerical form so that the model can receive and process it, while reducing storage space and computational complexity, improving data processing efficiency, and providing a basic data format for subsequent model training.
[0066] Furthermore, the encoded value is expanded by one-hot encoding. Assuming that there are n different value categories after hash encoding (such as the above protocol type has 2 categories), for a certain encoded value H(D i ), whose unique hot encoding vector is a vector O of length n. If H(D i ) belongs to the jth category, then in vector O, the value of the jth position is 1, and the rest of the positions are 0. For 56 after "TCP" encoding, if there are a total of 2 categories ("TCP" and "UDP"), its one-hot encoding vector is [1,0]; for 89 after "UDP" encoding, its one-hot encoding vector is [0,1].
[0067] It should be understood that the effect of one-hot encoding is to represent the category features in a sparse and clear way, which can highlight the differences between different categories and enhance the model's sensitivity and recognition ability to category features. During the training process of the deep convolutional neural network model, this encoding method enables the model to more clearly distinguish the impact of non-numerical features such as different protocol types on the encrypted traffic detection results, thereby more accurately identifying the type, source and potential malicious behavior of encrypted traffic, and closely cooperating with the feature extraction, fusion, model construction and training links in the overall method to improve the performance and accuracy of the entire encrypted traffic detection system.
[0068] Furthermore, the data enhancement is to generate multiple enhanced data sets by performing random perturbations, adding noise, data transformation and other operations on metadata and feature vectors, including randomly stretching or compressing the time interval of traffic packets, slightly randomly increasing or decreasing the packet size, and randomly replacing or modifying certain fields in the protocol features to improve the generalization ability of the model;
[0069] Furthermore, the time interval of the traffic packets is randomly stretched or compressed, and the original time interval sequence is assumed to be T = [t1, t2, ..., t n], by introducing a random perturbation factor α, the value range of α is between [-β, β], β is a positive number set according to the actual data, such as 01, and randomly stretching or compressing each time interval to obtain the enhanced time interval sequence:
[0070] T′=[(1+α1)t1,(1+α2)t2,…,(1+α n )t n ]
[0071] If the original time interval sequence is [1, 2, 3, 4, 5], and the randomly generated perturbation factor sequence is [0.05, -0.03, 0.08, -0.02, 0.01], then the enhanced time interval sequence is:
[0072] [1×(1+0.05),2×(1-0.03),3×(1+0.08),4×(1-0.02),5×(1+0.01)]
[0073] =[1.05, 1.94, 3.24, 3.92, 5]
[0074] It should be understood that the effect of this operation is to simulate the time jitter that may occur in the actual transmission process of network traffic, so that the model will not rely too much on fixed time interval patterns, thereby improving its ability to identify encrypted traffic with different time characteristics, better adapting to traffic changes in real network environments, and closely integrating with the feature extraction and model training links in the overall framework to enhance the generalization performance of the model in practical applications.
[0075] Furthermore, when the packet size is slightly increased or decreased randomly, the original packet size sequence is assumed to be S = [s1, s2, ..., s m ], and also introduce a random increase or decrease factor γ in the range of [-δ,δ], where δ is a suitable positive number set according to the data and model requirements, such as 5, to generate the enhanced packet size sequence S′=[s1+γ1,s2+γ2,…,s m +γ m For example, if the original packet size sequence is [100, 120, 110, 90, 130], and the randomly generated increase and decrease factor sequence is [-2, 3, -1, 4, -3], then the enhanced packet size sequence is [98, 123, 109, 94, 127].
[0076] It should be understood that the purpose of doing so is to enable the model to adapt to fluctuations in packet size within a certain range, avoid detection errors due to slight changes in packet size, improve the robustness of the model to the characteristics of encrypted traffic packet size, and thereby enhance the ability of the entire model to accurately identify encrypted traffic in complex network traffic environments, provide more diverse and representative data for subsequent deep convolutional neural network training, and enhance the generalization ability of the model, so that it can maintain a high level of detection accuracy when facing different network environments and encrypted traffic characteristics.
[0077] Furthermore, some fields in the protocol features are randomly replaced or modified. In the header information of the network protocol, some optional fields may have different value taking methods. Suppose a field value sequence of the original protocol header is F = [f1, f2, ..., f k ], there exists a value set V = {v1,v2,…,v l}(l is the number of possible values of the field), by randomly selecting elements in the value set, the value of probability P is set according to the actual situation, and the original field value is replaced to obtain the enhanced field value sequence F′.
[0078] For example, for a protocol field, its original value sequence is [A,B,A,C,B], and its value set is {A,B,C,D}. After the random replacement operation, the enhanced field value sequence may be [A,B,D,C,D], that is, the 3rd and 5th elements are replaced this time.
[0079] By simulating the changes in protocol characteristics that may occur when different network devices or software follow protocol specifications, the model can learn various manifestations of protocol characteristics and enhance its understanding and recognition of protocol characteristics. In the overall encrypted traffic detection method, it helps the model to more accurately distinguish different types of encrypted traffic and potential malicious behaviors from the protocol level, improve the applicability and accuracy of the model in the actual network environment, and enhance the reliability and effectiveness of the entire system in detecting encrypted traffic.
[0080] Furthermore, S102: the deep convolutional neural network model comprises: a convolutional layer, a pooling layer, a feature fusion layer and a fully connected layer connected in sequence, wherein the convolutional layer is used to automatically extract local features in metadata and each modal feature, the pooling layer is used to reduce the resolution of the feature map and reduce the amount of calculation, and the fully connected layer is used to integrate and classify the features; in the input layer of the network, independent input channels are set for the metadata and each modal feature, respectively, so as to receive multiple types of data and features at the same time.
[0081] Furthermore, the feature fusion layer first calculates the preliminary weights of the features through a multi-head attention mechanism, including:
[0082] The initial weight of each feature is calculated through the attention mechanism. The attention mechanism dynamically calculates the weight based on the correlation between features and the degree of contribution to the model prediction results. The multi-head attention mechanism is used to map the input multimodal features to multiple different subspaces for attention calculation. Then, the attention results of multiple subspaces are spliced and linearly transformed to obtain the initial attention weight.
[0083] Using a multi-head attention mechanism, assuming that the input multimodal feature set is X = {x1, x2, ..., x n}, where x i Represents the i-th eigenvector, and each eigenvector has a dimension of d.
[0084] Through linear mapping, they are projected into h different subspaces to obtain h groups of feature representations.
[0085]
[0086] in is the corresponding learnable projection weight matrix.
[0087] For multimodal features of encrypted traffic, such as packet size feature vector, time interval feature vector and protocol feature vector, they are mapped into multiple subspaces respectively to analyze the correlation between features from different perspectives.
[0088] In each subspace, the attention score is calculated using the dot product attention formula:
[0089]
[0090] Among them, d k It's K j The softmax function is used to normalize the attention scores so that the sum of the attention weights of all features is 1.
[0091] The purpose of this formula is to assign attention weights based on the correlation between features. Features with higher correlation will be given greater weights during the fusion process. For example, for certain types of encrypted traffic, there may be a strong correlation between the packet size and the time interval. Through the attention mechanism, this correlation can be automatically learned during the model training process, and these two features can be given higher attention weights, thereby highlighting their contribution to traffic type judgment during feature fusion.
[0092] Then, the attention results of the h subspaces are concatenated to obtain Z = [Z 1 ; Z 2 ;…;Z h ], where Z jis the attention output of the jth subspace. Then it passes through a linear transformation layer, that is, O = ZW O , where W O It is the linear transformation weight matrix, and the feature representation O is obtained after the preliminary attention weight allocation.
[0093] This combination of multi-head attention mechanism and linear transformation can integrate the attention information of multiple subspaces and more comprehensively capture the complex relationship between features. It enables the model to flexibly adjust the weight of each feature when facing different encrypted traffic patterns, thereby improving the recognition accuracy of encrypted traffic types, sources, and potential malicious behaviors. It works closely with the overall feature extraction, data preprocessing, and deep convolutional neural network model training to enhance the performance and reliability of the entire encrypted traffic detection system.
[0094] Furthermore, the use of a multimodal factor analysis algorithm to adjust the preliminary weights to obtain fine-tuning weights includes:
[0095] Multimodal factor analysis is used to reduce the dimensionality of each modality feature and extract common factors and special factors. Common factors reflect the common features among the modalities, while special factors retain the unique information of each modality. According to the importance of common factors and special factors, the initial attention weight is adjusted twice, so that the fusion process can fully utilize the commonalities of each modality and highlight its uniqueness.
[0096] For a multimodal feature set, assume that there are M modes and the feature matrix of each mode is X m (m=1,2,…,M), whose dimension is n m ×p, where n m is the number of samples, and p is the feature dimension.
[0097] Multimodal factor analysis is used for dimensionality reduction by constructing a factor model:
[0098] X m =μ m +L m F+∈ m ,
[0099] where μ m is the mean vector of mode m, L m is the factor loading matrix (dimension is p×q, q is the number of common factors, and q<p), F is the common factor matrix (dimension is q×n m ),∈ m is a special factor matrix (dimension p×n m ).
[0100] The maximum likelihood estimation method is used to solve the factor model, and the factor loading matrix L is estimated through an iterative optimization algorithm.m And the common factor matrix F.
[0101] For example, among the three modes of encrypted traffic, namely packet size, time interval and protocol characteristics, through multimodal factor analysis, it may be found that there is a common factor related to traffic periodicity, which is reflected in both packet size and time interval modes. This is the common feature among the modes; and for the protocol characteristic mode, there may be some special factors that reflect information such as the unique header field settings of the protocol. These are the unique information of each mode.
[0102] According to the importance of common factors and special factors, the initial attention weights are adjusted twice. Assume that the initial attention weight vector is w init , dimension is p;
[0103] Calculate the common factor importance score:
[0104]
[0105] Among them, F i is the i-th row vector of the common factor matrix F;
[0106] Special factor importance score:
[0107]
[0108] where ∈ m,i is a special factor matrix∈ m The i-th row vector of .
[0109] Then, through the linear weighting method, the quadratically adjusted attention weight vector w is obtained adj :
[0110]
[0111] Among them, α and β m It is a set weight coefficient used to balance the influence of common factors and special factors on the adjustment of attention weight.
[0112] Taking a specific encrypted traffic detection scenario as an example, for the traffic generated by a specific encrypted application, its packet size and time interval may show periodic changes to a certain extent. This common feature is extracted through common factors. When fusing features, the model pays more attention to these periodicity-related features by adjusting the attention weights. At the same time, some special field settings in the protocol features (such as specific port numbers or protocol options) are also highlighted in the fusion process as special factors, so that when the model determines the type of encrypted traffic, it can not only use the common features of each modality to quickly locate possible traffic categories, but also make more accurate identification based on the unique information of each modality, thereby improving the adaptability and detection accuracy of the entire encrypted traffic detection system to different encrypted traffic modes, and working in coordination with data collection, preprocessing, and the construction and training of deep convolutional neural network models to further optimize the detection effect of encrypted traffic, reduce false alarm and missed alarm rates, better identify potential malicious encrypted traffic behaviors, and ensure network security.
[0113] Furthermore, the training set is used to train the deep convolutional neural network model to obtain the trained model, which is to train the deep convolutional neural network through a back propagation algorithm to minimize the cross entropy loss between the prediction results and the actual traffic type and source label. During the training process, the stochastic gradient descent algorithm is used to optimize the model parameters, and the early stopping method is used to prevent overfitting, that is, when the performance of the model on the validation set does not improve for multiple consecutive training cycles, the training is stopped and the current optimal model parameters are saved;
[0114] Specifically, during the model training process, the cross entropy loss function is used to measure the difference between the predicted results and the actual traffic type and source label. Assume that the predicted output of the model is y pred y pred It is the probability distribution vector normalized by the softmax function, with a dimension of C, where C is the number of categories and the actual label is y true y true is a one-hot encoded vector with a dimension of C. The cross entropy loss function formula is:
[0115]
[0116] The formula is used to quantify the inaccuracy of the model's predictions. By minimizing this loss value, the model is prompted to learn more accurate feature representations and classification decision boundaries. When distinguishing between normal encrypted traffic and malicious encrypted traffic, if the model predicts a malicious traffic sample as malicious traffic with a low probability (close to 0), and the actual label is malicious traffic (the corresponding position is 1), then the loss value calculated by the cross-entropy loss function will be large, and the model will receive a strong signal to adjust the parameters during the back-propagation process to improve the prediction accuracy of such samples.
[0117] Next, the stochastic gradient descent algorithm (SGD) is used to optimize the model parameters. In each iteration, a mini-batch of data is randomly extracted from the training data set, the average gradient of the batch of data is calculated, and then the model parameters are updated according to the learning rate. For example, for a weight parameter w in a deep convolutional neural network, the update formula is:
[0118]
[0119] in is the gradient of the loss function L with respect to the weight w. The learning rate η controls the step size of each parameter update. If it is set too large, the model may skip the optimal solution during training, resulting in failure to converge. If it is set too small, the model training speed will be very slow. In actual training, a learning rate decay strategy is adopted to gradually reduce the learning rate as the training progresses to balance the convergence speed and accuracy of the model.
[0120] To prevent overfitting, the Early Stopping method is used. During the training process, the data set is divided into a training set, a validation set, and a test set. The model is trained on the training set, and the performance is evaluated on the validation set. The accuracy of the model on the validation set is recorded every 5 epochs. If the performance of the model on the validation set does not improve after 10 consecutive cycles, it is considered that the model may begin to overfit. At this time, the training is stopped, and the model parameters with the best performance in the current validation set are saved as the final model. For example, in the detection training of encrypted traffic, as the training cycle increases, the accuracy of the model on the training set may continue to increase, but the accuracy on the validation set may first increase and then decrease. When it is found that the accuracy of the validation set has not improved for 10 consecutive cycles, the training is stopped and the model at this time is saved.
[0121] Furthermore, the S103: obtains the encrypted traffic to be predicted, processes the encrypted traffic to be predicted, inputs the processed encrypted traffic to be predicted into the trained model, obtains the detection result of the encrypted traffic, and inputs the encrypted traffic to be detected into the trained deep convolutional neural network model after the data collection step, the preprocessing step and the feature extraction step. The model outputs the type and source information of the encrypted traffic and the probability assessment of potential malicious behavior. When the probability of potential malicious behavior exceeds a preset threshold, an alarm is issued and relevant traffic information is recorded for further analysis.
[0122] Furthermore, the data collection step includes:
[0123] For long-term network traffic data, a sliding window method is used for data collection and feature extraction. The window size and sliding step are dynamically adjusted according to the rate change of network traffic, the distribution density of data packets and the statistical laws of historical traffic data. By establishing a decision model based on the dynamic characteristics of traffic, various indicators of network traffic are monitored in real time. When the traffic rate increases and the distribution of data packets becomes denser, the window size is appropriately reduced and the sliding step is increased. Conversely, the window size is increased and the sliding step is reduced to ensure the integrity of the data and the accuracy of feature extraction, while improving the real-time performance and efficiency of the system.
[0124] Let R(t) be the network traffic rate at time t, which is obtained by statistically calculating the number of bytes of data packets per unit time; D(t) represents the distribution density of data packets per unit time at time t, that is, the ratio of the number of data packets to time; H(t) is the traffic characteristic statistic calculated based on historical traffic data, such as the mean and variance of the traffic rate and the distribution characteristics of the data packet size over a period of time. These statistics can be updated and maintained by methods such as sliding average to reflect the long-term trends and laws of network traffic.
[0125] Establish a decision model based on the dynamic characteristics of traffic, and dynamically adjust the window size W(t) and sliding step size S(t): when R(t)>αR avg (t) and D(t)>βD avg (t), reduce the window size, where α and β are pre-set thresholds, R avg (t) and D avg (t) are the average values of traffic rate and packet distribution density calculated based on historical traffic data.
[0126] The reduced window size is: W(t)=W(t-1)-γ; wherein γ is the window reduction step size set according to the actual situation.
[0127] At the same time, the sliding step size is increased, S(t)=S(t-1)+δ, where δ is the increase in the sliding step size.
[0128] This is because when the traffic rate is faster and the data packets are densely distributed, a smaller window can capture the rapid changes in traffic more timely, while a larger sliding step can reduce the amount of data processing and improve the real-time performance of the system while ensuring data representativeness. For example, when a distributed denial of service attack (DDoS) occurs, the network traffic rate will soar instantly and a large number of data packets will pour in. At this time, appropriately reducing the window size and increasing the sliding step can quickly collect key features of the attack traffic, such as a large number of data packets from the same source IP in a short period of time. These features are crucial for the subsequent identification of malicious encrypted traffic behavior.
[0129] On the contrary, when R(t)<αR avg (t) and D(t)<βD avg When (t), increase the window size, W(t) = W(t-1) + γ, and reduce the sliding step size, S(t) = S(t-1) - δ c .
[0130] In this way, when the traffic is relatively stable, more comprehensive traffic feature information can be obtained, ensuring data integrity and avoiding the loss of some potentially important features due to a small window size. For example, under normal network traffic conditions, appropriately increasing the window size can cover features such as the trend of traffic packet size changes over a longer period of time and the periodicity of time intervals, which helps the model learn the pattern of normal encrypted traffic more accurately, thereby being able to more accurately distinguish normal traffic from abnormal traffic during detection.
[0131] Embodiment 2
[0132] This embodiment provides a malicious encrypted traffic detection system based on multimodal feature fusion, including:
[0133] A construction module is configured to: construct a data set, wherein the data set is encrypted traffic of known malicious or non-malicious detection results;
[0134] The training module is configured to: construct a deep convolutional neural network model, use a training set to train the deep convolutional neural network model, and obtain a trained model; the deep convolutional neural network model includes: a convolution layer, a pooling layer, a feature fusion layer, and a fully connected layer connected in sequence, the feature fusion layer first calculates the initial weight of the feature through a multi-head attention mechanism, and then uses a multimodal factor analysis algorithm to adjust the initial weight to obtain a fine-tuning weight, and finally uses the fine-tuning weight to perform weighted fusion on the features to obtain a fused feature, and uses a fully connected layer to classify the fused feature to obtain a classification result;
[0135] The detection module is configured to: obtain the encrypted traffic to be predicted, process the encrypted traffic to be predicted, input the processed encrypted traffic to be predicted into the trained model, and obtain the detection result of the encrypted traffic.
[0136] It should be noted that the above-mentioned construction module, training module and detection module correspond to steps S101 to S103 in Embodiment 1, and the examples and application scenarios implemented by the above-mentioned modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned modules as part of the system can be executed in a computer system such as a set of computer executable instructions.
[0137] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0138] The proposed system can be implemented in other ways. For example, the system embodiment described above is only illustrative, and the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0139] Embodiment 3
[0140] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the above embodiment one.
[0141] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0142] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0143] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software.
[0144] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0145] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0146] Embodiment 4
[0147] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is completed.
[0148] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A malicious encrypted traffic detection method based on multimodal feature fusion, characterized in that: include: Constructing a data set, wherein the data set is encrypted traffic of known malicious or non-malicious detection results; Constructing a deep convolutional neural network model, using a training set to train the deep convolutional neural network model to obtain a trained model; the deep convolutional neural network model includes: a convolutional layer, a pooling layer, a feature fusion layer and a fully connected layer connected in sequence, the feature fusion layer first calculates the initial weight of the feature through a multi-head attention mechanism, and then uses a multimodal factor analysis algorithm to adjust the initial weight to obtain a fine-tuning weight, and finally uses the fine-tuning weight to perform weighted fusion on the features to obtain a fused feature, and uses a fully connected layer to classify the fused feature to obtain a classification result; The encrypted traffic to be predicted is obtained, the encrypted traffic to be predicted is processed, and the processed encrypted traffic to be predicted is input into the trained model to obtain the detection result of the encrypted traffic.
2. The method for detecting malicious encrypted traffic based on multimodal feature fusion as claimed in claim 1, characterized in that: include: The constructing of the data set comprises: Capture encrypted traffic data packets from the network communication structure, synchronously collect metadata of each data packet, perform multimodal feature extraction on the encrypted traffic, and obtain statistical features and time series features; Performing preprocessing operations on the collected metadata and multimodal features, wherein the preprocessing operations include: data cleaning, data normalization, data encoding, and data enhancement; The data cleaning includes: removing abnormal data by using statistical analysis methods, performing cluster analysis on the data by using clustering algorithms, removing abnormal data and invalid data; and filling in missing data.
3. The malicious encrypted traffic detection method based on multimodal feature fusion as described in claim 2 is characterized in that: include: The method of removing abnormal data by using statistical analysis includes: First, we use statistical analysis methods to calculate the mean μ and standard deviation σ of the data. c ; The 3σ principle is used, that is, data points that exceed the mean plus or minus 3 times the standard deviation are considered abnormal to preliminarily screen abnormal data; for the traffic packet size feature, if its mean is μ s , with standard deviation σ s , then it is greater than μ s +3σ or less than μ s -3σ s The relevant traffic packet size data will be marked as suspected abnormal points; The clustering algorithm is used to perform cluster analysis on the data to remove abnormal data and invalid data, including: Use K-Means clustering to perform cluster analysis on the data, divide the suspected outliers into outlier clusters and remove them; The filling of missing data includes: using a mean or median filling method based on adjacent data; for the time interval feature sequence T = [t1, t2, ..., t n ], if t i For missing values, calculate the mean of its adjacent data As the filling value; if it is at both ends of the sequence, such as t1 is missing, t2 is used as the filling value; if t n If missing, use t n-1 As fill value; The data normalization includes: using adaptive normalization to automatically select a suitable normalization method according to the dynamic range and distribution characteristics of the data, using Z-score normalization for data with a near normal distribution, and using Min-Max normalization for data within a set interval, to map each eigenvalue to a unified value range; The data encoding includes: for non-numeric data, using an encoding method based on a hash function to convert it into a numerical form suitable for model input, and at the same time, performing a one-hot encoding expansion on the encoded numerical value.
4. The method for detecting malicious encrypted traffic based on multimodal feature fusion as claimed in claim 2, characterized in that: include: The data enhancement is to generate multiple enhanced data sets by randomly perturbing metadata and feature vectors, adding noise, and performing data transformation operations, including randomly stretching or compressing the time interval of traffic packets, slightly randomly increasing or decreasing the packet size, and randomly replacing or modifying certain fields in the protocol features; The time interval of the traffic packet is randomly stretched or compressed. Assume that the original time interval sequence is T = [t1, t2, ..., t n ], by introducing a random disturbance factor α, the value range of α is between [-β, β], β is a positive number set according to the actual data, and randomly stretching or compressing each time interval to obtain the enhanced time interval sequence: T′=[(1+α1)t1,(1+α2)t2,…,(1+α n )t n ].
5. The method for detecting malicious encrypted traffic based on multimodal feature fusion as claimed in claim 1, characterized in that: include: The deep convolutional neural network model includes: a convolutional layer, a pooling layer, a feature fusion layer and a fully connected layer connected in sequence, wherein the convolutional layer is used to automatically extract metadata and local features in each modal feature, the pooling layer is used to reduce the resolution of the feature map and reduce the amount of calculation, and the fully connected layer is used to integrate and classify the features; in the input layer of the network, independent input channels for metadata and each modal feature are respectively set so as to simultaneously receive multiple types of data and features.
6. The method for detecting malicious encrypted traffic based on multimodal feature fusion as claimed in claim 1, characterized in that: include: The feature fusion layer first calculates the initial weights of the features through a multi-head attention mechanism, including: The initial weight of each feature is calculated through the attention mechanism. The attention mechanism dynamically calculates the weight based on the correlation between features and the degree of contribution to the model prediction results. The multi-head attention mechanism is used to map the input multimodal features to multiple different subspaces for attention calculation. Then, the attention results of multiple subspaces are spliced and linearly transformed to obtain the initial attention weight. Using a multi-head attention mechanism, assuming that the input multimodal feature set is X = {x1, x2, ..., x n }, where x i represents the i-th eigenvector, each eigenvector has a dimension of d; Through linear mapping, they are projected into h different subspaces to obtain h groups of feature representations. Where j = 1, 2, ..., h, is the corresponding learnable projection weight matrix; For the packet size feature vector, time interval feature vector and protocol feature vector multimodal features of encrypted traffic, they are mapped to multiple subspaces respectively to analyze the correlation between features from different perspectives; In each subspace, the attention score is calculated using the dot product attention formula: Among them, d k It's K j The softmax function is used to normalize the attention scores so that the sum of the attention weights of all features is 1; Then, the attention results of the h subspaces are concatenated to obtain Z = [Z 1 ; Z 2 ;…;Z h ], where Z j is the attention output of the jth subspace; Then pass through a linear transformation layer, O = ZW O , where W O It is the linear transformation weight matrix, and the feature representation O is obtained after the preliminary attention weight allocation.
7. The malicious encrypted traffic detection method based on multimodal feature fusion as claimed in claim 1 is characterized in that: include: The method of using a multimodal factor analysis algorithm to adjust the preliminary weights to obtain fine-tuning weights includes: For a multimodal feature set, assume that there are M modes and the feature matrix of each mode is X m , whose dimension is n m ×p, where n m is the number of samples, p is the feature dimension; Multimodal factor analysis is used for dimensionality reduction by constructing a factor model: X m =μ m +L m F+∈ m , where μ m is the mean vector of mode m, L m is the factor loading matrix, F is the common factor matrix, ∈ m is a special factor matrix; The maximum likelihood estimation method is used to solve the factor model, and the factor loading matrix L is estimated through an iterative optimization algorithm. m and the common factor matrix F; According to the importance of common factors and special factors, the initial attention weights are adjusted twice. Assume that the initial attention weight vector is w init , dimension is p; Calculate the common factor importance score: where F i is the i-th row vector of the common factor matrix F; Special factor importance score: where ∈ m,i is a special factor matrix∈ m The i-th row vector of ; Then, through the linear weighting method, the quadratically adjusted attention weight vector w is obtained adj : Among them, α and β m It is a set weight coefficient used to balance the influence of common factors and special factors on the adjustment of attention weight.
8. The method for detecting malicious encrypted traffic based on multimodal feature fusion as claimed in claim 1, characterized in that: include: The training set is used to train the deep convolutional neural network model to obtain the trained model. The deep convolutional neural network is trained by a back propagation algorithm to minimize the cross entropy loss between the prediction results and the actual traffic type and source label. During the training process, the stochastic gradient descent algorithm is used to optimize the model parameters, and the early stopping method is used to prevent overfitting. When the performance of the model on the validation set does not improve for multiple consecutive training cycles, the training is stopped and the current optimal model parameters are saved.
9. The method for detecting malicious encrypted traffic based on multimodal feature fusion as claimed in claim 1, characterized in that: include: Acquire the encrypted traffic to be predicted, process the encrypted traffic to be predicted, input the processed encrypted traffic to be predicted into the trained model, obtain the detection result of the encrypted traffic, input the encrypted traffic to be detected into the trained deep convolutional neural network model after the data collection step, the preprocessing step and the feature extraction step, and the model outputs the type and source information of the encrypted traffic and the probability evaluation of the potential malicious behavior, wherein when the probability of the potential malicious behavior exceeds a preset threshold, issue an alarm and record the relevant traffic information for further analysis; The data collection step includes: for long-term network traffic data, a sliding window is used to collect data and extract features, and the window size and sliding step length are dynamically adjusted according to the rate change of network traffic, the distribution density of data packets and the statistical law of historical traffic data, and a decision model based on the dynamic characteristics of traffic is established to monitor various indicators of network traffic in real time. When the traffic rate is accelerated and the distribution of data packets becomes dense, the window size is reduced and the sliding step length is increased, otherwise the window size is increased and the sliding step length is reduced; Establish a decision model based on the dynamic characteristics of traffic, and dynamically adjust the window size W(t) and sliding step size S(t): when R(t)>αR avg (t) and D(t)>βD avg (t), reduce the window size, where α and β are pre-set thresholds, R avg (t) and D avg (t) are the average values of the flow rate and packet distribution density calculated based on historical flow data respectively; the reduced window size is: W(t) = W(t-1) - γ; where γ is the window reduction step size set according to the actual situation; at the same time, the sliding step size is increased, S(t) = S(t-1) + δ; where δ is the increase in the sliding step size; On the contrary, when R(t)<αR avg (t) and D(t)<βD avg When (t), increase the window size, W(t) = W(t-1) + γ, and reduce the sliding step size, S(t) = S(t-1) - δ c .
10. A malicious encrypted traffic detection system based on multimodal feature fusion, characterized in that: include: A construction module is configured to: construct a data set, wherein the data set is encrypted traffic of known malicious or non-malicious detection results; The training module is configured to: construct a deep convolutional neural network model, use a training set to train the deep convolutional neural network model, and obtain a trained model; the deep convolutional neural network model includes: a convolution layer, a pooling layer, a feature fusion layer, and a fully connected layer connected in sequence, the feature fusion layer first calculates the initial weight of the feature through a multi-head attention mechanism, and then uses a multimodal factor analysis algorithm to adjust the initial weight to obtain a fine-tuning weight, and finally uses the fine-tuning weight to perform weighted fusion on the features to obtain a fused feature, and uses a fully connected layer to classify the fused feature to obtain a classification result; The detection module is configured to: obtain the encrypted traffic to be predicted, process the encrypted traffic to be predicted, input the processed encrypted traffic to be predicted into the trained model, and obtain the detection result of the encrypted traffic.
Citation Information
Cited By
Encrypted traffic threat detection method and device based on multi-modal feature fusion
CN120415907A
Intelligent data management system for train fault monitoring
CN120833061A
Method and system for detecting legal network traffic in complex environment
CN121000523A
Concealed transmission channel encrypted traffic classification detection method and device based on multi-mode multi-task semi-supervised learning
CN121705918A