Encrypted traffic threat detection method and device based on multi-modal feature fusion
Through the combination of multimodal feature fusion and GRU model, the problem of low accuracy of single-modal feature detection is solved, and efficient and accurate threat detection of encrypted traffic is achieved, especially the identification of emerging threats.
Patent Information
- Application Number
- CN202510905353.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
In the prior art, the encryption traffic detection method relies on a single modal feature, resulting in limited detection accuracy and difficulty in identifying emerging or highly concealed threat encrypted traffic.
By extracting the statistical features and content features of encrypted traffic, a multimodal feature fusion model is constructed, combining optimization algorithms to find weight matrix and bias, using the GRU model for timing analysis, and combining the Emperor Penguin optimization algorithm and grid search optimization learning rate to achieve multimodal feature fusion and threat detection.
It improves the accuracy and efficiency of encrypted traffic detection, enhances the ability to identify emerging threats, reduces the rate of missed and false alarms, and realizes real-time and accurate threat detection.
Smart Images

Figure CN120415907A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security. Specifically, it particularly relates to an encrypted traffic threat detection method and device based on multi-modal feature fusion. Background Art
[0002] With the wide application of network encryption technology, while encrypted traffic protects data privacy, it also provides a hidden channel for malicious attacks. Traditional encrypted traffic threat detection methods mainly rely on single-modal features for analysis, and there are still limitations in the existing technology.
[0003] Chinese Patent CN114363020B discloses an encrypted traffic detection method, system, device, and storage medium. In the present invention, before the encrypted traffic reaches the target system, the traffic packet is captured and analyzed, and the query domain name and IP address are extracted. If the domain name matches the threat intelligence library, the corresponding IP is added to the local blacklist. Subsequently, after the traffic enters the system, the target access address and the certificate information in the session connection are further extracted and matched with the local blacklist and the threat intelligence library respectively. According to the matching results, the system intelligently determines whether the traffic constitutes a threat and adds the suspicious certificate information to the blacklist.
[0004] In the existing technology, there are problems such as insufficient utilization of encrypted traffic feature information, and it is difficult for single-modal features to comprehensively represent complex threat behaviors, resulting in limited detection accuracy. Simply relying on the threat intelligence library and certificate information for detection, it is often difficult to achieve accurate identification when facing emerging or highly concealed threat encrypted traffic. Summary of the Invention
[0005] Aiming at the problems in the related technology, the present invention provides an encrypted traffic threat detection method based on multi-modal feature fusion. Through feature fusion technology and deep learning technology, the present invention solves the problems of insufficient utilization of encrypted traffic feature information and low detection accuracy of threat encrypted traffic.
[0006] To solve the above technical problems, the present invention is realized through the following technical solutions: S1. Extract the statistical features and content features of the encrypted traffic to obtain a multi-modal feature training set; S2. Construct a multi-modal feature fusion model, combine the multi-modal feature training set and an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model, and obtain the optimal solution; use the optimal solution as the weight matrix and bias of the multi-modal feature fusion model to obtain a multi-modal encrypted traffic feature fusion model; S3. Build a GRU model, collect historical encrypted traffic and historical encrypted traffic threat labels; extract the statistical features and content features of the historical encrypted traffic, and input them into the multi-modal encrypted traffic feature fusion model to obtain the multi-modal fusion features of the historical encrypted traffic; use the multi-modal fusion features of the historical encrypted traffic and the historical encrypted traffic threat labels to train the GRU model to obtain the GRU encrypted traffic threat monitoring model; S4. Collect real-time encrypted traffic, extract the statistical features and content features of the real-time encrypted traffic, and input them into the multi-modal encrypted traffic feature fusion model to obtain the multi-modal fusion features of the real-time encrypted traffic and input them into the GRU encrypted traffic threat monitoring model to obtain the threat detection results of the real-time encrypted traffic.
[0007] Preferably, the S1 includes the following steps: S11. Capture the encrypted traffic through the mirror port to obtain the original encrypted traffic; perform hash desensitization on the sensitive fields in the original encrypted traffic to obtain the encrypted traffic; S12. Extract the statistical features and content features of the encrypted traffic to obtain the multi-modal feature training set of the encrypted traffic; The above steps capture the original encrypted traffic through the mirror port, perform hash desensitization processing on the sensitive fields, extract the statistical features and content features of the desensitized traffic, and construct a multi-modal feature training set; rich feature information is collected.
[0008] Preferably, the S12 includes the following steps: S121. Extract the number of data packets, average packet length, traffic entropy, variance of time interval, and protocol type of the encrypted traffic to obtain session-level features; extract the session frequency of the same source IP, destination port distribution, and traffic burstiness of the encrypted traffic to obtain host-level features; the session-level features and host-level features together constitute the statistical features of the encrypted traffic; S122. Intercept the first N bytes of the encrypted traffic, fill in zeros for the insufficient part, and normalize the byte sequence of the encrypted traffic to obtain the content features of the encrypted traffic; S123. Correlate and store the statistical features and content features to obtain a multi-modal training set; The above steps extract key statistical indicators such as the number of data packets and average packet length from the session level, analyze features such as the session frequency of the source IP at the host level, thus forming comprehensive statistical features; capture the content features by intercepting and normalizing the byte sequence of the encrypted traffic; correlate and store these two types of features to construct a multi-modal training set; deeply mine the multi-dimensional feature information of the encrypted traffic.
[0009] Preferably, the S2 includes the following steps: S21. Construct a multi-modal feature fusion model and set the weight matrix and bias of the multi-modal feature fusion model; the multi-modal feature fusion model includes a statistical feature encoder, a content feature encoder, and a cross-modal projection layer; S22. Add Gaussian noise to the time intervals in the statistical features in the training set and randomly discard some features to obtain enhanced statistical features; perform random masking and block permutation operations on the byte sequences in the content features in the training set to obtain enhanced content features; S23. Through the statistical feature encoder, content feature encoder, and cross-modal projection layer, vectorize the enhanced statistical features, enhanced content features, positive sample pairs, and negative sample pairs to obtain enhanced statistical feature vectors, enhanced content feature vectors, positive sample pair vectors, and negative sample pair vectors; S24. Match the enhanced statistical features and enhanced content features of the same session to obtain positive sample pairs; Randomly match the enhanced statistical features and enhanced content features of different sessions in the training set to obtain negative sample pairs; S25. Calculate the difference between the enhanced statistical feature vector and the enhanced content feature vector through a modal alignment loss function to obtain a modal alignment difference; Calculate the ratio of the similarity between positive sample pairs to the similarity between positive sample pairs and all negative sample pairs through a cross-modal contrast loss function to obtain a cross-modal contrast difference; Use an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model to minimize the modal alignment difference and the cross-modal contrast difference, and obtain the optimal solution; use the optimal solution as the weight matrix and bias of the multi-modal feature fusion model to obtain an improved multi-modal feature fusion model; Set the feature fusion formula of the improved multi-modal feature fusion model; The above steps construct a multi-modal feature fusion model including a statistical feature encoder, a content feature encoder, and a cross-modal projection layer, and set the initial weight matrix and bias; perform enhancement processing on the statistical features and content features of the training set, such as adding Gaussian noise, randomly discarding features, and random masking and block permutation operations, to improve the robustness of the model; vectorize the enhanced features and samples through different components of the model; form positive sample pairs by matching the enhanced features of the same session, and randomly match the features of different sessions to form negative sample pairs; use the modal alignment loss function and the cross-modal contrast loss function to calculate the differences between feature vectors, and adjust the model parameters through an optimization algorithm to minimize these differences, thereby obtaining a multi-modal encrypted traffic feature fusion model; the multi-modal encrypted traffic feature fusion model provides a more powerful and accurate tool for the analysis, classification, and anomaly detection of encrypted traffic, which helps to improve the accuracy and efficiency of security analysis.
[0010] Preferably, in S25, the optimization algorithm is used to find the weight matrix and bias of the multi-modal feature fusion model, and obtaining the optimal solution includes the following steps: S251. Construct a population of emperor penguins, set the size of the population of emperor penguins, and set the maximum number of optimization iterations; S252. Set the initial positions of the population of emperor penguins according to the weight matrix and bias to obtain the initial position set of the population of emperor penguins; S253. Define a fitness function according to the modality alignment difference and the cross-modal contrast difference; S254. Perform iterative operations on the initial position set of the population of emperor penguins, calculate the fitness value of each position in the initial position set of the population of emperor penguins, update the positions of each emperor penguin in the initial position set of the population of emperor penguins, and obtain the best individual position of the emperor penguin in the population of emperor penguins and the global best position of the emperor penguin during each round of iteration; the higher the fitness value, the better the position; during each round of iteration, according to the fitness function; S255. Repeat S254. When the maximum number of test iterations is reached, stop the iteration to obtain the optimal solution; The above steps finely adjust the weight matrix and bias of the multi-modal feature fusion model through the emperor penguin optimization algorithm to ensure the optimal performance of the model; through steps of constructing a population of emperor penguins, setting the initial positions, defining the fitness function, and iteratively updating the positions, this algorithm can efficiently search the parameter space to find the optimal solution that minimizes the modality alignment difference and the cross-modal contrast difference; it not only improves the accuracy and efficiency of the model in processing encrypted traffic, but also enhances its generalization ability, enabling the model to maintain stable performance in complex and changeable scenarios.
[0011] Preferably, S3 includes the following steps: S31. Construct a GRU model and set the initial learning rate of the GRU; S32. Collect historical encrypted traffic, extract the statistical features and content features of the historical encrypted traffic, set the normal label and abnormal label of the historical encrypted traffic data to obtain the threat label of the historical encrypted traffic; input the statistical features and content features of the historical encrypted traffic into the improved multi-modal feature fusion model to obtain the multi-modal fusion features of the historical encrypted traffic; S33. Use the multi-modal fusion features of the historical encrypted traffic and the threat label of the historical encrypted traffic to train the GRU model. During the training process, combine grid search and cross-validation algorithm to find the learning rate of the GRU model, obtain the optimal learning rate, use the optimal learning rate as the learning rate of the GRU model, and obtain the GRU encrypted traffic threat detection model; The above steps include constructing a GRU model and initializing the learning rate; collecting historical encrypted traffic data, extracting its statistical and content features, and attaching threat labels, and obtaining fused features using an improved multi-modal feature fusion model; training the GRU model using these fused features and labels, and optimizing the learning rate by combining grid search and cross-validation algorithms to obtain an optimal GRU encrypted traffic threat detection model; improving the accuracy and efficiency of threat detection.
[0012] Preferably, in the training process in S33, finding the learning rate of the GRU model by combining grid search and cross-validation algorithms and obtaining the optimal learning rate includes the following steps: S331. Set the parameter search space, the maximum number of random searches, and the number of cross-validation folds of the grid search combined with the cross-validation algorithm to X; S332. Randomly set the initial learning rate set; Initialize the best performance of the GRU model as U; sequentially select the learning rate u i from the initial learning rate set as the learning rate of the GRU model, divide the historical encrypted traffic multi-modal fused features and historical encrypted traffic threat labels into X folds, for each fold x, {x|1≤x≤X}; use all historical encrypted traffic multi-modal fused features and historical encrypted traffic threat label data except the x-th fold to train the GRU model, use the training data of the x-th fold to evaluate the model, obtain a set of performance metrics, and calculate the mean of the performance metrics of all folds to obtain the average performance F; If the average performance F > the best performance U, then update the best performance to K and update the optimal learning rate to u i ; otherwise, keep the original best performance and optimal learning rate; S333. Repeat S332, and when the maximum number of random searches is reached, stop the iteration to obtain the optimal learning rate; The above steps systematically find the optimal learning rate of the GRU model by combining grid search and cross-validation algorithms, effectively improving the performance and generalization ability of the model in encrypted traffic threat detection; avoiding the subjectivity and randomness of learning rate selection, ensuring the stable performance of the model on different data folds, and enabling the model to more accurately learn and predict threat patterns through the determination of the optimal learning rate, improving the accuracy and efficiency of detection.
[0013] Preferably, S4 includes the following steps: S41. Collect real-time encrypted traffic and extract the statistical features and content features of the real-time encrypted traffic; S42. Input the statistical features and content features of the real-time encrypted traffic into the multi-modal encrypted traffic feature fusion model to obtain the real-time encrypted traffic multi-modal fused features; S43. Input the multi-modal fusion features of real-time encrypted traffic into the GRU encrypted traffic threat detection model to obtain the encrypted traffic threat detection result; The above steps collect real-time encrypted traffic and extract its statistical features and content features; input these features into the multi-modal encrypted traffic feature fusion model to obtain the multi-modal fusion features of real-time encrypted traffic; input the fusion features into the pre-trained GRU encrypted traffic threat detection model, realizing the real-time detection of encrypted traffic threats; not only improving the real-time performance and accuracy of threat detection, but also enhancing the depth and breadth of detection through the combination of multi-modal feature fusion and the GRU model, providing strong and reliable technical support for the real-time defense of encrypted traffic threats.
[0014] An encrypted traffic threat detection system based on multi-modal feature fusion for implementing the above-mentioned encrypted traffic threat detection method based on multi-modal feature fusion, including a multi-modal feature extraction module, a multi-modal feature fusion optimization module, a GRU threat detection model construction module, and a real-time encrypted traffic threat detection module; The multi-modal feature extraction module is used to capture encrypted traffic through a mirror port, desensitize sensitive fields by hashing, and extract statistical features and content features to obtain a multi-modal feature training set; The multi-modal feature fusion optimization module is used to construct a fusion model including a statistical feature encoder, a content feature encoder, and a cross-modal projection layer, and use data augmentation technology to improve the model robustness for the multi-modal feature training set. By defining a modal alignment loss function and a cross-modal contrast loss function, and combining the emperor penguin optimization algorithm to globally search for the optimal weight matrix and bias parameters, a multi-modal encrypted traffic feature fusion model is obtained; The GRU threat detection model construction module is used to adopt a GRU network and use its gating mechanism to capture the long-term temporal dependence relationship of encrypted traffic; train the model through the multi-modal encrypted traffic fusion features of historical encrypted traffic and their threat labels, and combine grid search and cross-validation to optimize the learning rate parameter to obtain a GRU encrypted traffic threat detection model; The real-time encrypted traffic threat detection module is used to extract its statistical features and content features after real-time collecting encrypted traffic, input them into the multi-modal encrypted traffic feature fusion model to generate a fusion feature vector; input the fusion feature vector into the GRU encrypted traffic threat detection model for dynamic temporal analysis, and output the threat determination result of real-time traffic.
[0015] An encrypted traffic threat detection device based on multi-modal feature fusion, on which a program is stored, and when the program is executed by a processor, it is used to implement the above-mentioned encrypted traffic threat detection method based on multi-modal feature fusion.
[0016] Beneficial effects The present invention has the following beneficial effects: The present invention realizes the enhancement of multimodal feature complementarity and improves the comprehensiveness of detection. By simultaneously extracting the statistical features and content features of encrypted traffic and combining the cross-modal projection layer and the contrast learning mechanism, the macro behavior pattern of traffic and the micro payload information are effectively fused; the statistical features can capture the temporal patterns and abnormal fluctuations of traffic, while the content features can identify potential threat patterns in the encrypted payload; the complementarity of the two reduces the information deviation of a single modality, especially when facing emerging or low-feature threats, and can reduce the false negative and false positive rates through multi-dimensional feature cross-verification.
[0017] The present invention realizes the cross-modal fusion optimization mechanism and improves the feature discriminability. By introducing the modality alignment loss function and the cross-modal contrast loss function, through the similarity optimization of positive and negative sample pairs, the consistency constraint and difference retention of statistical features and content features are adaptively balanced; it can suppress redundant noise and at the same time strengthen the cross-modal association of threat-related features, thereby generating more discriminative fusion features and providing high-quality input for subsequent temporal modeling.
[0018] The present invention realizes dynamic temporal modeling and lightweight design, taking into account both accuracy and real-time performance; uses the GRU model to perform temporal modeling on the multimodal fusion features, and its gating mechanism can effectively capture the long-term dependencies in encrypted traffic, while reducing the number of parameters and the computational complexity; combined with the learning rate optimized by grid search and cross-validation, it further accelerates the model convergence, enabling it to adapt to real-time traffic monitoring scenarios and meet the low-latency requirements while ensuring high detection accuracy.
[0019] The present invention realizes the drive of intelligent optimization algorithms and improves the generalization ability of the model. In the feature fusion and model training stages, the emperor penguin optimization algorithm and the grid search cross-validation strategy are respectively introduced to realize the automatic optimization of the weight matrix, bias parameters and learning rate; the emperor penguin optimization algorithm avoids the problem that the traditional gradient descent method is prone to fall into local optima through population iteration and fitness function design, ensuring the global optimality of the parameters of the multimodal fusion model; while the cross-validation mechanism reduces the overfitting risk through multi-fold evaluation and enhances the generalization ability of the model to unknown traffic.
[0020] Of course, it is not necessary for any product implementing the present invention to achieve all the above advantages simultaneously. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0022] Figure 1 This is a schematic flowchart of a method for detecting encrypted traffic threats based on multimodal feature fusion according to the present invention; Figure 2 This is a schematic diagram of the modules of a system for detecting encrypted traffic threats based on multimodal feature fusion according to the present invention. Specific embodiments
[0023] Next, the technical solutions in the embodiments of the invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the invention. Obviously, the described embodiments are only a part of the embodiments of the invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the invention without creative efforts belong to the scope of protection of the invention.
[0024] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "top", "middle", "inner", etc. indicating orientation or positional relationships are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0025] Embodiment 1 Please refer to Figure 1 , the present invention discloses a method for detecting encrypted traffic threats based on multimodal feature fusion, including the following steps: S1. Extract the statistical features and content features of the encrypted traffic to obtain a multimodal feature training set; The S1 includes the following steps: S11. Capture the encrypted traffic through the mirror port to obtain the original encrypted traffic; desensitize the sensitive fields in the original encrypted traffic by hashing to obtain the encrypted traffic; the original encrypted traffic such as HTTPS, SSH, VPN traffic, and the sensitive fields such as IP address, MAC address; S12. Extract the statistical features and content features of the encrypted traffic to obtain a multimodal feature training set of the encrypted traffic; The S12 includes the following steps: S121. Extract the number of data packets, average packet length, traffic entropy, time interval variance, and protocol type of the encrypted traffic to obtain session-level features; extract the session frequency of the same source IP, destination port distribution, and traffic burstiness of the encrypted traffic to obtain host-level features; the session-level features and host-level features together constitute the statistical features of the encrypted traffic; S122. Intercept the first N bytes of the encrypted traffic, pad the insufficient part with zeros, and normalize the byte sequence of the encrypted traffic to obtain the content features of the encrypted traffic; S123. Associate and store the statistical features with the content features to obtain a multi-modal training set; S2. Construct a multi-modal feature fusion model, and combine the multi-modal feature training set and an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model to obtain an optimal solution; Use the optimal solution as the weight matrix and bias of the multi-modal feature fusion model to obtain a multi-modal encrypted traffic feature fusion model; The S2 includes the following steps: S21. Construct a multi-modal feature fusion model and set the weight matrix and bias of the multi-modal feature fusion model; The multi-modal feature fusion model includes a statistical feature encoder, a content feature encoder, and a cross-modal projection layer; S22. Add Gaussian noise to the time intervals in the statistical features in the training set and randomly discard some features to obtain enhanced statistical features; Perform random masking and block permutation operations on the byte sequences in the content features in the training set to obtain enhanced content features; S23. Through the statistical feature encoder, content feature encoder, and cross-modal projection layer, vectorize the enhanced statistical features, enhanced content features, positive sample pairs, and negative sample pairs to obtain enhanced statistical feature vectors, enhanced content feature vectors, positive sample pair vectors, and negative sample pair vectors; S24. Match the enhanced statistical features and enhanced content features of the same session to obtain positive sample pairs; For example, the original session: statistical feature S + content feature C, after enhancement: statistical feature S' (with added noise) + content feature C' (masked) to obtain positive sample pair (S', C'); Generate positive sample pairs and negative sample pairs; Randomly match the enhanced statistical features and enhanced content features of different sessions in the training set to obtain negative sample pairs; For example, the enhanced statistical feature S_A' of session A + the enhanced content feature C_B' of session B to obtain negative sample pair (S_A', C_B'); S25. Through the modal alignment loss function, calculate the difference between the enhanced statistical feature vector and the enhanced content feature vector to obtain the modal alignment difference; The formula of the modal alignment loss function is as follows, ; where H represents the modal alignment loss function, h mm represents the enhanced statistical feature vector, h bb represents the enhanced content feature vector; Through the cross-modal contrast loss function, calculate the ratio of the similarity between positive sample pairs to the similarity between positive sample pairs and all negative sample pairs to obtain the cross-modal contrast difference; The formula of the cross-modal contrast loss function is as follows, ; where, L represents the cross-modal contrast loss function, hp Denote the positive sample pair as h q Denote the negative sample pair as s( ), where s( ) represents the cosine similarity function, and N represents the total number of negative sample pairs; Use an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model, minimize the modal alignment difference and cross-modal contrast difference to obtain the optimal solution; use the optimal solution as the weight matrix and bias of the multi-modal feature fusion model to obtain an improved multi-modal feature fusion model; Set the feature fusion formula of the improved multi-modal feature fusion model as follows ; where α represents the attention weight, W and b represent the weight matrix and bias of the improved multi-modal feature fusion model respectively, [h m ,h b represents the concatenation of the statistical feature vector and the content feature vector, and h n represents the multi-modal fusion feature; The step of using an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model in S25 to obtain the optimal solution includes the following steps: S251. Construct a king penguin population, set the size of the king penguin population as g, then the king penguin population is denoted as d = {d1, d2, …, d i , …, d g}, where n i represents the i-th king penguin in the king penguin population; set the maximum number of optimization iterations; S252. According to the weight matrix and bias, set the initial positions of the king penguin population to obtain the initial position set of the king penguin population as e = {(f 11 , f 12 ), (f 21 , f 22 ), …, (f i1 , f i2 ), …, (f g1 , f g2 ), where x i1 represents the first-dimensional position of the i-th king penguin in the king penguin population, and x i2 represents the second-dimensional position of the i-th king penguin in the king penguin population; S253. Define a fitness function according to the modal alignment difference and cross-modal contrast difference. The fitness function formula is as follows ; S254. Perform iterative operations on the initial position set of the emperor penguin population. The higher the fitness value, the better the position. During each round of iteration, according to the fitness function, calculate the fitness value of each position in the initial position set of the emperor penguin population, and update the positions of each emperor penguin in the initial position set of the emperor penguin population from high to low according to the fitness value. And during each round of iteration, obtain the position of the best emperor penguin individual and the global best emperor penguin position in the emperor penguin population; S255. Repeat S254. When the maximum number of test iterations is reached, stop the iteration to obtain the optimal solution.
[0026] S3. Build a GRU model, collect historical encrypted traffic and historical encrypted traffic threat labels; extract the statistical features and content features of historical encrypted traffic, and input them into the multi-modal encrypted traffic feature fusion model to obtain the multi-modal fusion features of historical encrypted traffic; use the multi-modal fusion features of historical encrypted traffic and historical encrypted traffic threat labels to train the GRU model to obtain the GRU encrypted traffic threat monitoring model; The S3 includes the following steps: S31. Build a GRU model and set the initial learning rate of the GRU; S32. Collect historical encrypted traffic, extract the statistical features and content features of historical encrypted traffic, set the normal label and abnormal label of historical encrypted traffic data to obtain historical encrypted traffic threat labels; input the statistical features and content features of historical encrypted traffic into the improved multi-modal feature fusion model to obtain the multi-modal fusion features of historical encrypted traffic; the threat labels include normal encrypted traffic, DDoS attack, data leakage, and encrypted mining; S33. Use the multi-modal fusion features of historical encrypted traffic and historical encrypted traffic threat labels to train the GRU model. During the training process, combine grid search with the cross-validation algorithm to find the learning rate of the GRU model, obtain the optimal learning rate, and use the optimal learning rate as the learning rate of the GRU model to obtain the GRU encrypted traffic threat detection model; The process of finding the learning rate of the GRU model by combining grid search with the cross-validation algorithm during the training in S33 and obtaining the optimal learning rate includes the following steps: S331. Set the parameter search space, the maximum number of random searches, and the number of cross-validation folds of the grid search combined with the cross-validation algorithm to X; S332. Randomly set the initial learning rate set u = {u1, u2, …, u i , …, u v}; where, u i represents the i-th learning rate in the initial learning rate set, and v represents the total number of learning rates in the initial learning rate set; The best performance of initializing the GRU model is U; select the learning rate u from the initial learning rate set in turn i As the learning rate of the GRU model, divide the historical encrypted traffic multi-modal fusion features and historical encrypted traffic threat labels into X folds. For each fold x, {x|1 ≤ x ≤ X}; use all historical encrypted traffic multi-modal fusion features and historical encrypted traffic threat label data except the x-th fold to train the GRU model, and use the training data of the x-th fold to evaluate the model to obtain the performance metric set y = {y1, y2, …, y x , …, y X}, calculate the mean of the performance metrics for all folds to obtain the average performance F. The calculation formula is as follows ; If the average performance F > the best performance U, update the best performance to H and update the optimal learning rate to u i ; otherwise, keep the original best performance and optimal learning rate; S333. Repeat S332. When the maximum random search number is reached, stop the iteration to obtain the optimal learning rate; S4. Collect real-time encrypted traffic, extract the statistical features and content features of the real-time encrypted traffic, and input them into the multi-modal encrypted traffic feature fusion model to obtain the real-time encrypted traffic multi-modal fusion features and input them into the GRU encrypted traffic threat monitoring model to obtain the threat detection result of the real-time encrypted traffic; The S4 includes the following steps: S41. Collect real-time encrypted traffic and extract the statistical features and content features of the real-time encrypted traffic; S42. Input the statistical features and content features of the real-time encrypted traffic into the multi-modal encrypted traffic feature fusion model to obtain the real-time encrypted traffic multi-modal fusion features; S43. Input the real-time encrypted traffic multi-modal fusion features into the GRU encrypted traffic threat detection model to obtain the encrypted traffic threat detection result.
[0027] Embodiment 2 Please refer to Figure 2 , a multi-modal feature fusion-based encrypted traffic threat detection system for implementing the above multi-modal feature fusion-based encrypted traffic threat detection method, including a multi-modal feature extraction module, a multi-modal feature fusion optimization module, a GRU threat detection model construction module, and a real-time encrypted traffic threat detection module; The multi-modal feature extraction module is used to capture encrypted traffic through a mirror port, desensitize sensitive fields by hashing, and extract statistical features and content features to obtain a multi-modal feature training set; The multimodal feature fusion optimization module is used to construct a fusion model including a statistical feature encoder, a content feature encoder, and a cross-modal projection layer, and adopt data augmentation technology for the multimodal feature training set to improve the model robustness. By defining a modal alignment loss function and a cross-modal contrast loss function, and combining the emperor penguin optimization algorithm to globally search for the optimal weight matrix and bias parameters, a multimodal encrypted traffic feature fusion model is obtained; The GRU threat detection model construction module is used to adopt a GRU network and utilize its gating mechanism to capture the long-term time series dependence relationship of encrypted traffic; the model is trained through the multimodal encrypted traffic fusion features of historical encrypted traffic and their threat labels, and the learning rate parameters are optimized by combining grid search and cross-validation to obtain a GRU encrypted traffic threat detection model; The real-time encrypted traffic threat detection module is used to extract its statistical features and content features after real-time collecting encrypted traffic, and input them into the multimodal encrypted traffic feature fusion model to generate a fusion feature vector; the fusion feature vector is input into the GRU encrypted traffic threat detection model for dynamic time series analysis, and the threat determination result of the real-time traffic is output.
[0028] Embodiment III An encrypted traffic threat detection device based on multimodal feature fusion stores a program, and when the program is executed by a processor, it is used to implement the above-mentioned encrypted traffic threat detection method based on multimodal feature fusion.
[0029] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0030] The above-disclosed preferred embodiments of the invention are only used to help explain the invention. The preferred embodiments do not elaborate all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the invention, so that those skilled in the relevant technical field can understand and utilize the invention well.
Claims
1. A method for detecting encrypted traffic threats based on multi-modal feature fusion, characterized in that, Including the following steps: S1. Extract the statistical features and content features of encrypted traffic to obtain a multi-modal feature training set; S2. Construct a multi-modal feature fusion model, combine the multi-modal feature training set and an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model, and obtain the optimal solution; use the optimal solution as the weight matrix and bias of the multi-modal feature fusion model to obtain a multi-modal encrypted traffic feature fusion model; The S2 includes the following steps: S21. Construct a multi-modal feature fusion model and set the weight matrix and bias of the multi-modal feature fusion model; the multi-modal feature fusion model includes a statistical feature encoder, a content feature encoder, and a cross-modal projection layer; S22. Add Gaussian noise to the time intervals in the statistical features in the training set and randomly discard some features to obtain enhanced statistical features; perform random masking and block permutation operations on the byte sequences in the content features in the training set to obtain enhanced content features; S23. Through the statistical feature encoder, content feature encoder, and cross-modal projection layer, vectorize the enhanced statistical features, enhanced content features, positive sample pairs, and negative sample pairs to obtain enhanced statistical feature vectors, enhanced content feature vectors, positive sample pair vectors, and negative sample pair vectors; S24. Match the enhanced statistical features and enhanced content features of the same session to obtain positive sample pairs; Randomly match the enhanced statistical features and enhanced content features of different sessions in the training set to obtain negative sample pairs; S25. Through the modal alignment loss function, calculate the difference between the enhanced statistical feature vector and the enhanced content feature vector to obtain the modal alignment difference; Through the cross-modal contrast loss function, calculate the ratio of the similarity between positive sample pairs to the similarity between positive sample pairs and all negative sample pairs to obtain the cross-modal contrast difference; Use the optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model to minimize the modal alignment difference and the cross-modal contrast difference, and obtain the optimal solution; use the optimal solution as the weight matrix and bias of the multi-modal feature fusion model to obtain an improved multi-modal feature fusion model; Set the feature fusion formula of the improved multi-modal feature fusion model; S3. Construct a GRU model, collect historical encrypted traffic and historical encrypted traffic threat labels; extract the statistical features and content features of historical encrypted traffic and input them into the multi-modal encrypted traffic feature fusion model to obtain historical encrypted traffic multi-modal fusion features; use the historical encrypted traffic multi-modal fusion features and historical encrypted traffic threat labels to train the GRU model to obtain a GRU encrypted traffic threat detection model; S4. Collect real-time encrypted traffic, extract the statistical features and content features of real-time encrypted traffic, and input them into the multi-modal encrypted traffic feature fusion model to obtain real-time encrypted traffic multi-modal fusion features and input them into the GRU encrypted traffic threat monitoring model to obtain the threat detection result of real-time encrypted traffic.
2. The encryption traffic threat detection method based on multi-modal feature fusion according to claim 1, wherein, The S1 includes the following steps: S,11. Capture encrypted traffic through a mirror port to obtain the original encrypted traffic; perform hash desensitization on sensitive fields in the original encrypted traffic to obtain encrypted traffic; S12. Extract the statistical features and content features of the encrypted traffic to obtain a multi-modal feature training set of the encrypted traffic.
3. The method for detecting encrypted traffic threats based on multi-modal feature fusion according to claim 2, wherein The S12 includes the following steps: S121. Extract the number of data packets, average packet length, traffic entropy, variance of time interval, and protocol type of the encrypted traffic to obtain session-level features; extract the session frequency of the same source IP of the encrypted traffic, destination port distribution, and traffic burstiness to obtain host-level features; the session-level features and host-level features together constitute the statistical features of the encrypted traffic; S122. Intercept the first N bytes of the encrypted traffic, pad with zeros for the insufficient part, and normalize the byte sequence of the encrypted traffic to obtain the content features of the encrypted traffic; S123. Correlate and store the statistical features and content features to obtain a multi-modal training set.
4. A method for detecting encrypted traffic threats based on multi-modal feature fusion according to claim 1, characterized in that, The step of using an optimization algorithm to find the weight matrix and bias of the multi-modal feature fusion model to obtain the optimal solution in S25 includes the following steps: S251. Construct a population of emperor penguins, set the size of the population of emperor penguins, and set the maximum number of optimization iterations; S252. Set the initial positions of the population of emperor penguins according to the weight matrix and bias to obtain an initial position set of the population of emperor penguins; S253. Define a fitness function according to the modal alignment difference and cross-modal contrast difference; S254. Perform iterative operations on the initial position set of the population of emperor penguins, calculate the fitness value of each position in the initial position set of the population of emperor penguins, update the positions of each emperor penguin in the initial position set of the population of emperor penguins, and obtain the best individual position of the emperor penguin in the population of emperor penguins and the global best position of the emperor penguin in each round of iteration; the higher the fitness value, the better the position; according to the fitness function in each round of iteration; S255. Repeat S254, and stop the iteration when the maximum number of test iterations is reached to obtain the optimal solution.
5. A method for detecting encrypted traffic threats based on multi-modal feature fusion according to claim 1, characterized in that, The S3 includes the following steps: S31. Construct a GRU model and set the initial learning rate of the GRU; S32. Collect historical encrypted traffic, extract the statistical features and content features of the historical encrypted traffic, set normal labels and abnormal labels for the historical encrypted traffic data to obtain historical encrypted traffic threat labels; input the statistical features and content features of the historical encrypted traffic into the improved multi-modal feature fusion model to obtain historical encrypted traffic multi-modal fusion sample features; S33. Collect historical encrypted traffic, extract the statistical features and content features of the historical encrypted traffic, set normal labels and abnormal labels for the historical encrypted traffic data to obtain historical encrypted traffic threat labels; input the statistical features and content features of the historical encrypted traffic into the improved multi-modal feature fusion model to obtain historical encrypted traffic multi-modal fusion features; S33. Use the historical encrypted traffic multi-modal fusion features and historical encrypted traffic threat labels to train the GRU model. During the training process, combine grid search with cross-validation algorithm to find the learning rate of the GRU model to obtain the optimal learning rate, and use the optimal learning rate as the learning rate of the GRU model to obtain a GRU encrypted traffic threat detection model.
6. A method for detecting encrypted traffic threats based on multi-modal feature fusion according to claim 5, characterized in that, During the training process in S33, the learning rate of the GRU model is found by combining grid search with the cross-validation algorithm to obtain the optimal learning rate, including the following steps: S331. Set the parameter search space, the maximum number of random searches, and the number of cross-validation folds as X for the grid search combined with the cross-validation algorithm; S332. Randomly set the initial learning rate set; The best performance for initializing the GRU model is U; the learning rate u is sequentially selected from the set of initial learning rates i As the learning rate of the GRU model, the historical encrypted traffic multi-modal fusion features and the historical encrypted traffic threat labels are divided into X folds. For each fold x, {x|1 ≤ x ≤ X}; the GRU model is trained using all the historical encrypted traffic multi-modal fusion features and historical encrypted traffic threat label data except for the x-th fold, and the x-th fold of the training data is used to evaluate the model to obtain a set of performance metrics. The mean of the performance metrics for all folds is calculated to obtain the average performance F; If the average performance F > the best performance U, then update the best performance to K and update the optimal learning rate to u i ; otherwise, keep the original best performance and optimal learning rate; S333. Repeat S332. When the maximum number of random searches is reached, stop the iteration to obtain the optimal learning rate.
7. A method for detecting encrypted traffic threats based on multi-modal feature fusion according to claim 1, characterized in that S4 includes the following steps: S41. Collect real-time encrypted traffic and extract the statistical features and content features of the real-time encrypted traffic; S42. Input the statistical features and content features of the real-time encrypted traffic into the multi-modal encrypted traffic feature fusion model to obtain the multi-modal fusion features of the real-time encrypted traffic; S43. Input the multi-modal fusion features of the real-time encrypted traffic into the GRU encrypted traffic threat detection model to obtain the encrypted traffic threat detection result.
8. An encrypted traffic threat detection system based on multi-modal feature fusion is used to implement an encrypted traffic threat detection method according to any one of claims 1-7; the system includes a multi-modal feature extraction module, a multi-modal feature fusion and optimization module, a GRU threat detection model construction module, and a real-time encrypted traffic threat detection module.
9. An encrypted traffic threat detection device based on multimodal feature fusion, characterized in that, A program is stored thereon, and when the program is executed by a processor, it implements an encrypted traffic threat detection method according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-modal encrypted traffic classification method and system based on cross-modal attention
CN117521015A
Network security detection method and system based on deep learning
CN119011196A
Space competition situation threat level evaluation method and system based on credibility weighting
CN119538012A
Malicious encrypted traffic detection method and system based on multi-modal feature fusion
CN119966688A
Network threat multi-modal detection method based on large model
CN120185905A