Malicious sample automatic analysis method and system based on deep learning
By constructing network traffic correlation graphs and multi-scale statistical features, combined with interactive gating vectors and protocol deviation terms, the problem of low malicious sample identification accuracy in traditional methods is solved, and efficient malicious sample identification is achieved in complex network environments.
Patent Information
- Application Number
- CN202511300102.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In existing automatic analysis methods for malicious samples, traditional feature extraction only focuses on the independent features of a single sample, ignoring the correlation between samples, time decay characteristics, protocol differences, and multi-scale statistical laws. This results in low malicious sample identification accuracy and high misjudgment rate, making it difficult to accurately identify malicious samples in complex network environments.
By constructing a network traffic correlation graph, introducing a time decay factor, generating protocol-aware attention weights and neighborhood aggregation features, extracting multi-scale statistical features, combining interaction gating vectors and protocol deviation terms, generating encoding vectors, and performing cross-scale interaction feature reconstruction and error fusion, the recognition accuracy and reliability of malicious samples are improved.
It accurately captures the clustering and dynamics of samples, reduces the misjudgment rate, improves the precision and accuracy of identifying malicious samples, and can effectively identify malicious samples in complex network environments.
Smart Images

Figure CN120785668A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of sample analysis, and specifically refers to a malicious sample automatic analysis method and system based on deep learning. BACKGROUND
[0002] The malicious sample automatic analysis method is a kind of method based on deep learning technology, which mines the potential law of normal and malicious samples in historical network traffic sample data, accurately identifies the feature mode of malicious samples, and automatically analyzes whether the real-time network traffic sample is malicious, thereby providing intelligent and efficient detection means for network security protection, helping to intercept malicious attacks in time, and guaranteeing the safety and stability of network environment. However, in the existing malicious sample automatic analysis method, the traditional feature extraction only focuses on the independent features of a single sample, ignores the correlation between samples, time decay characteristics, protocol differences and multi-scale statistical rules, resulting in low recognition accuracy and high misjudgment rate of malicious samples; in the existing malicious sample automatic analysis method, the sample feature mining and utilization are insufficient, the relationship between samples, protocol characteristics and multi-time scale rules are not effectively associated, the representation and judgment ability of malicious sample features is insufficient, it is difficult to accurately identify malicious samples in complex network environment, and it is easy to appear misjudgment, which cannot reliably support the automatic analysis demand. SUMMARY
[0003] In view of the above, in order to overcome the defects of the prior art, the present application provides a malicious sample automatic analysis method and system based on deep learning, which is aimed at the problem that in the existing malicious sample automatic analysis method, the traditional feature extraction only focuses on the independent features of a single sample, ignores the correlation between samples, time decay characteristics, protocol differences and multi-scale statistical rules, resulting in low malicious sample recognition accuracy and high misjudgment rate. The present scheme calculates the edge weight through source IP, destination IP and protocol-port similarity, constructs a network traffic correlation graph, introduces a time decay factor to obtain dynamic neighbor aggregation features, converts protocol types into embedded vectors to generate protocol-aware attention weights, and obtains protocol-aware enhanced features through weighted splicing. Based on protocol neighbor attention, neighborhood aggregation features are obtained, and protocol-aware gating features are obtained according to the gating vector. Statistical feature vectors of three scales of short-term, medium-term and long-term are extracted, which provides a rich and high-discriminatory feature basis for accurately identifying complex and unknown malicious samples, and improves the recognition accuracy of malicious samples. In view of the problem that in the existing malicious sample automatic analysis method, the sample feature mining and utilization are insufficient, the relationship between samples, protocol characteristics and multi-time scale rules are not effectively associated, the representation and judgment ability of malicious sample features is insufficient, and it is difficult to accurately identify malicious samples in complex network environments, and it is easy to miss and misjudge, which cannot reliably support the automatic analysis demand, the present scheme obtains cross-scale interaction features based on interactive gating vectors, introduces protocol bias terms to calculate dynamic attention weights, generates encoding vectors, and four parallel branches respectively reconstruct short-term, medium-term and long-term statistical features and protocol-aware gating features. The weighted sum of the branch weights is reconstructed to obtain the error, the malicious probability is predicted based on the encoding vector, the total reconstruction error and the malicious probability are fused to obtain the comprehensive malicious score, the malicious sample is judged, the malicious sample in the complex network environment is accurately identified, and the accuracy and reliability of the automatic analysis of the malicious sample are improved.
[0004] The technical solutions adopted by the present application are as follows: The present application provides a malicious sample automatic analysis method based on deep learning, which comprises the following steps:
[0005] Step S1: sample data integration;
[0006] Step S2: sample feature extraction;
[0007] Step S3: constructing a malicious sample automatic analysis model;
[0008] Step S4: intelligent analysis.
[0009] Further, in step S1, the sample data integration is to collect historical network traffic sample data; the historical network traffic sample data includes the start time, duration, protocol type, source IP address, source port, destination IP address, destination port, TCP flag, packet number, byte number, packet length standard deviation, average packet length, minimum packet length, maximum packet length, arrival time interval standard deviation and sample type of network traffic session; the sample type includes normal sample and malicious sample; the historical network traffic sample data is preprocessed; the sample type is taken as a data label, and a network traffic sample data set with data labels is constructed.
[0010] Further, in step S2, the sample feature extraction specifically includes the following steps:
[0011] Step S21: network traffic association graph construction; based on the network traffic sample data set, each sample is regarded as a node in the graph, the edge weight between the sample nodes is calculated through multi-dimensional association feature fusion, and an adjacency matrix is constructed according to the association strength threshold, to obtain a network traffic association graph;
[0012] Step S22: dynamic neighbor feature aggregation; each sample node in the network traffic association graph is traversed, the start time difference between it and all neighbor sample nodes is calculated to obtain a time decay factor, the adjacency relationship and the time decay factor are fused, the features of the neighbor sample nodes are weighted and aggregated to obtain dynamic neighbor aggregation features;
[0013] Step S23: protocol-aware enhancement; the one-hot vector of the protocol type is converted into a protocol type embedding vector through an embedding matrix, a multi-layer perception machine is used to generate protocol-aware attention weights matching the original feature dimension, the original features are weighted element by element, and the weighted features are spliced with the dynamic neighbor aggregation features to obtain protocol-aware enhanced features;
[0014] Step S24: gated neighborhood enhancement; for the sample nodes in the network traffic association graph, the protocol type embedding vectors of the sample nodes themselves and their neighbor sample nodes are combined with the protocol-aware enhanced features to generate protocol neighbor attention, the protocol-aware enhanced features of the neighbor sample nodes are weighted and aggregated to obtain neighborhood aggregation features, the neighborhood aggregation features are spliced with the original features of the sample nodes, and then a linear transformation is performed to generate fusion intermediate features, the fusion intermediate features are linearly transformed and activated by Sigmoid to obtain a gating vector, which is multiplied element by element with the fusion intermediate features to obtain protocol-aware gated features;
[0015] Step S25: multi-scale statistical features; taking the start time of each network traffic sample data as a reference, sliding time windows are divided into three scales of short-term, medium-term and long-term, and statistical features are extracted from all samples contained in each window to obtain a statistical feature vector of each sample in three scales.
[0016] Further, in step S3, the construction of the malicious sample automatic analysis model is based on the protocol-aware gating feature and the three-scale statistical feature vector, and the construction of the malicious sample automatic analysis model is completed according to the neural network; specifically including the following steps:
[0017] Step S31: cross-protocol dynamic encoder; including the following steps:
[0018] Step S311: cross-scale feature interaction; linearly transform the statistical feature vector of each scale, calculate the forward interaction gating vector and the reverse interaction gating vector between the short-term and the medium-term, and the medium-term and the long-term scale, and obtain the cross-scale interaction feature based on the interaction gating vector;
[0019] Step S312: protocol-aware dynamic encoding; splice the cross-scale interaction feature, the protocol-aware gating feature and the protocol type embedding vector to obtain the unified fusion feature, linearly transform the unified fusion feature with 3 independent weight matrices to obtain the query vector, the key vector and the value vector of the attention mechanism, introduce the protocol bias term, calculate the dynamic attention weight, and weight sum the value vector to obtain the final encoding vector;
[0020] Step S32: multi-branch reconstruction decoder; including the following steps:
[0021] Step S321: multi-branch feature reconstruction; the multi-branch reconstruction decoder contains 4 parallel branches, which correspond to the reconstruction of the short-term, medium-term, long-term and protocol-aware gating feature based on the encoding vector, and each branch adopts the structure combined by the deconvolution layer, the batch normalization layer and the Sigmoid activation function;
[0022] Step S322: multi-branch reconstruction error; the reconstruction error of each branch is calculated respectively, and the weighted sum of the reconstruction errors is obtained according to the branch weight to obtain the total reconstruction error;
[0023] Step S33: malicious sample judgment; based on the encoding vector to predict the malicious probability, half of the sum of the total reconstruction error and the malicious probability after standardization is taken as the comprehensive malicious score, a score threshold is set, if the comprehensive malicious score is greater than or equal to the score threshold, the sample is determined as a malicious sample, otherwise the sample is determined as a normal sample, and the data label is output.
[0024] Further, in step S4, the intelligent analysis is to collect real-time network traffic sample data, after preprocessing the real-time network traffic sample data, the protocol-aware gating feature and the three-scale statistical feature vector are extracted therefrom, input into the malicious sample automatic analysis model for processing, and according to the output data label, the sample type corresponding to the real-time network traffic sample data is obtained.
[0025] The application provides a deep learning-based malicious sample automatic analysis system, which comprises a sample data integration module, a sample feature extraction module, a malicious sample automatic analysis model construction module and an intelligent analysis module.
[0026] The sample data integration module collects historical network flow sample data, pre-processes the historical network flow sample data, constructs a network flow sample data set, and sends the data to the sample feature extraction module.
[0027] The sample feature extraction module receives the data sent by the sample data integration module, calculates edge weights, constructs a network flow correlation graph, introduces a time decay factor to obtain dynamic neighbor aggregation features, converts protocol types into embedding vectors to generate protocol-aware attention weights, performs weighted splicing to obtain protocol-aware enhanced features, obtains neighborhood aggregation features based on protocol neighbor attention, obtains protocol-aware gating features according to a gating vector, extracts statistical feature vectors of three scales, and sends the data to the malicious sample automatic analysis model construction module.
[0028] The malicious sample automatic analysis model construction module receives the data sent by the sample feature extraction module, obtains cross-scale interaction features based on an interaction gating vector, introduces a protocol bias term to calculate dynamic attention weights, generates an encoding vector, respectively reconstructs short-term, medium-term and long-term statistical features and protocol-aware gating features, performs weighted summation reconstruction error according to branch weights, predicts a malicious probability based on the encoding vector, obtains a comprehensive malicious score, judges a malicious sample, and sends the data to the intelligent analysis module.
[0029] The intelligent analysis module receives the data sent by the malicious sample automatic analysis model construction module, obtains a sample type corresponding to real-time network flow sample data based on the output of the malicious sample automatic analysis model.
[0030] The application has the following beneficial effects:
[0031] (1) In view of the problem that in the existing automatic analysis method of malicious samples, the traditional feature extraction only focuses on the independent features of a single sample, ignores the correlation between samples, time decay characteristics, protocol differences and multi-scale statistical rules, resulting in low recognition accuracy and high misjudgment rate of malicious samples, the scheme calculates the edge weight through source IP, destination IP and protocol-port similarity, constructs a network flow correlation graph, introduces a time decay factor to obtain dynamic neighbor aggregation features, accurately captures the clustering and dynamics of samples, reduces the missed judgment caused by isolated sample analysis, and improves the recognition accuracy of malicious patterns; the protocol type is converted into an embedded vector to generate protocol-aware attention weight, and the protocol-aware enhanced features are obtained by weighted splicing, so that the feature processing is more targeted, which can effectively distinguish malicious behaviors under different protocols and reduce misjudgment caused by protocol differences; based on the protocol neighbor attention, the neighborhood aggregation features are obtained, the protocol-aware gating features are obtained according to the gating vector, the irrelevant neighbors and redundant information are filtered, the feature purity is improved, and high-quality input is provided for subsequent model construction; the statistical feature vectors of three scales of short-term, medium-term and long-term are extracted, which provides a rich and high-discriminative feature basis for accurately identifying complex and unknown malicious samples, and improves the recognition accuracy of malicious samples.
[0032] (2) In view of the problem that in the existing automatic analysis method of malicious samples, the sample feature mining and utilization are insufficient, the relationship between samples, protocol characteristics and multi-time scale rules are not effectively associated, the representation and judgment ability of malicious sample features is insufficient, it is difficult to accurately identify malicious samples in complex network environment, and misjudgment and misjudgment are prone to occur, which cannot reliably support the automatic analysis demand, the scheme obtains cross-scale interaction features based on the interaction gating vector, dynamically integrates complementary information of different scales, and solves the one-sidedness of single-scale analysis; the dynamic attention weight is calculated by introducing a protocol bias term to generate an encoding vector, which ensures that the encoding vector preferentially retains the core malicious information related to the protocol and more comprehensively reflects the multi-dimensional features of complex malicious samples; four parallel branches respectively reconstruct short-term, medium-term and long-term statistical features and protocol-aware gating features, ensure that the encoding vector does not lose key malicious information, more sensitively capture unknown malicious patterns, and reduce missed judgment; the errors are reconstructed by weighted summation according to branch weights to highlight the contribution of high-value errors to malicious judgment; the malicious probability is predicted based on the encoding vector, the total reconstruction error and the malicious probability are fused to obtain a comprehensive malicious score, and the malicious sample is judged to accurately identify the malicious sample in the complex network environment and improve the accuracy and reliability of the automatic analysis of malicious samples. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 A flowchart of the automatic analysis method of malicious samples based on deep learning provided by the application is shown in the figure.
[0034] Figure 2 A schematic diagram of the automatic analysis system of malicious samples based on deep learning provided by the application is shown in the figure.
[0035] Figure 3 is a flowchart of step S2;
[0036] Figure 4 is a flowchart of step S3.
[0037] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of this specification, illustrate embodiments of the application, and are included to further explain the application, and are not intended to limit the application. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the application.
[0039] In the description of the application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the application.
[0040] Embodiment one, refer to Figure 1 The application provides a malicious sample automatic analysis method based on deep learning, which comprises the following steps:
[0041] Step S1: sample data integration; collecting historical network traffic sample data, pre-processing the historical network traffic sample data, and constructing a network traffic sample data set;
[0042] Step S2: sample feature extraction; calculating edge weight, constructing a network traffic correlation graph, introducing a time decay factor, obtaining dynamic neighbor aggregation features, converting protocol types into embedded vectors, generating protocol-aware attention weights, weighting and splicing to obtain protocol-aware enhanced features, obtaining neighborhood aggregation features based on protocol neighbor attention, obtaining protocol-aware gating features according to gating vectors, and extracting three-scale statistical feature vectors;
[0043] Step S3: Constructing the automatic analysis model of malicious samples; obtaining cross-scale interaction features based on the interaction gate vector, introducing a protocol bias term to calculate dynamic attention weights, generating an encoding vector, respectively reconstructing short-term, medium-term and long-term statistical features and protocol-aware gate features, weighting and summing the errors according to branch weights, predicting the malicious probability based on the encoding vector, obtaining the comprehensive malicious score, and judging the malicious samples;
[0044] Step S4: Intelligent analysis; based on the output of the automatic analysis model of malicious samples, obtaining the sample type corresponding to the real-time network traffic sample data.
[0045] Embodiment two, refer to Figure 1 This embodiment is based on the above embodiment, in step S1, the sample data integration is to collect historical network traffic sample data, pre-process the historical network traffic sample data, take the sample type as the data label, and construct a network traffic sample data set with data label;
[0046] The historical network traffic sample data includes the start time, duration, protocol type, source IP address, source port, destination IP address, destination port, TCP flag, packet number, byte number, packet length standard deviation, average packet length, minimum packet length, maximum packet length, arrival time interval standard deviation and sample type of network traffic session;
[0047] The sample type includes normal samples and malicious samples;
[0048] The pre-processing includes missing value processing, encoding processing and standardization processing;
[0049] The missing value processing is to fill the missing values with the median of the same protocol type sample;
[0050] The encoding processing is to convert the category type data into numerical type data using one-hot encoding;
[0051] The standardization processing is to unify the numerical type data to the range of [0, 1] using the max-min scaling method.
[0052] Embodiment three, refer to Figure 1 and Figure 3 This embodiment is based on the above embodiment, in step S2, the sample feature extraction specifically includes the following contents:
[0053] Step S21: Network traffic association graph construction; traditional feature extraction only focuses on the independent features of a single sample, ignoring the association relationship between samples. Malicious samples are often not isolated. If only a single sample is analyzed, it is easy to produce a false judgment. Discrete network traffic samples are converted into structured association graphs. The association strength between samples is quantified through edge weights. The sample cluster with association is selected; based on the network traffic sample data set, each sample is regarded as a node in the graph. The edge weight between the sample nodes is calculated through multi-dimensional association feature fusion. According to the association strength threshold , the adjacency matrix is constructed. When the edge weight is greater than or equal to the association strength threshold, the corresponding value of the adjacency matrix is 1, otherwise 0, and the network traffic association graph is obtained; the formula used is as follows:
[0054]
[0055]
[0056] In the formula, and are the edge weight between the i-th sample and the j-th sample and the protocol-port similarity, respectively, sIP i and sIP j are the source IP addresses of the i-th sample and the j-th sample, respectively, dIP i and dIP j are the destination IP addresses of the i-th sample and the j-th sample, respectively, p i and p j are the protocol types of the i-th sample and the j-th sample, respectively, dP i and dP j are the destination ports of the i-th sample and the j-th sample, respectively, is a logical and operator; is an indicator function. If the condition in the parentheses is true, then , otherwise ;
[0057] Step S22: Dynamic neighbor feature aggregation; network traffic has a time sequence, and the association strength between samples will decay over time. However, the association graph does not consider the time factor, which may include unrelated samples across time windows as neighbors, resulting in aggregated features containing noise, reducing the purity of malicious patterns, and introducing a time decay factor to dynamically adjust the weight of neighbor samples. Only the neighbor features close in time to the current sample are aggregated to improve the timeliness and relevance of the aggregated features; each sample node in the network traffic association graph is traversed, and the start time difference between it and all neighbor sample nodes is calculated to obtain the time decay factor. The adjacency relationship and the time decay factor are fused, and the features of the neighbor sample nodes are weighted and aggregated to obtain the dynamic neighbor aggregation features; the formula used is as follows:
[0058] ;
[0059] ;
[0060] wherein, and are the start time of the sample node i and its neighbor sample node j, respectively, is the time decay coefficient, , is the element in the adjacency matrix A, is the time decay factor of the sample node i and its neighbor sample node j, f j is the original feature of the sample node j, and M is the number of samples in the network traffic sample dataset, is the smoothing term, , is the dynamic neighbor aggregated feature of the sample node i;
[0061] Step S23: Protocol-aware enhancement; the network traffic feature laws of different protocols are significantly different, and the manifestation of malicious behavior is also strongly related to the protocol. However, the traditional feature extraction adopts a unified feature processing method for all protocols, which cannot highlight the protocol-specific malicious patterns. The protocol type is converted into an embedding vector, and protocol-specific attention weights are generated to strengthen the dimensions of the original features related to the malicious patterns of the protocol. At the same time, the dynamic neighbor aggregated features are fused to form protocol-aware enhanced features. The one-hot vector of the protocol type is converted into a protocol type embedding vector through an embedding matrix, and the protocol-aware attention weights matching the dimensions of the original features are generated through a multi-layer perception machine. The original features are weighted element by element, and the weighted features are spliced with the dynamic neighbor aggregated features to obtain protocol-aware enhanced features. The formula used is as follows:
[0062] ;
[0063] ;
[0064] wherein, e i is the protocol type embedding vector of the sample node i, h i is the one-hot vector of the protocol type of the sample node i, W t is the embedding matrix, f i is the original feature of the sample node i, is the Sigmoid activation function, is the multi-layer perception machine, is the feature splicing operation, is the element-wise multiplication, is the protocol-aware enhanced feature of the sample node i;
[0065] Step S24: Gated neighborhood enhancement; neighbor aggregation does not consider the protocol relevance and feature importance of neighbor samples, quantifies the importance of neighbor samples to the current sample based on protocol embedding and protocol-aware enhanced features, filters irrelevant neighbors, combines neighborhood aggregation features with original features, retains single-sample details and neighbor cluster information, dynamically selects effective information in the intermediate feature based on the gating vector, and suppresses noise; for sample nodes in the network traffic association graph, combine their own and their neighbor sample nodes' protocol type embedding vectors and protocol-aware enhanced features to generate protocol neighbor attention, aggregate the protocol-aware enhanced features of neighbor sample nodes by weighting to obtain neighborhood aggregation features, concatenate the neighborhood aggregation features with the original features of the sample nodes, and then generate the fusion intermediate features through linear transformation, perform linear transformation on the fusion intermediate features and activate them through Sigmoid to obtain the gating vector and multiply it with the fusion intermediate features element by element to obtain the protocol-aware gating features; the formulas used are as follows:
[0066] ;
[0067] ;
[0068] ;
[0069] ;
[0070] In the formula, is the protocol neighbor attention of sample node i to neighbor sample node j, e j and e k are the protocol type embedding vectors of sample node j and sample node k, respectively, and are the protocol-aware enhanced features of sample node j and sample node k, respectively, T is the transpose operation, a is the attention parameter vector, is the neighbor set of sample node i, , and are the neighborhood aggregation features, fusion intermediate features and protocol-aware gating features of sample node i, respectively, W agg and b agg are the aggregation weight matrix and aggregation bias term, respectively, W middle and b middle are the fusion weight matrix and fusion bias term, respectively, W g and b g are the gating weight matrix and gating bias term, respectively, is the leaky linear rectifier function, is the rectified linear unit activation function;
[0071] Step S25: Multi-scale statistical features; malicious samples have large differences in time span, and may only present local patterns in single-sample features, which are difficult to capture through a single time scale. Single-sample features cannot reflect the statistical rules of traffic. According to short-term, medium-term and long-term three time windows, the statistical features of each sample in the window are extracted to supplement the global rules of the time dimension missing in single-sample features. Taking the start time of each network traffic sample data as the benchmark, the sliding time window is divided according to the short-term H S , medium-term H M and long-term H L three scales respectively. In each window, 8-dimensional statistical features are extracted for all samples contained, obtaining an 8-dimensional statistical feature vector of each sample in three scales. The statistical features include packet number mean, packet number standard deviation, byte number mean, byte number standard deviation, duration mean, maximum packet length, sample number and arrival time interval standard deviation mean. The formulas used are as follows:
[0072] ;
[0073] where H S , H M and H L are the sizes of the short-term, medium-term and long-term sliding time windows respectively, H S = 60 seconds, H M = 300 seconds, H L = 600 seconds, is the 8-dimensional statistical feature vector of the i-th sample in the window with size H S , , , , , , , and are the packet number mean, packet number standard deviation, byte number mean, byte number standard deviation, duration mean, maximum packet length, sample number and arrival time interval standard deviation mean of all samples in the window with size H S starting from the start time of the i-th sample.
[0074] By performing the above operation, for the existing malicious sample automatic analysis method, the traditional feature extraction only focuses on the independent features of a single sample, ignores the correlation between samples, time decay characteristics, protocol differences and multi-scale statistical rules, resulting in low malicious sample recognition accuracy and high misjudgment rate. The scheme calculates the edge weight through the source IP, destination IP and protocol-port similarity, constructs a network flow correlation graph, introduces a time decay factor to obtain dynamic neighbor aggregation features, accurately captures the clustering and dynamics of samples, reduces the missed judgment caused by isolated sample analysis, and improves the malicious pattern recognition accuracy. The protocol type is converted into an embedded vector to generate a protocol-aware attention weight, and the protocol-aware enhanced features are obtained by weighted splicing, so that the feature processing is more targeted, which can effectively distinguish malicious behaviors under different protocols and reduce misjudgment caused by protocol differences. Based on the protocol neighbor attention, the neighborhood aggregation features are obtained, the protocol-aware gating features are obtained according to the gating vector, the irrelevant neighbors and redundant information are filtered, the feature purity is improved, and high-quality input is provided for subsequent model construction. Extract the statistical feature vectors of three scales of short-term, medium-term and long-term to provide a rich and high-discriminative feature basis for accurately identifying complex and unknown malicious samples, and improve the recognition accuracy of malicious samples.
[0075] In an embodiment, referring to Figure 1 and Figure 4 , the embodiment is based on the above embodiment. In step S3, the malicious sample automatic analysis model is constructed based on the protocol-aware gating features and the statistical feature vectors of three scales, and the construction of the malicious sample automatic analysis model is completed according to the neural network; specifically including the following steps:
[0076] Step S31: Cross-protocol dynamic encoder; including the following steps:
[0077] Step S311: Cross-scale feature interaction; there may be complementary information between different scale features, and independent use will lose the associated value, and single scale feature may have noise, which is easy to cause model misjudgment. Through the interaction of the gating vector, the bidirectional information fusion of different scale features is realized, the complementary information is strengthened, and the single scale noise is suppressed. Linearly transform the statistical feature vector of each scale to unify the dimension, calculate the forward interaction gating vector and the reverse interaction gating vector between the short-term and the medium-term, and between the medium-term and the long-term scale, and obtain the cross-scale interaction features based on the interaction gating vector; the formula is as follows:
[0078] ;
[0079] ;
[0080] ;
[0081] ;
[0082] ;
[0083] ;
[0084] ;
[0085] where, , and are the results of linear transformation of the statistical feature vectors of the i-th sample within the windows of size H S , H M and H L , respectively, W SM and b SM are the trainable weight matrix and bias term of short-to-medium scale interaction, respectively, W MS and b MS are the trainable weight matrix and bias term of medium-to-short scale interaction, respectively, W ML and b ML are the trainable weight matrix and bias term of medium-to-long scale interaction, respectively, W LM and b LM are the trainable weight matrix and bias term of long-to-medium scale interaction, respectively, and are the forward interaction gating vectors of short-to-medium and medium-to-long scale, respectively, and are the reverse interaction gating vectors of medium-to-short and long-to-medium scale, respectively, , and are the short, medium and long cross-scale interaction features of the i-th sample, respectively;
[0086] Step S312: protocol-aware dynamic encoding; the traditional encoder adopts a unified attention mechanism for all samples, without considering the differences between protocol types. The importance of malicious features of different protocols is different when encoding. Unified encoding will lead to the averaging of protocol-specific malicious patterns, reducing the discrimination of the encoding vector for malicious samples. A protocol bias term is introduced to dynamically adjust the attention weight, generating a protocol-specific encoding vector to provide high-quality feature representation for subsequent reconstruction and prediction. The cross-scale interaction features, protocol-aware gating features and protocol type embedding vectors are spliced to obtain unified fusion features. Three independent weight matrices are used to linearly transform the unified fusion features to obtain the query vector, key vector and value vector of the attention mechanism. A protocol bias term is introduced to calculate the dynamic attention weight, and the value vector is weighted and summed to obtain the final encoding vector. The formula used is as follows:
[0087] ;
[0088] ;
[0089] ;
[0090] wherein, is the result of linear transformation, is the unified fusion feature of the i-th sample, W e is the weight matrix of the protocol bias term, is the protocol bias term, Q t , K t and V t are the components of the query vector, the key vector and the value vector in the t-th attention head, d q is the dimension of the query vector, is the dynamic attention weight of the i-th sample in the t-th attention head, t max is the number of attention heads;
[0091] Step S32: multi-branch reconstruction decoder; comprising the following steps:
[0092] Step S321: multi-branch feature reconstruction; the traditional single-branch decoder only reconstructs a single feature, and cannot comprehensively verify the preservation degree of the encoding vector for multi-dimensional malicious features. If the encoding vector loses malicious information of a certain dimension, the single-branch reconstruction cannot find it, resulting in weak recognition ability of the model for malicious samples relying on the dimension. By reconstructing the core features through four parallel branches respectively, the integrity of the multi-dimensional information preserved by the encoding vector is verified; the multi-branch reconstruction decoder includes four parallel branches, which correspond to the reconstruction of the short-term, medium-term, long-term and protocol perception gating features based on the encoding vector. Each branch adopts the structure combined by the deconvolution layer, the batch normalization layer and the Sigmoid activation function. Each branch shares the encoding vector, but the weight parameters and the bias parameters of the deconvolution layer and the scaling parameters and the translation parameters of the batch normalization layer of each branch are independent of each other. The used formula is as follows:
[0093] ;
[0094] wherein, is the reconstructed feature of the i-th sample in the short-term branch, and are the deconvolution layer and the batch normalization layer of the short-term branch, respectively, z i is the final encoding vector of the i-th sample;
[0095] Step S322: multi-branch reconstruction error; the reconstruction error of different branches has different indicative significance for malicious samples, the reconstruction error of long-term statistical characteristics has higher value for malicious judgment than that of short-term statistical characteristics, if all errors are equally weighted, the contribution of high-value errors will be diluted, and the recognition sensitivity of malicious samples will be reduced, the reconstruction error of each branch is calculated and weighted sum is calculated according to the weight, so as to highlight the contribution of high-value errors to malicious judgment; the reconstruction error of each branch is calculated respectively, the reconstruction error is weighted and summed according to the weight of each branch, and the total reconstruction error is obtained; the formula used is as follows:
[0096] ;
[0097] ;
[0098] In the formula, , , and are the reconstruction errors of the four branches, ω S , ω M , ω L and ω F are the weights of the four branches, ω S =0.35, ω M =0.35, ω L =0.1, ω F =0.2, is the square of L2 norm, Err i is the total reconstruction error of the i-th sample;
[0099] Step S33: malicious sample judgment; in automatic analysis of malicious samples, if only the malicious probability is relied on, it is easy to be disturbed by normal samples with malicious similar characteristics, if only the reconstruction error is relied on, the normal samples may produce high error due to characteristic fluctuation, leading to misjudgment of normal samples as malicious, through fusion of malicious probability prediction and total reconstruction error standardization, the advantages of the two kinds of judgment basis are combined to form a quantitative and interpretable judgment standard, so as to realize accurate classification of malicious samples; the malicious probability is predicted based on the encoding vector, half of the sum of the total reconstruction error and the normalized malicious probability is taken as the comprehensive malicious score, the score threshold is set, if the comprehensive malicious score is greater than or equal to the score threshold, the sample is judged as a malicious sample, otherwise the sample is judged as a normal sample, and the data label is output; the formula used is as follows:
[0100] ;
[0101] ;
[0102] Wherein, is the malicious probability of the i-th sample, and respectively, the first full connection layer and the second full connection layer, U is a comprehensive malicious score, and respectively, Err i and standardized results.
[0103] By performing the above operations, in view of the problems that in existing automatic analysis methods of malicious samples, sample feature mining and utilization are insufficient, relationships between samples, protocol characteristics and multi-time scale rules are not effectively associated, the representation and judgment ability of malicious sample features is insufficient, it is difficult to accurately identify malicious samples in a complex network environment, false negatives and false positives are prone to occur, and automatic analysis requirements cannot be reliably supported, the scheme obtains cross-scale interaction features based on an interaction gating vector, dynamically integrates complementary information of different scales, solves the one-sidedness of single-scale analysis, introduces a protocol bias term to calculate dynamic attention weights, generates an encoding vector, ensures that the encoding vector preferentially retains core malicious information related to the protocol, more comprehensively reflects the multi-dimensional features of complex malicious samples, four parallel branches respectively reconstruct short-term, medium-term and long-term statistical features and protocol-aware gating features, ensure that the encoding vector does not lose key malicious information, more sensitively captures unknown malicious patterns, and reduces false negatives, weights the sum of errors according to branch weights to highlight the contribution of high-value errors to malicious judgment, predicts a malicious probability based on the encoding vector, fuses the total reconstruction error and the malicious probability, obtains a comprehensive malicious score, judges malicious samples, accurately identifies malicious samples in a complex network environment, and improves the accuracy and reliability of automatic analysis of malicious samples.
[0104] Embodiment five, refer to Figure 1 This embodiment is based on the above-mentioned embodiments, in step S4, the intelligent analysis is to collect real-time network traffic sample data; the real-time network traffic sample data includes the start time, duration, protocol type, source IP address, source port, destination IP address, destination port, TCP flag, packet number, byte number, packet length standard deviation, average packet length, minimum packet length, maximum packet length and arrival time interval standard deviation of the network traffic session, after preprocessing the real-time network traffic sample data, the protocol-aware gating features and the statistical feature vectors of three scales are extracted therefrom and input into the malicious sample automatic analysis model for processing, according to the output data label, the sample type corresponding to the real-time network traffic sample data is obtained.
[0105] Embodiment six, refer to Figure 2 This embodiment is based on the above-mentioned embodiments, and the malicious sample automatic analysis system based on deep learning provided by the present application comprises a sample data integration module, a sample feature extraction module, a malicious sample automatic analysis model construction module and an intelligent analysis module.
[0106] The sample data integration module collects historical network traffic sample data, pre-processes the historical network traffic sample data, constructs a network traffic sample data set, and sends the data to the sample feature extraction module;
[0107] The sample feature extraction module receives the data sent by the sample data integration module, calculates edge weights, constructs a network traffic correlation graph, introduces a time decay factor to obtain dynamic neighbor aggregation features, converts protocol types into embedding vectors to generate protocol-aware attention weights, and weights and splices to obtain protocol-aware enhanced features, obtains neighborhood aggregation features based on protocol neighbor attention, obtains protocol-aware gating features according to a gating vector, extracts statistical feature vectors of three scales, and sends the data to the malicious sample automatic analysis model construction module;
[0108] The malicious sample automatic analysis model construction module receives the data sent by the sample feature extraction module, obtains cross-scale interaction features based on the interaction gating vector, introduces a protocol bias term to calculate dynamic attention weights, generates an encoding vector, respectively reconstructs short-term, medium-term and long-term statistical features and protocol-aware gating features, weights and sums the errors according to branch weights, predicts a malicious probability based on the encoding vector, obtains a comprehensive malicious score, judges a malicious sample, and sends the data to the intelligent analysis module;
[0109] The intelligent analysis module receives the data sent by the malicious sample automatic analysis model construction module, and obtains the sample type corresponding to the real-time network traffic sample data based on the output of the malicious sample automatic analysis model.
[0110] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such a process, method, article or device.
[0111] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the present application.
[0112] The above describes the present application and its embodiments, which are not limited, and the drawings only show one of the embodiments of the present application, and the actual structure is not limited thereto. In general, if a person skilled in the art is inspired thereby, without departing from the purpose of the present application, without creative design, similar structure and embodiments of the technical solution are not creative, and should belong to the protection scope of the present application.
Claims
1. A deep learning-based automated analysis method for malicious samples, characterized by: The method comprises the following steps: Step S1: Sample data integration: collect historical network traffic sample data, pre-process the historical network traffic sample data, and construct a network traffic sample data set; Step S2: Sample feature extraction: calculate edge weights, construct a network traffic association graph, introduce a time decay factor, obtain dynamic neighbor aggregation features, convert protocol types into embedding vectors, generate protocol-aware attention weights, perform weighted concatenation to obtain protocol-aware enhanced features, obtain neighborhood aggregation features based on protocol neighbor attention, obtain protocol-aware gating features based on gating vectors, and extract statistical feature vectors at three scales. Step S3: Construct an automatic analysis model for malicious samples. Based on the interaction gating vector, cross-scale interaction features are obtained. The protocol deviation term is introduced to calculate the dynamic attention weight, generate an encoding vector, reconstruct the short-term, medium-term, and long-term statistical features and the protocol-aware gating features, respectively, and sum the reconstruction errors according to the weighted branch weights. The malicious probability is predicted based on the encoding vector to obtain a comprehensive malicious score and judge the malicious sample. Step S4: Intelligent analysis: Based on the output of the malicious sample automatic analysis model, the sample type corresponding to the real-time network traffic sample data is obtained.
2. The method for automatic malicious sample analysis based on deep learning according to claim 1, characterized in that: In step S2, the sample feature extraction specifically includes the following steps: Step S21: constructing a network traffic association graph; based on the network traffic sample dataset, each sample is regarded as a node in the graph, the edge weights between the sample nodes are calculated by fusing multi-dimensional association features, and an adjacency matrix is constructed according to the association strength threshold to obtain a network traffic association graph; Step S22: Dynamic neighbor feature aggregation: traverse each sample node in the network traffic association graph, calculate the start time difference between it and all neighbor sample nodes, obtain the time decay factor, fuse the adjacency relationship and the time decay factor, and weightedly aggregate the features of the neighbor sample nodes to obtain the dynamic neighbor aggregation feature; Step S23: protocol perception enhancement: The one-hot vector of the protocol type is converted into a protocol type embedding vector through the embedding matrix, and the protocol perception attention weight matching the original feature dimension is generated through the multi-layer perceptron. The original features are weighted element by element, and the weighted features are spliced with the dynamic neighbor aggregation features to obtain the protocol perception enhancement features. Step S24: gated neighborhood enhancement; Step S25: multi-scale statistical features.
3. The method for automatic malicious sample analysis based on deep learning according to claim 2, characterized in that: In step S24, the gated neighborhood enhancement is to generate protocol neighbor attention for the sample node in the network traffic association graph, combining the protocol type embedding vector and the protocol perception enhancement feature of the sample node itself and its neighbor sample nodes, and obtain the neighborhood aggregation feature by weighted aggregation of the protocol perception enhancement feature of the neighbor sample nodes. After splicing the neighborhood aggregation feature with the original feature of the sample node, a fusion intermediate feature is generated through linear transformation, the fusion intermediate feature is linearly transformed and activated by Sigmoid to obtain a gating vector and multiply it element-wise by the fusion intermediate feature to obtain the protocol perception gated feature.
4. The method for automatic malicious sample analysis based on deep learning according to claim 2, characterized in that: In step S25, the multi-scale statistical features are based on the start time of each network traffic sample data, and the sliding time window is divided into three scales: short-term, medium-term and long-term. In each window, statistical features are extracted for all samples contained therein to obtain the statistical feature vectors of each sample at three scales.
5. The method for automatic malicious sample analysis based on deep learning according to claim 1, characterized in that: In step S3, the construction of the malicious sample automatic analysis model is based on the protocol-aware gating features and the statistical feature vectors of the three scales, and the construction of the malicious sample automatic analysis model is completed according to the neural network; specifically, the following steps are included: Step S31: cross-protocol dynamic encoder; Step S32: multi-branch reconstruction decoder; comprising the following steps: Step S321: Multi-branch feature reconstruction; the multi-branch reconstruction decoder contains four parallel branches, which correspond to the reconstruction of short-term, medium-term, long-term and protocol-aware gated features based on the encoding vector. Each branch adopts a structure that combines a deconvolution layer, a batch normalization layer and a sigmoid activation function. Step S322: multi-branch reconstruction error; calculate the reconstruction error of each branch separately, perform weighted summation of the reconstruction errors according to the weight of each branch, and obtain the total reconstruction error; Step S33: Malicious sample judgment: The malicious probability is predicted based on the coding vector, and half of the normalized sum of the total reconstruction error and the malicious probability is used as the comprehensive malicious score. A score threshold is set. If the comprehensive malicious score is greater than or equal to the score threshold, the sample is judged to be a malicious sample, otherwise the sample is judged to be a normal sample, and the data label is output.
6. The method for automatic malicious sample analysis based on deep learning according to claim 5, characterized in that: In step S31, the cross-protocol dynamic encoder specifically includes the following steps: Step S311: Cross-scale feature interaction: linearly transform the statistical feature vector of each scale, calculate the positive interaction gating vector and the negative interaction gating vector between the short-term and medium-term, and medium-term and long-term scales, and obtain the cross-scale interaction feature based on the interaction gating vector; Step S312: protocol-aware dynamic encoding; concatenate the cross-scale interaction features, protocol-aware gating features, and protocol type embedding vectors to obtain a unified fusion feature, perform a linear transformation on the unified fusion feature using three independent weight matrices to obtain the query vector, key vector, and value vector of the attention mechanism, introduce a protocol bias term, calculate the dynamic attention weight, and perform weighted summation on the value vector to obtain the final encoding vector.
7. The method for automatic malicious sample analysis based on deep learning according to claim 1, characterized in that: In step S1, the sample data integration is to collect historical network traffic sample data; the historical network traffic sample data includes the start time, duration, protocol type, source IP address, source port, destination IP address, destination port, TCP flag, number of packets, number of bytes, standard deviation of packet length, average packet length, minimum packet length, maximum packet length, standard deviation of arrival time interval and sample type of the network traffic session; the sample type includes normal samples and malicious samples; the historical network traffic sample data is preprocessed; the sample type is used as a data label to construct a network traffic sample data set with data labels.
8. The method for automatic malicious sample analysis based on deep learning according to claim 1, characterized in that: In step S4, the intelligent analysis collects real-time network traffic sample data, pre-processes the real-time network traffic sample data, extracts protocol-aware gating features and statistical feature vectors of three scales, and inputs them into the malicious sample automatic analysis model for processing. According to the output data label, the sample type corresponding to the real-time network traffic sample data is obtained.
9. A deep learning-based automatic malicious sample analysis system, configured to implement the deep learning-based automatic malicious sample analysis method according to any one of claims 1 to 8, characterized in that: It includes sample data integration module, sample feature extraction module, malicious sample automatic analysis model building module and intelligent analysis module; The sample data integration module collects historical network traffic sample data, preprocesses the historical network traffic sample data, constructs a network traffic sample data set, and sends the data to the sample feature extraction module; The sample feature extraction module receives data sent by the sample data integration module, calculates edge weights, constructs a network traffic association graph, introduces a time decay factor, obtains dynamic neighbor aggregation features, converts protocol types into embedding vectors, generates protocol-aware attention weights, performs weighted concatenation to obtain protocol-aware enhanced features, obtains neighborhood aggregation features based on protocol neighbor attention, obtains protocol-aware gating features based on gating vectors, extracts statistical feature vectors at three scales, and sends the data to the module for constructing a malicious sample automatic analysis model; The module for constructing a malicious sample automatic analysis model receives data sent by the sample feature extraction module, obtains cross-scale interaction features based on the interaction gating vector, introduces a protocol deviation term to calculate the dynamic attention weight, generates an encoding vector, reconstructs short-term, medium-term, and long-term statistical features and protocol-aware gating features, sums the reconstruction errors weighted by branch weights, predicts the malicious probability based on the encoding vector, obtains a comprehensive malicious score, determines the malicious sample, and sends the data to the intelligent analysis module; The intelligent analysis module receives data sent by the module for constructing the automatic analysis model for malicious samples, and obtains the sample type corresponding to the real-time network traffic sample data based on the output of the automatic analysis model for malicious samples.
Citation Information
Patent Citations
Network attack dynamic detection and security protection method and system based on artificial intelligence
CN120342748A
Network traffic anomaly detection strategy generation method based on machine learning
CN120415800A
Network security malicious traffic tracing method based on generative adversarial network
CN120415910A
Malicious behavior identification method and system for weighted heterogeneous graph, and storage medium
US20230362175A1
Cited By
Multi-dimensional data flow monitoring and exception interception protection method and system
CN121012694A
Calculation cluster job operation time-oriented estimation method
CN122285461A