A method for detecting scanning traffic in an intranet and determining a scanning tool based on multi-feature fusion
Patent Information
- Application Number
- CN202510669131.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2025-05-23
- Publication Date
- 2026-09-29
AI Technical Summary
这两种方法相对较为简单,当攻击者检测到规则后很容易进行绕过
[0061]1.本发明提出了一种基于多特征融合的内网扫描流量检测与扫描工具判定方法,采用深度学习方法对会话流统计特征和流量图进行特征提取,利用特征融合方法充分混合两种特征,使用融合后的特征对扫描流量和扫描工具进行识别。
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer network security, and in particular to a method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion. Background Technology
[0002] With the continuous development of the internet and the increasing number of active online users, the overall network security situation has become more severe, with various forms of cyberattacks emerging one after another. Many attacks targeting corporate internal networks begin with some form of reconnaissance, and one of the most well-known methods for achieving this is port scanning. As a preliminary step in internal network attacks, network probing and scanning to gather information about the internal network significantly influence the success rate of cyberattacks.
[0003] Port scan detection began as early as 1990, and currently, detection methods can be categorized into rule-based, threshold-based, statistical, and AI-based methods. Rule-based methods filter scan traffic by detecting packet flags or simple statistical measures. Threshold-based methods detect scans by comparing the behavioral characteristics of source IPs with preset thresholds. These two methods are relatively simple, and attackers can easily bypass them once they detect the rules. Statistical methods, by extracting more refined statistical features, increase the complexity of bypass detection, but determining the corresponding detection thresholds is more difficult. In recent years, AI-based methods have also emerged, but most of them identify port scans as one of many attacks, lacking further research into port scan detection, such as scanning tool identification.
[0004] Different port scanning tools differ in functionality, providing varying information and posing different threats. Therefore, identifying these tools not only helps in obtaining more accurate threat intelligence but also reveals attackers' intentions and techniques. This is of great significance in helping internal security personnel identify attacker intent and implement precise defenses. Summary of the Invention
[0005] The main problem addressed by this invention is to identify potential scanning traffic within an intranet and further identify the tools initiating the scan. Therefore, a method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion is proposed. This method extracts communication features from network traffic, constructs the communication behaviors of different scanning tools, and ultimately achieves the goal of identifying the scanning tools.
[0006] The technical solution of this invention is: a method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion, the steps of which are as follows:
[0007] A. Using the five-tuple of data packets (source IP, source port, destination IP, destination port, protocol) as identifiers, data packets are aggregated to form a session stream.
[0008] B. For each session flow, extract its typical statistical features, including: one-way flow features, session flow features, time features, etc.
[0009] C. For the extracted typical statistical features, normalization is first performed, then mutual information and recursive feature elimination strategy based on random forest cross-validation are used to filter features, and finally, multilayer perceptron is used to extract statistical features.
[0010] D. For each session flow, its session traffic is visualized according to certain rules to form a traffic graph.
[0011] E. For the image information of each session flow—the flow graph—Swin Transform is used to extract deep features from the image information to form flow graph features.
[0012] F. For the statistical features and flow graph features obtained from deep extraction, a low-rank feature fusion strategy is used to deeply fuse the two types of features while preserving the original information, forming the final feature vector for prediction.
[0013] G. For the final fused feature vector, a feedforward neural network is used to detect the scanning traffic and determine the scanning tool.
[0014] The typical statistical characteristics are composed of the following:
[0015] a. Unidirectional flow characteristics include: the number of forward and reverse data packets, the total size, maximum, minimum, average size, and standard deviation of the payload bytes of forward and reverse data packets, statistical characteristics of the header bytes of forward and reverse data packets, the ratio of data packets transmitted in reverse / forward, and the ratio of payload bytes of data packets transmitted in reverse / forward, etc.
[0016] b. Session flow characteristics include: the number of session flow packets, the total number of session flow payload bytes, maximum value, minimum value, average value and standard deviation, the number of packets with various flags, the number of packets with payload, etc.
[0017] c. Time characteristics include: stream duration, number of data packets transmitted per second, number of data packet payload bytes transmitted per millisecond, and maximum, minimum, average, and standard deviation of data packet intervals.
[0018] The process for identifying typical statistical features involves first normalization, then feature elimination using mutual information and a recursive feature elimination strategy based on random forest cross-validation, and finally, deep extraction of statistical features using a multilayer perceptron. The steps are as follows:
[0019] a. For the extracted traffic statistics features, use minimum and maximum value normalization to normalize the statistical features.
[0020] b. Calculate the mutual information of each feature according to its feature type using the mutual information calculation formulas for discrete and continuous features respectively.
[0021] c. Set candidate thresholds for mutual information and use each candidate threshold to filter candidate features.
[0022] d. Construct a random forest model using the selected candidate features, select the threshold with the best average cross-validation score as the final threshold for mutual information, and select the final candidate features based on the final threshold.
[0023] e. Construct a random forest model using the final candidate features, and then use a cross-validation recursive feature elimination strategy to further filter features.
[0024] f. Construct a multilayer perceptron to perform deep feature extraction on the selected features, forming the final typical statistical features.
[0025] The formula for normalizing the minimum and maximum values is as follows:
[0026]
[0027] Where min and max are the minimum and maximum values in the sample data, respectively.
[0028] The formula for calculating the discrete feature mutual information is as follows:
[0029]
[0030] Where f i Let be the set of possible values for the i-th discrete feature, Y be the set of possible classification labels, p(x,Y) be the joint probability distribution function of feature value x and classification label Y, p(x) be the marginal probability density function of feature value x, and p(y) be the marginal probability of classification label y.
[0031] The formula for calculating the continuous feature mutual information is as follows:
[0032]
[0033] Where p(f) i ,y) is the flow characteristic f i The joint probability density function of the result Y, p(f i p(y) and p(y) are the flow characteristics f i The marginal probability density function of the classification result Y. I(f) i ;Y) is the flow characteristic f iThe mutual information between the classification result Y and the classification result Y.
[0034] The threshold with the optimal average cross-validation score is selected as the final threshold for mutual information, where the score represents the classification accuracy, and its average accuracy formula is:
[0035]
[0036] Where n is the number of cross-validations, and TP, TN, FP, and FN represent the number of correctly predicted positive classes, correctly predicted negative classes, incorrectly predicted positive classes, and incorrectly predicted negative classes in each validation, respectively.
[0037] The steps for constructing a random forest model using the final candidate features and then further filtering features using a cross-validation recursive feature elimination strategy are as follows:
[0038] a. Construct a random forest model using the complete feature set.
[0039] b. Train the random forest model using 5-fold cross-validation and calculate the average prediction accuracy.
[0040] c. Based on the importance of each feature given by the random forest, remove the feature with the lowest importance and reformat the feature set.
[0041] d. Repeat the above steps until the feature set size meets the predetermined size, and select the feature set with the highest average prediction accuracy as the final feature set.
[0042] The steps for the graphical processing of the session traffic are as follows:
[0043] a. For each data packet in the session stream, take the first 32 rows (32 bits per row) to form a 32*32 image with channel 1. When the data packet length is insufficient, pad with bits 0 at the end; when the data packet length is too long, truncate the data packet.
[0044] b. For a session stream, take 8 communication data packets to form an image, and finally form an image data with 8 channels for the session stream. If a session stream does not meet the requirement of 8 communication data packets, fill the image with bits 0 to make up 8 communication data packets.
[0045] The steps for extracting deep features from image information using Swin Transform to form flow map features are as follows:
[0046] a) Divide the session stream image data into 4 windows, each window is 16*16 in size, and each window contains 16 patches, each patch is 4*4 in size.
[0047] b. Calculate the attention information within each window.
[0048] c. Disrupt the original window layout and move two patches to the upper left to form a new window, thus merging the attention limited to each window with the attention in the new window formed after the movement.
[0049] d. Repeat steps b and c until the predetermined number of repetitions is reached. Then, calculate the mean of the image data in the channel dimension and form a one-dimensional vector as the image feature for depth extraction.
[0050] The statistical features and flow graph features obtained from deep extraction are then fused using a low-rank feature fusion strategy to retain the original information and form the final feature vector for prediction. The steps are as follows:
[0051] a) For the extracted statistical feature vector and flow map feature vector, append 1 to the end of both feature vectors to ensure that the original feature information is preserved after the outer product.
[0052] b. Perform an outer product between the statistical feature vector (with 1 appended) and the flow graph feature vector to form a compact multi-feature fusion tensor.
[0053] c. Perform a linear transformation on the fusion tensor after the outer product to obtain the final fusion feature vector.
[0054] The formula for taking the outer product of the statistical feature vector with the appended 1 and the flow graph feature vector is as follows:
[0055]
[0056] Among them, z m This represents the feature vector after adding 1, and M represents the number of different feature types.
[0057] The fusion tensor after the outer product is linearly transformed to obtain the final fusion feature vector. The linear transformation formula is as follows:
[0058]
[0059] Where W is the weight and b is the offset. The weights for rank-based decomposition. This represents the element-wise product over the fusion tensor.
[0060] The beneficial effects of this invention are:
[0061] 1. This invention proposes a method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion. It uses deep learning to extract features from session flow statistics and traffic graphs, and uses a feature fusion method to fully mix the two features. The fused features are then used to identify scanning traffic and scanning tools.
[0062] 2. This invention proposes a network flow representation method, named Flow Graph (TG). It represents network flow characteristics by forming a multi-channel image from multiple data packets within a session.
[0063] 3. By fully integrating statistical features and flow graph features, this invention enhances the model's detection performance in different application scenarios and enables it to detect unknown scan flow. Attached Figure Description
[0064] Figure 1 This is the flowchart of the method.
[0065] Figure 2 A diagram illustrating the recursive feature elimination process using RF-based cross-validation.
[0066] Figure 3 The diagram shows the performance of this method in two scenarios.
[0067] Figure 4 This is a graph showing the performance of this method on the CICIDS2017 dataset. Detailed Implementation
[0068] Exemplary embodiments will now be described in more detail, and the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention.
[0069] The process flow shown in the embodiments of the present invention is as follows: Figure 1 As shown. First, in the data processing module, the captured traffic data file (Pcap) is processed to extract traffic statistical features and traffic image features. The specific steps of this process are as follows: 1. Using the packet 5-tuple (source IP, source port, destination IP, destination port, protocol) as an identifier, the packets are aggregated to form a session flow. 2. For each session flow, its typical statistical features are extracted, including: unidirectional flow features, session flow features, time features, etc. At the same time, its session traffic is imaged according to certain rules to form a traffic graph.
[0070] The specific steps for forming the traffic graph are as follows: 1. For each data packet in the session flow, take the first 32 rows (32 bits per row) to form a 32*32 image with channel 1. If the data packet length is insufficient, pad with bits 0 at the end; if the data packet length is too long, truncate the data packet. 2. For a session flow, take 8 communication data packets to form an image, ultimately forming an image data with channel 8 for the session flow. If a session flow does not have 8 communication data packets, pad with bits 0 to make up the 8 communication data packet images.
[0071] Subsequently, in the core functional module, typical statistical features are screened and deep features are extracted, and then fused with the deeply extracted flow graph features to construct a good scanning feature representation.
[0072] For typical statistical features, normalization is first performed. Then, mutual information and a recursive feature elimination strategy based on random forest cross-validation are used to filter features. Finally, a multilayer perceptron is used to deeply extract statistical features. The specific steps are as follows: 1. For the extracted traffic statistical features, minimum-maximum value normalization is used to normalize the statistical features. 2. According to the feature type, the mutual information of each feature is calculated using the mutual information calculation formula for discrete and continuous features respectively. 3. Candidate thresholds for mutual information are set, and candidate features are filtered using each candidate threshold. 4. A random forest model is built using the filtered candidate features. The threshold with the best average cross-validation score is selected as the final threshold for mutual information, and the final candidate features are filtered based on the final threshold. 5. A random forest model is built using the final candidate features, and a recursive feature elimination strategy based on cross-validation is used to filter features again. The filtering process is as follows. Figure 2 As shown. 6. Construct a multilayer perceptron to perform deep feature extraction on the selected features, forming the final typical statistical features. In this embodiment, the candidate thresholds are selected as [0.1, 0.2, 0.3, 0.4], and the selected cross-validation score is the recognition accuracy. To address this problem, it is necessary to introduce the minimum and maximum value normalization formula, the discrete mutual information calculation formula, the continuous mutual information calculation formula, and the classification recognition average accuracy formula.
[0073]
[0074] In formula (1), min and max are the minimum and maximum values in the sample data, respectively.
[0075]
[0076] In formula (2) f i Let be the set of possible values for the i-th discrete feature, Y be the set of possible classification labels, p(x,Y) be the joint probability distribution function of feature value x and classification label Y, p(x) be the marginal probability density function of feature value x, and p(y) be the marginal probability of classification label y.
[0077]
[0078] In formula (3), p(f) i ,y) is the flow characteristic f i The joint probability density function of the result Y, p(f i p(y) and p(y) are the flow characteristics f i The marginal probability density function of the classification result Y. I(f)i ;Y) is the flow characteristic f i The mutual information between the classification result Y and the classification result Y.
[0079]
[0080] In formula (4), n is the number of cross-validations, and TP, TN, FP, and FN represent the number of correctly predicted positive, correctly predicted negative, incorrectly predicted positive, and incorrectly predicted negative classes in each validation, respectively.
[0081] For the generated traffic graph, the Swing Transform is used to extract deep features from the image information, forming traffic graph features. The specific steps are as follows: 1. Divide the session flow image data into four windows, each 16*16 pixels, containing 16 patches, each 4*4 pixels. 2. Calculate the attention information within each window. 3. Shuffle the original window layout, moving two patches to the upper left to form a new window, and fuse the attention information confined to each window with the attention information in the new window. 4. Repeat steps 2 and 3 until a predetermined number of repetitions are reached. Then, average the image data along the channel dimension to form a one-dimensional vector as the deep-extracted image features.
[0082] For the statistical features and flow graph features obtained from deep extraction, a low-rank feature fusion strategy is used to deeply fuse the two types of features while preserving the original information, forming the final feature vector for prediction. The specific steps are as follows: 1. Append 1s to the end of both the extracted statistical feature vector and the flow graph feature vector to ensure that the original feature information is preserved after the outer product. 2. Perform an outer product between the appended statistical feature vector and the flow graph feature vector to form a compact multi-feature fusion tensor. 3. Perform a linear transformation on the fusion tensor after the outer product to obtain the final fused feature vector. This problem requires the introduction of vector outer product formulas and linear transformation fusion formulas.
[0083]
[0084] In formula (6), z m This represents the feature vector after adding 1, and M represents the number of different feature types.
[0085]
[0086] In formula (7), W is the weight and b is the offset. It is the weight of the rank decomposition. This represents the element-wise product over the fusion tensor.
[0087] Finally, for the final fused feature vector, a feedforward neural network is used to detect the scan traffic and determine the scanning tool.
[0088] The constructed detection method and recognition model were trained using 70% of the collected traffic from various scanning tools, and evaluated using the remaining 30%. The method's Macor-Accuracy, Macor-Precision, and Macor-Recall scores were 0.9981, 0.9925, and 0.9924, respectively. To verify the model's robustness, transfer testing was performed. In the new scenario, the method's Macor-Accuracy, Macor-Precision, and Macor-Recall scores were 0.9839, 0.9235, and 0.9153, respectively. The results are as follows... Figure 3 As shown. To test the identification of unknown scanning behavior, the method was validated using the scan traffic in the public dataset CICIDS2017. The accuracy, precision, and recall scores of the method were 0.9740, 0.9924, and 0.9774, respectively. The performance on the public dataset is as follows. Figure 4 As shown.
[0089] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion, comprising the following steps: 1) Data packets are aggregated to form a session stream, identified by the 5-tuple of the data packet (source IP, source port, destination IP, destination port, protocol); 2) For each session stream, extract its typical statistical features. include: Unidirectional flow characteristics, session flow characteristics, time characteristics, etc.; 3) For the extracted typical statistical features, normalization is performed first, then mutual information and cross-validation recursive feature elimination strategy based on random forest are used to filter features, and finally, multilayer perceptron is used to extract statistical features. 4) For each session flow, its session traffic is visualized according to certain rules to form a traffic graph; 5) For the image information of each session flow—the flow graph—Swin Transform is used to extract deep features from the image information to form flow graph features; 6) For the statistical features and flow map features obtained by deep extraction, a low-rank feature fusion strategy is used to deeply fuse the two types of features while retaining the original information, forming the final feature vector for prediction. 7) For the final fused feature vector, a feedforward neural network is used to detect the scanning flow and determine the scanning tool.
2. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 1, characterized in that: The typical statistical characteristics are as follows: 1) Unidirectional flow characteristics include: number of forward and reverse data packets, total size, maximum, minimum, average size and standard deviation of forward and reverse data packet payload bytes, statistical characteristics of forward and reverse data packet header bytes, ratio of data packets transmitted in reverse / forward direction, ratio of data packets transmitted in reverse / forward direction payload bytes, etc. 2) Session flow characteristics include: number of session flow packets, total number of session flow payload bytes, maximum value, minimum value, average value and standard deviation, number of packets with various flags, number of packets with payload, etc. 3) Time characteristics include: stream duration, number of data packets transmitted per second, number of data packet payload bytes transmitted per millisecond, maximum, minimum, average, and standard deviation of data packet interval time.
3. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 1, characterized in that: First, normalization is performed. Then, mutual information and a recursive feature elimination strategy based on random forest cross-validation are used to filter features. Finally, a multilayer perceptron is used to extract statistical features. The steps are as follows: 1) For the extracted traffic statistics features, the minimum and maximum value normalization is used to normalize the statistical features; 2) Calculate the mutual information of each feature according to its feature type using the mutual information calculation formulas for discrete and continuous features respectively; 3) Set candidate thresholds for mutual information, and use each candidate threshold to filter candidate features; 4) Construct a random forest model using the selected candidate features, select the threshold with the best average cross-validation score as the final threshold for mutual information, and select the final candidate features based on the final threshold. 5) Construct a random forest model using the final candidate features, and then use a cross-validation recursive feature elimination strategy to further filter features; 6) Construct a multilayer perceptron to perform deep feature extraction on the selected features, forming the final typical statistical features.
4. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion as described in claim 1, characterized in that: The steps for graphical processing of session traffic are as follows: 1) For each data packet in the session stream, take the first 32 rows (32 bits per row) to form a 32*32 image with channel 1. If the data packet length is insufficient, pad with bits 0 at the end; if the data packet length is too long, truncate the data packet. 2) For a session stream, take 8 communication data packets to form an image, and finally form an image data with 8 channels for the session stream. If a session stream does not meet the requirement of 8 communication data packets, fill the image with bits 0 to make up 8 communication data packets.
5. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 1, characterized in that: The steps for extracting deep features from image information using the Swing Transform to form flow graph features are as follows: 1) Divide the session stream image data into 4 windows, each window is 16*16 in size, each window contains 16 patches, and each patch is 4*4 in size; 2) Calculate attention information within each window; 3) Disrupt the original window layout and move two patches to the upper left to form a new window, thus merging the attention that is confined within each window with the attention within the new window formed after the movement; 4) Repeat steps b and c until the predetermined number of repetitions is reached. Then, calculate the mean of the image data in the channel dimension and form a one-dimensional vector as the image feature for depth extraction.
6. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 1, characterized in that: For the statistical features and flow graph features obtained from deep extraction, a low-rank feature fusion strategy is used to deeply fuse the two types of features while preserving the original information, forming the final feature vector for prediction. The steps are as follows: 1) For the extracted statistical feature vector and flow map feature vector, append 1 to the end of the two feature vectors to ensure that the original feature information is preserved after the outer product; 2) Perform an outer product between the statistical feature vector with the appended 1 and the flow graph feature vector to form a compact multi-feature fusion tensor; 3) Perform a linear transformation on the fusion tensor after the outer product to obtain the final fusion feature vector.
7. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 3, characterized in that: For the extracted traffic statistics features, minimum-maximum value normalization is used to normalize the statistical features. The normalization formula is as follows: Where min and max are the minimum and maximum values in the sample data, respectively.
8. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 3, characterized in that: The mutual information of each feature is calculated using the mutual information calculation formulas for discrete and continuous features, respectively, based on its feature type. The mutual information calculation formula for discrete features is as follows: Where f i Let be the set of possible values for the i-th discrete feature, Y be the set of possible classification labels, p(x,Y) be the joint probability distribution function of feature value x and classification label Y, p(x) be the marginal probability density function of feature value x, and p(y) be the marginal probability of classification label y.
9. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 3, characterized in that: For the extracted traffic statistics features, the mutual information of each feature is calculated using the mutual information calculation formulas for discrete and continuous features, respectively, according to the feature type. The mutual information calculation formula for continuous features is as follows: Where p(f) i ,y) is the flow characteristic f i The joint probability density function of the result Y, p(f i p(y) and p(y) are the flow characteristics f i The marginal probability density function of the classification result Y. I(f) i ;Y) is the flow characteristic f i The mutual information between the classification result Y and the classification result Y.
10. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 3, characterized in that: The threshold with the optimal average cross-validation score is selected as the final threshold for mutual information, where the score represents the classification accuracy, and its average accuracy formula is: Where n is the number of cross-validations, and TP, TN, FP, and FN represent the number of correctly predicted positive, correctly predicted negative, incorrectly predicted positive, and incorrectly predicted negative classes in each validation, respectively.
11. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 3, characterized in that: A random forest model is constructed using the final candidate features, and a cross-validation recursive feature elimination strategy is used to further filter the features. The steps are as follows: 1) Construct a random forest model using the complete feature set; 2) The random forest model was trained using 5-fold cross-validation, and the average prediction accuracy was calculated. 3) Based on the importance of each feature given by the random forest, remove the feature with the lowest importance and reformat the feature set; 4) Repeat the above steps until the feature set size meets the predetermined size, and select the feature set with the highest average prediction accuracy as the final feature set.
12. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 6, characterized in that: The formula for outer product between the statistical feature vector (with the appended 1) and the flow graph feature vector is as follows: Among them, z m This represents the feature vector after adding 1, and M represents the number of different feature types.
13. The method for detecting intranet scanning traffic and determining scanning tools based on multi-feature fusion according to claim 6, characterized in that: The fused tensor after the outer product is linearly transformed to obtain the final fused feature vector. The linear transformation formula is as follows: Where W is the weight and b is the offset. The weights for rank-wise decomposition. This represents the element-wise product over the fusion tensor.