A Network Security Attack Classification and Detection Method and System Based on Random Forest
Through time-based sliding window algorithm and improved random forest algorithm, combined with feature importance analysis, the problems of high computational complexity and low efficiency of the detection method of the network security attack classification in the prior art are solved, and fast response and efficient network security attack classification are achieved.
Patent Information
- Application Number
- CN202411740569.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The existing network security attack classification detection methods are highly complex and time-consuming when processing large-scale and multi-dimensional network traffic data, making it difficult to respond to sudden network attacks quickly. They lack in-depth analysis of the importance of features, resulting in inefficient classification and inability to meet real-time classification requirements.
The time-based sliding window algorithm is used to determine the target network traffic data, extract multiple network traffic characteristics and calculate the average entropy value, and classify and detect it through the improved random forest algorithm, combine feature importance analysis and dynamic feature selection to optimize the decision tree structure to improve classification efficiency.
It effectively reduces the computational complexity, improves classification efficiency, can quickly respond to sudden network attacks, meets real-time classification needs, and improves the identification and response speed of network security threats.
Smart Images

Figure CN119583172B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a method and system for classifying and detecting network security attacks based on random forest. Background Art
[0002] With the rapid development of Internet technology, the types and complexity of network attacks are also increasing continuously. Network security issues have become the core issues of common concern in various industries. Especially under the promotion of emerging technologies such as cloud computing, big data, and the Internet of Things, the scale of network traffic shows exponential growth, and at the same time, the traffic structure is becoming increasingly complex, resulting in more diverse and more concealed forms of network attacks. Therefore, how to effectively and accurately classify and detect network security attacks has become a key technical challenge that urgently needs to be solved in the field of network security.
[0003] In the existing technical solutions, when traditional network security attack classification and detection methods process large-scale and multi-dimensional network traffic data, the computational complexity is relatively high, and they cannot quickly respond to sudden network attacks.
[0004] In addition, in the face of massive network traffic data, the existing detection methods lack in-depth analysis of the importance of features, resulting in low classification efficiency, a long classification process, and it is difficult to meet the requirements of real-time classification of network security attacks. Summary of the Invention
[0005] In order to solve the technical problems that when traditional network security attack classification and detection methods process large-scale and multi-dimensional network traffic data, the computational complexity is relatively high, the classification process takes a long time, they cannot quickly respond to sudden network attacks, and in the face of massive network traffic data, they lack in-depth analysis of the importance of features, resulting in low classification efficiency and it is difficult to meet the requirements of real-time classification of network security attacks, the present invention provides a method and system for classifying and detecting network security attacks based on random forest.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] A method for classifying and detecting network security attacks based on random forest provided by an embodiment of the present invention includes:
[0009] S1: Collect network traffic data;
[0010] S2: Determine target network traffic data through a time-based sliding window algorithm;
[0011] S3: Extract multiple network traffic features from the target network traffic data;
[0012] S4: Calculate the average entropy value of the network traffic characteristics;
[0013] S5: Determine whether the average entropy value of the network traffic characteristics is within a preset entropy value range; if so, preliminarily determine the target network traffic data as normal traffic data, and return to S1 for continued detection; otherwise, preliminarily determine the target network traffic data as abnormal traffic data, and proceed to the next step;
[0014] S6: Process the abnormal traffic data;
[0015] S7: Through the improved random forest algorithm, perform classification detection on the processed abnormal traffic data to obtain an accurate detection result, and return to S1 for continued detection.
[0016] Second aspect:
[0017] A network security attack classification detection system based on random forest provided by an embodiment of the present invention includes: a memory and one or more processors;
[0018] One or more application programs are stored in the memory, and the one or more application programs are adapted to be executed by the one or more processors to implement the above-mentioned network security attack classification detection method based on random forest.
[0019] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0020] In the present invention, through the time-based sliding window algorithm, the target network traffic data is determined. When processing large-scale and multi-dimensional network traffic data, the data processing scale is effectively reduced, and the computational complexity is reduced. By extracting multiple network traffic characteristics from the target network traffic data and calculating the average entropy value of the network traffic characteristics, the target network traffic data is preliminarily determined as normal traffic data or abnormal traffic data, which can quickly respond to sudden network attacks. Through the improved random forest algorithm for classification detection, in-depth analysis of the feature importance can be carried out, with high classification efficiency and short classification process time, which can meet the requirements of real-time classification of network security attacks. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic flowchart of a network security attack classification detection method based on random forest provided by an embodiment of the present invention;
[0023] Figure 2 This is a schematic structural diagram of a network security attack classification and detection system based on random forest provided by an embodiment of the present invention. Detailed implementation manners
[0024] The following describes the technical solutions in the present invention with reference to the accompanying drawings.
[0025] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" aims to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.
[0026] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0027] Refer to the accompanying drawings of the specification Figure 1 , which shows a schematic flow diagram of a network security attack classification and detection method based on random forest provided by an embodiment of the present invention.
[0028] The embodiments of the present invention provide a network security attack classification and detection method based on random forest. This method can be implemented by a network security attack classification and detection device based on random forest. The network security attack classification and detection device based on random forest can be a terminal or a server. The processing flow of the network security attack classification and detection method based on random forest can include the following steps:
[0029] S1: Collect network traffic data.
[0030] Optionally, collect network traffic data transmitted through the network in real time through a network packet capture tool (such as Wireshark or tcpdump).
[0031] Among them, Wireshark is a graphical network protocol analysis tool that can capture and display the content of network packets transmitted in the network in real time. It supports the parsing of hundreds of network protocols and provides powerful filtering and search functions, enabling users to quickly locate the network packets or network events of interest. The user interface of Wireshark is intuitive, making it convenient for network administrators, security experts and network application developers to perform network debugging and network security analysis.
[0032] Among them, tcpdump is a packet analysis tool with a command-line interface that runs on various Unix-like operating systems. It can capture packets on a network interface and provide options to filter and display detailed information about these captured packets. tcpdump is more suitable for script usage or network monitoring and problem troubleshooting in environments where the GUI is not available.
[0033] In the present invention, using a network packet capture tool can capture the data stream in the network in real time, ensuring that the obtained data is the latest, thus more accurately reflecting the network state and traffic characteristics. This is crucial for timely detecting and responding to network security threats. The packet capture tool can capture detailed information of each packet (such as source IP, destination IP, port number, protocol type, etc.), making subsequent data analysis and feature extraction more meticulous, which helps to improve the accuracy of network security detection.
[0034] S2: Determine the target network traffic data through a time-based sliding window algorithm.
[0035] It should be noted that the time-based sliding window algorithm is a dynamic data processing technology used for real-time monitoring and analysis of data streams. In this algorithm, data is managed and processed in time windows of a fixed size. Whenever new data arrives, the window slides forward, retaining the latest data and removing the outdated data. This method can effectively capture and analyze the time series characteristics of data and is applicable to scenarios such as traffic monitoring, anomaly detection, and trend analysis, ensuring that the analysis results can reflect the current state rather than historical data, thereby improving the response speed and accuracy of the system.
[0036] In the present invention, the sliding window algorithm can instantaneously process and analyze network traffic data, ensuring that the system can quickly respond to changes in the network state and timely identify potential security threats or abnormal activities. The dynamic nature of the sliding window enables the algorithm to capture the changing trends and patterns of traffic in real time, providing a more accurate analysis of network traffic behavior, especially in the face of a rapidly changing network environment. By removing outdated data, the sliding window ensures that the analysis results reflect the current network state rather than historical data, thereby improving the timeliness and accuracy of decision-making.
[0037] In a possible implementation manner, S2 specifically includes sub-steps S201 and S202:
[0038] S201: Insert the network traffic data into the sliding window.
[0039] S202: Determine whether the collection time of the network traffic data is greater than the preset duration of the sliding window. If so, remove the network traffic data in the sliding window. Otherwise, retain the network traffic data in the sliding window and use the retained network traffic data as the target network traffic data.
[0040] It should be noted that those skilled in the art can set the size of the preset duration according to actual needs, and the present invention does not make any limitations in this regard.
[0041] In the present invention, by inserting network traffic data into the sliding window in real time, the system can quickly capture the latest traffic information, ensure that the analysis is based on the current state, thereby improving the response speed to network attacks and abnormal traffic. Judging whether the data collection time exceeds the preset duration enables the sliding window to be automatically updated and removes outdated data. In this way, the system can focus on the latest traffic characteristics and reduce the impact of redundant data. The mechanism for retaining valid data ensures that the analysis is based on the most relevant traffic information, enhances the accuracy of traffic monitoring and anomaly detection, and reduces the risks of false alarms and missed alarms.
[0042] S3: Extract multiple network traffic characteristics from the target network traffic data.
[0043] It should be noted that network traffic characteristics refer to various attributes or metrics used to describe and analyze network traffic data, and these characteristics can help identify the behavior patterns of traffic and their potential security threats.
[0044] Optionally, the network traffic characteristics specifically include: source IP address, destination IP address, source port, destination port, transport protocol (such as TCP, UDP), packet size, traffic duration, traffic direction (inbound or outbound), number of sessions, flag bits (such as TCP flag bits), and byte and packet counts in the traffic.
[0045] In the present invention, extracting multiple network traffic characteristics enables the system to analyze network traffic from multiple perspectives, which helps to comprehensively understand the behavior patterns of traffic and identify the differences between normal and abnormal traffic. Through the analysis of the characteristics, potential security threats such as DDoS attacks, intrusion behaviors, and data leaks can be identified in a timely manner. This feature extraction ability can help the network security protection system respond to various attacks more quickly and effectively. Different traffic characteristics can help build models to identify abnormal behaviors in traffic and further enhance the system's adaptability in the face of unknown threats.
[0046] S4: Calculate the average entropy value of the network traffic characteristics.
[0047] In the present invention, the entropy value, as an important index in information theory, is used to measure the uncertainty and complexity of network traffic characteristics. By calculating the average entropy value, the diversity and variation degree of traffic data can be effectively understood, which helps to identify potential abnormal behaviors. By calculating the average entropy value, the complex information of multiple traffic characteristics can be condensed into a single value, thereby simplifying the data analysis process and improving the data processing efficiency.
[0048] In a possible implementation, S4 specifically includes sub-steps S401 to S403:
[0049] S401: Calculate the entropy value of each network traffic feature through the Shannon entropy formula.
[0050] It should be noted that the Shannon entropy formula is a mathematical tool in information theory used to measure the uncertainty or randomness of information, proposed by Claude Shannon. This formula quantifies the complexity and amount of information by calculating the probability distribution of all possible values of a random variable.
[0051] Optionally, calculate the entropy value of each network traffic feature according to the following formula:
[0052]
[0053] Among them, H(O) represents the entropy value of the network traffic feature O, O represents the network traffic feature in the target network traffic data, p(o i ) represents the probability that the i-th eigenvalue of the network traffic feature appears, o i represents the i-th eigenvalue of the network traffic feature, and n represents the total number of eigenvalues of the network traffic feature.
[0054] S402: Normalize the entropy value of each network traffic feature to obtain the normalized entropy value of each network traffic feature.
[0055] Optionally, normalize the entropy value of each network traffic feature according to the following formula to obtain the normalized entropy value of each network traffic feature:
[0056]
[0057] Among them, H0(O) represents the normalized entropy value of the network traffic feature O.
[0058] S403: Calculate the average entropy value of the network traffic feature according to the normalized entropy value of each network traffic feature.
[0059] Optionally, calculate the average entropy value of the network traffic feature according to the following formula:
[0060]
[0061] Among them, H avg represents the average entropy value of the network traffic feature, M represents the total number of network traffic features, O m represents the m-th network traffic feature.
[0062] In the present invention, by calculating the entropy value, abnormal patterns in traffic can be effectively identified. When the entropy value of certain features is significantly higher than the normal level, it may indicate the existence of potential security threats, such as network attacks. This information can be used for real-time monitoring and response. Normalizing the entropy value can eliminate the scale differences between different features, making the features more comparable in subsequent analyses. This standardization helps to improve the performance and stability of machine learning models. By calculating the average entropy value, the information of multiple features is condensed into one metric, simplifying the data processing process. The average entropy value provides a measure of feature importance, helps to screen out the most useful features for the classification task, optimize the construction of the model, and improve the prediction accuracy.
[0063] S5: Determine whether the average entropy value of the network traffic features is within the preset entropy value range. If so, preliminarily determine the target network traffic data as normal traffic data, and return to S1 for continued detection. Otherwise, preliminarily determine the target network traffic data as abnormal traffic data, and proceed to the next step.
[0064] It should be noted that those skilled in the art can set the size of the preset entropy value range according to actual needs, and the present invention does not limit it here.
[0065] In the present invention, by setting the entropy value range, the normality and abnormality of network traffic can be quickly judged, helping the system to timely identify potential security threats, such as network attacks or data leaks, thereby improving the response speed. By real-time monitoring and evaluating the entropy value of network traffic features and comparing it with the preset threshold, potential abnormal or attack behaviors can be identified early. This early detection helps to prevent the spread and damage of security threats. By determining whether the traffic is normal, the system can effectively allocate resources and attention, focus on real threats, and avoid excessive intervention in normal activities, thereby optimizing the use of security resources.
[0066] S6: Perform data processing on the abnormal traffic data.
[0067] In the present invention, through processing steps such as data cleaning and normalization, noise and inconsistent data can be effectively removed, errors can be corrected, and missing values can be filled. This not only improves the quality of the data, but also provides accurate and clean data input for subsequent analyses and machine learning models, thereby improving the reliability and accuracy of the analysis results.
[0068] In a possible implementation manner, S6 specifically includes sub-steps S601 and S602:
[0069] S601: Perform data cleaning on the abnormal traffic data.
[0070] Optionally, data cleaning includes: removing duplicate records, handling missing values, and removing noise data.
[0071] Among them, removing duplicate records specifically involves checking whether there are duplicate data records in the abnormal traffic data and deleting the redundant copies to ensure the uniqueness of each piece of data. Handling missing values specifically involves identifying the missing values in the data and processing them according to the specific situation. Common methods include filling in the missing values (such as using the mean, median, or mode to fill), or directly deleting the records containing missing values. Removing noise data specifically involves identifying and deleting the outliers or incorrect data that may interfere with the analysis results, which may be caused by sensor failures, data transmission errors, etc.
[0072] S602: Normalize the abnormal traffic data after data cleaning.
[0073] Optionally, normalize the abnormal traffic data after data cleaning according to the following formula:
[0074]
[0075] where represents the i-th eigenvalue of the network traffic feature after normalization processing of represents the i-th eigenvalue of the network traffic feature of represents the network traffic feature in the abnormal traffic data, min represents taking the minimum value, and max represents taking the maximum value.
[0076] In the present invention, by removing duplicate records, handling missing values and noise data, the uniqueness and accuracy of the data are ensured, and the overall quality of the data set is improved. This cleaning process removes the inaccurate factors that may affect the analysis results, making the subsequent data processing and analysis more reliable. Cleaning the noise and inconsistencies in the data helps to reduce errors and biases and improve the accuracy of model training and prediction. Data normalization processing makes all eigenvalue on the same scale, which is crucial for most machine learning algorithms because these algorithms usually assume that all input data have the same scale. The normalized data helps to avoid the algorithm being biased towards certain features due to the difference in feature scales, thereby improving the overall performance of the model.
[0077] S7: Use the improved random forest algorithm to perform classification detection on the processed abnormal traffic data to obtain accurate detection results, and return to S1 to continue detection.
[0078] It should be noted that the random forest algorithm is an ensemble learning method mainly used for classification and regression tasks. It improves the accuracy and robustness of the model by constructing multiple decision trees and combining their prediction results. During the training process, the random forest uses the bootstrap method to randomly sample from the original dataset to generate multiple subsets and constructs independent decision trees for each subset. In addition, when splitting the nodes of each tree, it also randomly selects features to increase the diversity of the model and reduce overfitting. Finally, the random forest determines the final prediction by voting or averaging the results of all decision trees, thereby improving the stability and accuracy of the model. This algorithm performs well in dealing with high-dimensional data and large-scale datasets and is widely used in fields such as finance, healthcare, and cybersecurity.
[0079] Furthermore, the random forest algorithm is improved by introducing the importance evaluation of features and the dynamic feature selection mechanism. Specifically, the improvement methods include calculating the contribution degree of each feature to classification in real time, thereby dynamically adjusting the feature set so that each tree can use the most relevant features for splitting. This method can not only improve the classification performance of the model but also reduce the computational overhead and the risk of overfitting. At the same time, by combining the Out-Of-Bag Error to optimize the number and structure of decision trees, the algorithm can automatically adjust parameters on different training sets, further enhancing the adaptability and accuracy of the model.
[0080] Among them, the decision tree is a machine learning model for classification and regression, representing the decision-making process through a tree structure. Each internal node represents a test of a feature, the branches represent the test results, and the leaf nodes represent the final classification or prediction results. The decision tree starts from the root node, splits the data according to the values of the features, and recursively constructs layer by layer until the stopping condition is met (such as reaching the maximum depth or the number of samples in the node is less than the threshold). Due to its intuitive, easy-to-interpret, and visualizable characteristics, the decision tree is widely used in data analysis, feature selection, and model interpretation. However, a single decision tree is prone to overfitting, so it is often combined with other techniques (such as random forests) to improve the robustness and accuracy of the model.
[0081] In the present invention, by introducing the importance evaluation of features and the dynamic feature selection mechanism, the algorithm can select the most relevant features when constructing each decision tree, which helps to improve the model's ability to recognize complex data patterns, thereby improving the overall classification accuracy. Calculating the contribution degree of features in real time and dynamically adjusting the feature set can avoid using unnecessary features in the construction of each tree, which not only improves the computational efficiency but also reduces the consumption of memory and processing time. By combining the Out-Of-Bag Error to evaluate the model performance and dynamically adjusting the number and structure of decision trees, the model can better adapt to different datasets and changing data distributions, enhancing the generalization ability of the model.
[0082] In a possible implementation, S7 specifically includes sub-steps S701 to S715:
[0083] S701: Construct an initial forest composed of multiple decision trees, where the decision trees include multiple nodes.
[0084] S702: Calculate the entropy value of each node:
[0085]
[0086] where E represents the entropy value of the node, and p(c) represents the probability of class label c in the node.
[0087] S703: Randomly select multiple network traffic characteristics from the abnormal traffic data.
[0088] S704: According to the entropy value of each node, use the randomly selected network traffic characteristics to divide the node, obtaining the left child node and the right child node of each node.
[0089] S705: According to the entropy value of the left child node and the entropy value of the right child node, calculate the division quality of the network traffic characteristics for the node:
[0090] Q(i,j) = exp{-(E l +E r )}
[0091] where Q(i,j) represents the division quality of the j-th network traffic characteristic for the i-th node, E l represents the entropy value of the left child node, and E r represents the entropy value of the right child node.
[0092] S706: According to the division quality of the network traffic characteristics for the node, calculate the local weight of the network traffic characteristic in the decision tree:
[0093]
[0094] where w kj represents the local weight of the j-th network traffic characteristic in the k-th decision tree, Q k (i,j) represents the division quality of the j-th network traffic characteristic for the i-th node in the k-th decision tree, and N represents the total number of nodes in the decision tree.
[0095] S707: Based on the out-of-bag error, calculate the normalized weight of the decision tree:
[0096]
[0097] where η kDenote the normalized weight of the $k$-th decision tree as $\delta$ k Denote the out-of-bag error of the $k$-th decision tree, and $\max$ represents taking the maximum value.
[0098] Among them, the out-of-bag error (OOB Error) is a method used to evaluate the performance of models in random forests and other ensemble learning algorithms. Since random forests use the bootstrap method to randomly sample data when constructing each decision tree, it means that not all samples are used when training each tree. Usually, about one-third of the data is not selected, and these unselected data are called out-of-bag samples. The out-of-bag error provides an unbiased model evaluation metric by predicting the out-of-bag samples of each tree and calculating the classification or regression errors of these samples. Compared with traditional cross-validation methods, the out-of-bag error has the advantages of high computational efficiency and simplicity of implementation, so it is widely used in the model optimization and selection of random forests.
[0099] S708: Calculate the global weight of the network traffic feature according to the local weight of the network traffic feature in the decision tree and the normalized weight of the decision tree:
[0100]
[0101] Among them, $\omega$ j Denote the global weight of the $j$-th network traffic feature, and $K$ represents the total number of decision trees.
[0102] S709: Sort each network traffic feature according to the global weight of the network traffic feature.
[0103] S710: According to the sorting result, mark the first $u_0$ network traffic features as important features and construct the first important feature set $\Gamma_1$. Mark the remaining network traffic features as irrelevant features and construct the first irrelevant feature set $U_1$.
[0104] S711: Discard some network traffic features in the first irrelevant feature set $U_1$ according to the first preset condition and construct the second irrelevant feature set $U_2$.
[0105] In a possible implementation manner, the first preset condition is specifically:
[0106] $\omega$ a $<$ $\mu$ n $- 2\sigma$ n
[0107] Among them, $\omega$ a Denote the global weight of the $a$-th network traffic feature in the first irrelevant feature set, $\mu$ n Denote the mean value of the global weights of all network traffic features in the first irrelevant feature set, $\sigma$ nRepresents the standard deviation of the global weights of all network traffic features in the first set of irrelevant features.
[0108] S712: According to the second preset condition, add some network traffic features in the second set of irrelevant features U2 to the first set of important features Γ1, and construct the third set of irrelevant features U3 and the second set of important features Γ2.
[0109] In a possible implementation manner, the second preset condition is specifically:
[0110] ω b ≥min(ω c )
[0111] Where ω b represents the global weight of the b-th network traffic feature in the second set of irrelevant features, ω c represents the global weight of the c-th network traffic feature in the first set of important features, and min represents taking the minimum value.
[0112] S713: Determine that the change in the number of network traffic features between the third set of irrelevant features U3 and the first set of irrelevant features U1 is Δv, and the change in the number of network traffic features between the second set of important features Γ2 and the first set of important features Γ1 is Δu.
[0113] S714: According to the change in the number of network traffic features between the third set of irrelevant features U3 and the first set of irrelevant features U1 and the change in the number of network traffic features between the second set of important features Γ2 and the first set of important features Γ1, determine the increase or decrease value of the number of decision trees:
[0114] B new =B0 + λ u Δu + λ v Δv
[0115] Where B new represents the increase or decrease value of the number of decision trees, B0 represents the initial number of decision trees, λ u represents the decision coefficient of the change in the number of network traffic features between the second set of important features and the first set of important features, Δu represents the change in the number of network traffic features between the second set of important features and the first set of important features, λ v represents the decision coefficient of the change in the number of network traffic features between the third set of irrelevant features and the first set of irrelevant features, and Δv represents the change in the number of network traffic features between the third set of irrelevant features and the first set of irrelevant features.
[0116] S715: Repeat sub-steps S702 to S714 until the number of network traffic features in the third set of irrelevant features is less than the preset feature number, and then end the algorithm to obtain the accurate detection result.
[0117] It should be noted that those skilled in the art can set the size of the number of preset features according to actual needs, and the present invention does not make any limitations in this regard.
[0118] In the present invention, by calculating the entropy value and partition quality of each feature, and accordingly determining the local and global weights of the features, the features that are most helpful for classification can be more accurately identified. This method helps to optimize the construction of the decision tree and avoid the noise and complexity introduced by irrelevant features. Evaluating and adjusting the weights of each tree according to the out-of-bag error, and adjusting the number of decision trees according to the actual classification effect, enables the model to more precisely adapt to different data characteristics and changes, thereby improving the generalization ability of the model and the accuracy of overall classification. By repeatedly adjusting and optimizing the composition of the decision tree, it is ensured that each iteration strives in the direction of improving classification accuracy, and ultimately the purpose of accurately detecting abnormal traffic is achieved.
[0119] In a possible implementation manner, the network security attack classification and detection method based on random forest further includes:
[0120] S8, aiming to minimize the error between the actual detection result and the true detection result of the abnormal traffic data, optimize the parameters of the improved random forest algorithm through the cuckoo search algorithm.
[0121] It should be noted that the cuckoo search algorithm is an optimization algorithm based on the parasitic breeding behavior of cuckoos in nature, mainly used to solve complex global optimization problems. The algorithm simulates the behavior of cuckoos laying eggs in other bird nests, and explores the search space and selects the optimal solution. In the algorithm, each solution is regarded as a bird nest, which is randomly generated initially, and the cuckoo selects high-quality nests to replace low-quality nests according to the fitness function. The algorithm uses the "discovery probability" and "Lévy flight" strategies for global search and local search, so as to achieve a better balance between exploration and exploitation during the search process. Due to its strong global search ability and adaptability, the cuckoo search algorithm has been widely used in the fields of optimization, machine learning, image processing, etc.
[0122] Furthermore, the cuckoo search algorithm is improved by introducing the mechanism of an adaptive step factor and a dynamic discovery probability. Specifically, the improvement method automatically adjusts the step factor according to the current iteration number and search state, so that the algorithm can conduct extensive exploration in the initial stage of the search, and intensify the fine search for the local optimal solution in the convergence stage.
[0123] In the present invention, the cuckoo search algorithm is renowned for its efficient global search ability and can quickly locate potential optimal solutions in the parameter space. By adaptively adjusting the step size factor and dynamically discovering probabilities, the algorithm can flexibly switch between extensive exploration and fine search, which helps to converge to the optimal parameter configuration more quickly, thereby reducing computational time and resource consumption. By precisely adjusting the parameters of the random forest algorithm, such as the number of decision trees, the method of feature selection, and the depth of the trees, etc., the cuckoo search algorithm helps to significantly improve the classification accuracy and robustness of the random forest model. The optimized parameter settings enable the model to better adapt to the data characteristics and improve the detection ability for abnormal traffic.
[0124] In a possible implementation manner, S8 specifically includes sub-steps S801 and S802:
[0125] S801: With the goal of minimizing the error between the actual detection result and the true detection result of the abnormal traffic data, construct an objective function:
[0126]
[0127] where MSE represents the objective function, θ represents the set of parameters of the improved random forest algorithm, R represents the total number of abnormal traffic data, y r represents the true detection result of the r-th abnormal traffic data, represents the actual detection result of the r-th abnormal traffic data.
[0128] S802: According to the objective function, optimize the parameters of the improved random forest algorithm through the cuckoo search algorithm.
[0129] In the present invention, through the objective function (MSE) constructed with the goal of minimizing the error between the actual detection result and the true detection result, the cuckoo search algorithm can more precisely adjust the parameters of the random forest algorithm. This refined optimization helps to improve the performance of the algorithm on a specific dataset, ensuring high accuracy and low error rate of the classification results. The global optimization ability of the cuckoo search algorithm enables it to explore a wide range of regions in the parameter space, thereby finding the optimal or near-optimal parameter configuration, which directly enhances the performance of the random forest model. The optimized parameter settings make the model more adaptable to the data characteristics and can effectively improve the accuracy and generalization ability of the model.
[0130] In a possible implementation manner, S802 specifically includes sub-steps S8021 to S8028:
[0131] S8021: Initialize the parameters and set the maximum number of iterations of the cuckoo search algorithm.
[0132] S8022: Randomly generate an initial population, where the initial population includes multiple bird nests, and each bird nest represents a set of parameters for a feasible improved random forest algorithm.
[0133] S8023: Use the objective function as the fitness function, calculate the fitness values of each bird nest, and take the bird nest with the maximum fitness value as the current bird nest.
[0134] S8024: According to the current iteration number, adaptively adjust the step size factor, discovery probability, and scaling factor:
[0135]
[0136] where α t represents the step size factor at the t-th iteration, α max represents the maximum value of the step size factor, T represents the maximum number of iterations, α min represents the minimum value of the step size factor, represents the discovery probability of the i-th bird nest at the t-th iteration, P max represents the maximum value of the discovery probability, P min represents the minimum value of the discovery probability, f i t represents the fitness value of the i-th bird nest at the t-th iteration, represents the maximum fitness value at the t-th iteration, the scaling factor of the i-th bird nest at the t-th iteration, γ max represents the maximum value of the scaling factor, γ min represents the minimum value of the scaling factor, represents the minimum fitness value at the t-th iteration.
[0137] S8025: According to the adaptively adjusted step size factor, update the current bird nest through Levy flight operation:
[0138]
[0139] where, represents the i-th bird nest at the (t + 1)-th iteration, represents the i-th bird nest at the t-th iteration, α represents the adaptively adjusted step size factor, represents the dot product operation, levy(β) represents the random step size generated by the Levy flight operation.
[0140] It should be noted that the Lévy flight operation is a random step used in optimization algorithms, characterized by the fact that its step size follows the Lévy distribution, which allows the algorithm to perform long-distance exploration. Compared with the standard random walk, the Lévy flight can more effectively cover a vast search space, helping the algorithm avoid getting trapped in local optima and accelerating the process of finding the global optimum. This method is particularly useful in solving complex optimization problems because it can explore farther regions rather than just local neighborhoods.
[0141] S8026: Calculate the fitness value of the updated bird nest and compare it with the fitness value of the current bird nest. When the fitness value of the updated bird nest is greater than that of the current bird nest, replace the current bird nest with the updated one. When the fitness value of the updated bird nest is less than or equal to that of the current bird nest, keep the current bird nest unchanged.
[0142] S8027: Generate a random number within the range of 0 to 1 and compare it with the adaptively adjusted discovery probability. When the random number is greater than the adaptively adjusted discovery probability, the current bird nest is discovered, and the current bird nest is updated according to the adaptively adjusted scaling factor. When the random number is less than or equal to the adaptively adjusted discovery probability, the current bird nest is not discovered, and the current bird nest remains unchanged.
[0143] Among them, when the random number in S8027 is greater than the adaptively adjusted discovery probability, the current bird nest is discovered, and the current bird nest is updated according to the adaptively adjusted scaling factor. Specifically:
[0144]
[0145] Among them, represents the i-th bird nest at the (t + 1)-th iteration after the current bird nest is discovered, represents the i-th bird nest at the t-th iteration after the current bird nest is discovered, and γ represents the adaptively adjusted scaling factor, and represent two individuals l1 and l2 randomly selected from the population at the t-th iteration after the current bird nest is discovered.
[0146] S8028: Repeat sub-steps S8025 to S8027 until the maximum number of iterations is reached.
[0147] In the present invention, the cuckoo search algorithm effectively finds the optimal or near-optimal parameter settings through the combination of global and local searches, which can significantly improve the performance of the random forest model on a specific dataset, enhancing classification accuracy and robustness. By adaptively adjusting the step size factor and discovery probability, the algorithm can flexibly adjust the search strategy according to the current iteration situation, adapt to various complex optimization scenarios, and improve the search efficiency. By using an effective global optimization algorithm, the number of trial-and-error attempts and time required for parameter adjustment are reduced, thereby reducing the consumption of computing resources. With the optimized parameters, the random forest algorithm can respond to actual network security threats faster and more accurately, detecting and dealing with abnormal behaviors in a timely manner.
[0148] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0149] In the present invention, through the time-based sliding window algorithm, the target network traffic data is determined. When processing large-scale and multi-dimensional network traffic data, the data processing scale is effectively reduced, and the computational complexity is lowered. By extracting multiple network traffic characteristics from the target network traffic data and calculating the average entropy value of the network traffic characteristics, it is preliminarily determined whether the target network traffic data is normal traffic data or abnormal traffic data, enabling a quick response to sudden network attacks. Through the improved random forest algorithm for classification detection, in-depth analysis of feature importance can be carried out, with high classification efficiency and short classification process time, meeting the requirements for real-time classification of network security attacks.
[0150] Refer to the attached Figure 2 illustrates a schematic structural diagram of a network security attack classification and detection system based on random forest provided by the present invention.
[0151] The present invention also provides a network security attack classification and detection system 30 based on random forest, including: a memory 303 and one or more processors 301.
[0152] One or more application programs are stored in the memory 303, and the one or more application programs are adapted to be executed by the one or more processors 301 to implement the network security attack classification and detection method based on random forest described in the method embodiments.
[0153] The network security attack classification and detection system 30 based on random forest includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as through a bus 302.
[0154] The structure of the network security attack classification and detection system 30 based on random forest does not constitute a limitation to the embodiments of the present invention.
[0155] The processor 301 can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 301 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0156] The bus 302 can include a path for transmitting information between the above components. The bus 302 can be a PCI bus, an EISA bus, or the like. The bus 302 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0157] The memory 303 can be a ROM or other type of static storage device that can store static information and instructions, a RAM, or other type of dynamic storage device that can store information and instructions. It can also be an EEPROM, a CD-ROM, or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic storage medium, or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0158] It should be noted that the network security attack classification detection system 30 based on random forest can implement the above-mentioned network security attack classification detection method based on random forest and can achieve the same or similar technical effects. To avoid repetition, the present invention will not be described in detail herein.
[0159] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0160] In the present invention, by using the time-based sliding window algorithm, the target network traffic data is determined. When processing large-scale and multi-dimensional network traffic data, the data processing scale is effectively reduced, and the computational complexity is reduced. By extracting multiple network traffic characteristics from the target network traffic data and calculating the average entropy value of the network traffic characteristics, it is initially determined whether the target network traffic data is normal traffic data or abnormal traffic data, which can quickly respond to sudden network attacks. Through the improved random forest algorithm for classification detection, in-depth analysis of the feature importance can be carried out, the classification efficiency is high, and the classification process takes a short time, which can meet the requirements of real-time classification of network security attacks.
[0161] As described above, this is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the said claims.
[0162] The following points need to be explained:
[0163] (1) The accompanying drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the normal design.
[0164] (2) For clarity, in the accompanying drawings used to describe the embodiments of the present invention, the thickness of the layer or region is enlarged or reduced, that is, these drawings are not drawn according to the actual proportion. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be an intermediate element.
[0165] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0166] As above, this is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A network security attack classification and detection method based on random forest, characterized in that, Including: S1: Collect network traffic data; S2: Determine target network traffic data through a time-based sliding window algorithm; S3: Extract multiple network traffic features from the target network traffic data; S4: Calculate the average entropy value of the network traffic features; S5: Determine whether the average entropy value of the network traffic features is within a preset entropy value range; if so, preliminarily determine the target network traffic data as normal traffic data and return to S1 for continued detection; Otherwise, preliminarily determine the target network traffic data as abnormal traffic data and enter S6; S6: Perform data processing on the abnormal traffic data; S7: Through an improved random forest algorithm, perform classification detection on the processed abnormal traffic data to obtain an accurate detection result and return to S1 for continued detection; Among them, the specific steps of S7 include: S706: Calculate the local weight of the network traffic feature in the decision tree according to the division quality of the nodes of the decision tree by the network traffic feature; S707: Calculate the standardized weight of the decision tree based on the out-of-bag error; S708: Calculate the global weight of the network traffic feature according to the local weight of the network traffic feature in the decision tree and the standardized weight of the decision tree; S709: Sort each network traffic feature according to the global weight of the network traffic feature; S710: According to the sorting result, mark the first u0 network traffic features as important features and construct the first important feature set Γ1; mark the remaining network traffic features as irrelevant features and construct the first irrelevant feature set U1; S711: Discard some network traffic features in the first irrelevant feature set U1 according to the first preset condition to construct the second irrelevant feature set U2; S712: According to the second preset condition, add some network traffic features in the second irrelevant feature set U2 to the first important feature set Γ1 to construct the third irrelevant feature set U3 and the second important feature set Γ2; S713: Determine that the change in the number of network traffic features between the third irrelevant feature set U3 and the first irrelevant feature set U1 is Δv, and the change in the number of network traffic features between the second important feature set Γ2 and the first important feature set Γ1 is Δu; S714: Determine the increase or decrease value of the number of decision trees according to the change in the number of network traffic features between the third irrelevant feature set U3 and the first irrelevant feature set U1 and the change in the number of network traffic features between the second important feature set Γ2 and the first important feature set Γ1; S715: Repeat the foregoing sub-steps until the number of network traffic features in the third irrelevant feature set is less than the preset feature number, and then end the algorithm to obtain the accurate detection result.
2. The network security attack classification and detection method based on random forest according to claim 1, characterized in that The specific steps of S2 include: S201: Insert the network traffic data into the sliding window; S202: Determine whether the collection time of the network traffic data is greater than the preset duration of the sliding window; if so, remove the network traffic data in the sliding window; otherwise, retain the network traffic data in the sliding window and use the retained network traffic data as the target network traffic data.
3. The network security attack classification and detection method based on random forest according to claim 1, characterized in that, The network traffic characteristics specifically include: source IP address, destination IP address, source port, destination port, transport protocol, packet size, traffic duration, traffic direction, number of sessions, flag bits, and byte and packet counts in the traffic.
4. The network security attack classification and detection method based on random forest according to claim 1, characterized in that The S4 specifically includes: S401: Calculate the entropy value of each of the network traffic characteristics through the Shannon entropy formula; S402: Perform normalization processing on the entropy values of each of the network traffic characteristics to obtain the normalized entropy values of each of the network traffic characteristics; S403: Calculate the average entropy value of the network traffic characteristics based on the normalized entropy values of each of the network traffic characteristics.
5. The network security attack classification and detection method based on random forest according to claim 1, characterized in that The S6 specifically includes: S601: Perform data cleaning on the abnormal traffic data; S602: Perform normalization processing on the abnormal traffic data after data cleaning.
6. The method for classifying and detecting network security attacks based on random forest according to claim 1, characterized in that, The S7 further includes: S701: Construct an initial forest composed of multiple decision trees, and the decision tree includes multiple nodes; S702: Calculate the entropy value of each of the nodes; S703: Randomly select multiple network traffic characteristics from the abnormal traffic data; S704: According to the entropy value of each of the nodes, use the randomly selected network traffic characteristics to divide the nodes to obtain the left child node and the right child node of each node; S705: Calculate the division quality of the network traffic characteristics for the nodes according to the entropy value of the left child node and the entropy value of the right child node. The entropy value of each of the nodes: Among them, E represents the entropy value of the node, and p(c) represents the probability of the class label c in the node; The division quality of the network traffic characteristics for the nodes: Q(i,j) = exp{-(E l + E r )} Among them, Q(i, j) represents the partitioning quality of the j-th network traffic feature for the i-th node, and E l represents the entropy value of the left child node, and E r represents the entropy value of the right child node; The local weight of the network traffic characteristics in the decision tree: where, w kj represents the local weight of the j-th network traffic feature in the k-th decision tree, and Q k (i, j) represents the partition quality of the j-th network traffic feature for the i-th node in the k-th decision tree, and N represents the total number of nodes in the decision tree; The normalized weight of the decision tree: Among them, η k represents the normalized weight of the k-th decision tree, and δ k represents the out-of-bag error of the k-th decision tree, and max represents taking the maximum value; The global weight of the network traffic characteristics: where ω j represents the global weight of the j-th network traffic feature, and K represents the total number of decision trees; The increment or decrement value of the number of decision trees: B new = B0 + λ u △u + λ v △v Among them, B new represents the increment or decrement value of the number of decision trees, B0 represents the initial number of decision trees, and λ u represents the decision coefficient of the change in the number of network traffic features between the second important feature set and the first important feature set, and λ v represents the decision coefficient of the change in the number of network traffic features between the third irrelevant feature set and the first irrelevant feature set.
7. The method for classifying and detecting network security attacks based on random forest according to claim 6, characterized in that, The first preset condition is specifically: ω a <μ n -2σ n Among them, ω a represents the global weight of the a-th network traffic feature in the first set of irrelevant features, and μ n represents the mean of the global weights of all network traffic features in the first set of irrelevant features, and σ n represents the standard deviation of the global weights of all network traffic features in the first set of irrelevant features.
8. The method for classifying and detecting network security attacks based on random forest according to claim 6, characterized in that, The second preset condition is specifically: ω b ≥ min(ω c ) where ω b represents the global weight of the b-th network traffic feature in the second set of irrelevant features, and ω c represents the global weight of the c-th network traffic feature in the first set of important features, and min represents taking the minimum value.
9. The network security attack classification and detection method based on random forest according to claim 1, characterized in that It further includes: S8, with the goal of minimizing the error between the actual detection result and the true detection result of the abnormal traffic data, optimize the parameters of the improved random forest algorithm through the cuckoo search algorithm.
10. A network security attack classification and detection system based on random forest, characterized in that, It includes: A memory and one or more processors; One or more application programs are stored in the memory, and the one or more application programs are adapted to be executed by the one or more processors to implement the random forest-based network security attack classification detection method according to any one of claims 1 to 9.
Citation Information
Patent Citations
A DDoS attack detection and defense method and system in a software-defined network
CN109005157A
Network attack detection system based on pattern recognition
CN118740521A