Abnormal network flow monitoring method based on data analysis
Patent Information
- Application Number
- CN202510648223.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-18
Smart Images

Figure CN120342746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network traffic monitoring, and particularly relates to an abnormal network traffic monitoring method based on data analysis. Background Art
[0002] In today's digital age, the application scope of the Internet has been continuously expanding, the network scale has been growing continuously, and network traffic has shown an explosive growth and is complex and changeable. Abnormal network traffic follows closely, such as distributed denial of service (DDoS) attacks, the hidden spread of malware, the illegal leakage of data, etc., which pose extremely severe challenges to network security in various fields.
[0003] Currently, traditional abnormal network traffic monitoring methods mainly rely on rule matching and threshold detection. Although the rule matching method can identify known attack patterns based on pre-set rules, due to the continuous innovation of network attack technologies, new attack means emerge in an endless stream and change rapidly, and the update of rules often lags behind the evolution of attack methods, resulting in a large number of new attacks being difficult to detect in a timely manner. For example, some attackers use zero-day vulnerabilities for attacks, and since there are no corresponding rules in the rule library, these attacks may run rampant in the network.
[0004] The threshold detection method determines whether the traffic is abnormal by setting a fixed traffic threshold. However, network traffic is not static, and it will change significantly over time and different business scenarios. During certain specific periods, such as peak business hours, normal network traffic may far exceed the usually set threshold and thus be misjudged as abnormal; while during off-peak business hours, even if there is a small-scale abnormal traffic, it may be ignored because it does not reach the threshold. This difficulty in setting thresholds leads to frequent false alarms and missed alarms, greatly increasing the workload of network administrators and seriously affecting the guarantee effect of network security. Summary of the Invention
[0005] The purpose of the present invention is to provide an abnormal network traffic monitoring method based on data analysis, aiming to effectively solve the problems in traditional monitoring methods that rules are difficult to cover new attacks and threshold setting difficulties lead to serious false alarms and missed alarms, and to achieve accurate and efficient monitoring of abnormal network traffic.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] An abnormal network traffic monitoring method based on data analysis, comprising the following steps:
[0008] Set monitoring at key network nodes to collect traffic data in the network in real time;
[0009] Clean the collected original traffic data to remove noise data, duplicate data, and incomplete data; and perform standardization processing on the cleaned data using the Z-score standardization formula to eliminate the dimension difference;
[0010] Extract the statistical features of the standardized data. The statistical features include mean, variance, maximum value, minimum value, deviation, and kurtosis; describe the traffic change trend through an autoregressive moving average model, extract periodic features using Fourier transform, and calculate the traffic change rate;
[0011] Count the number of connections and the mean connection duration between the source IP and the destination IP, calculate the connection frequency distribution, concurrency, and analyze the connection establishment and disconnection patterns; analyze the traffic proportion of different protocol types, extract specific field information in the protocol header, and analyze the data content characteristics for application layer protocols;
[0012] Construct an abnormal traffic monitoring model through machine learning algorithms, input the traffic data collected in real time, preprocessed, and feature-extracted into the trained model, and send an alarm message when abnormal traffic is detected.
[0013] As a further solution of the present invention: The network key nodes include core switches, border routers, and server clusters; The traffic data includes source IP address, destination IP address, port number, traffic size, number of data packets, traffic duration, data packet protocol type, detailed field information in the protocol header, transmission layer window size, network layer survival time value, and data packet timestamp.
[0014] As a further solution of the present invention: It also includes filling in missing values for the collected original traffic data, where:
[0015] Fill in a small number of missing values using the mean value, median value, or most frequent value; Fill in a large number of missing values using interpolation or machine learning algorithm prediction.
[0016] As a further solution of the present invention: Fourier transform is used to analyze the traffic spectrum characteristics to find the main periodic components, and the traffic change rate is calculated by the ratio of the traffic difference between adjacent time points to the time interval.
[0017] As a further solution of the present invention: The abnormal traffic monitoring model constructed through machine learning algorithms includes support vector machine, random forest, or long short-term memory network.
[0018] As a further solution of the present invention: In model training:
[0019] The support vector machine uses kernel functions such as linear kernel, polynomial kernel, and radial basis kernel to map the data to a high-dimensional space;
[0020] The random forest adjusts the number of decision trees, the maximum depth, and the feature selection ratio to optimize performance;
[0021] The long short-term memory network improves the generalization ability by adjusting the number of layers, the number of neurons, the learning rate, the number of iterations, and the dropout rate.
[0022] As a further solution of the present invention: The alarm information includes the source IP address, the destination IP address, the traffic size, and the anomaly type of the abnormal traffic.
[0023] As a further solution of the present invention: The anomaly type is analyzed by combining rule matching and machine learning, and the source is traced by querying the network topology structure, the IP address allocation table, and the log records to determine the source device and the affected scope.
[0024] The beneficial effects of the present invention: The present invention uses a variety of machine learning algorithms for abnormal traffic monitoring, getting rid of the excessive dependence on pre-defined rules. The machine learning model can automatically learn the feature patterns of normal traffic and abnormal traffic. Through continuous learning and updating, it can effectively identify new types of abnormal network traffic, overcoming the defect that traditional rule matching methods are difficult to cover new types of attacks;
[0025] Using multi-dimensional and multi-level traffic features for analysis, fully considering the dynamic change characteristics of network traffic. No longer limited to fixed threshold judgment, the machine learning model can perform adaptive learning and judgment according to the traffic characteristics in different time periods and different business scenarios, greatly reducing the false alarm and missed alarm rates caused by unreasonable threshold settings, and improving the accuracy and reliability of abnormal traffic monitoring;
[0026] Real-time monitoring, in-depth analysis, and accurate tracing of abnormal traffic, and timely issuing of alarms and taking response measures can quickly discover and handle abnormal network traffic, effectively ensuring the safe and stable operation of the network, and reducing potential losses caused by abnormal traffic. At the same time, detailed abnormal event records provide strong data support for the continuous improvement of network security. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present invention will be further described below with reference to the accompanying drawings.
[0028] Figure 1 It is a flowchart of a method for monitoring abnormal network traffic based on data analysis according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0030] Please refer to Figure 1 As shown, the present invention is an abnormal network traffic monitoring method based on data analysis, including the following steps:
[0031] Data collection:
[0032] The present invention adopts a multi-dimensional and all-round data collection strategy to ensure that the obtained network traffic data can comprehensively and truly reflect the actual operation status of the network. At each key node of the network, such as core switches, border routers, server clusters, etc., a variety of technical means such as network probes, switch port mirroring, intrusion detection systems (IDS), and traffic sensors are comprehensively used for data collection.
[0033] The collected data covers rich information, including not only basic information such as source IP address, destination IP address, port number, traffic size, number of data packets, traffic duration, etc., but also further extends to protocol types of data packets, detailed field information of protocol headers, window size of the transport layer, time to live (TTL) value of the network layer, etc. At the same time, in order to better analyze the dynamic changes of network traffic, the timestamp of each data packet will also be recorded, accurate to the millisecond level. Let the i-th traffic data sample collected be x i =[x i1 ,x i2 ,...,x in , where x ij represents the j-th feature of the i-th sample, j ∈ [1, n] and is a positive integer, n is the number of features, and n will continuously increase as the collected information becomes richer.
[0034] To ensure the continuity and stability of data collection, a distributed collection architecture is adopted. Multiple collection nodes are deployed at different positions of the network, and the collected data is transmitted to the data center in real time through high-speed network links for centralized processing. At the same time, in order to prevent loss and damage during data transmission, reliable data transmission protocols and data verification mechanisms are adopted to ensure the integrity and accuracy of the data.
[0035] Data preprocessing:
[0036] The data preprocessing stage is an important foundation for the entire monitoring method, which directly affects the effectiveness of subsequent feature extraction and model training. First, the collected original traffic data is comprehensively cleaned to remove noise data, duplicate data, and incomplete data. Noise data may be generated due to network interference, equipment failures, etc., such as checksum errors in data packets, out-of-order data packets, etc.; duplicate data may be caused by retransmission mechanisms in network transmission, which increases the burden of data processing and affects the accuracy of analysis results; incomplete data may be caused by interruptions or errors during data collection and cannot provide effective information.
[0037] For the cleaned data, standardization processing is performed using the Z-score standardization formula:
[0038]
[0039] where, is the standardized data, μ j is the mean of the j-th feature, and σ j is the standard deviation of the j-th feature. Through standardization processing, data in different formats and magnitudes are uniformly converted into a format convenient for analysis, eliminating the dimensional differences between data and making each feature equally important in model training.
[0040] In addition, to further improve the quality of data, missing value processing is also performed on the data. For a small number of missing values, methods such as mean filling, median filling, or most frequent value filling can be used for processing; for a large number of missing values, interpolation methods or machine learning algorithms are considered for predictive filling. At the same time, to detect and process outliers in the data, statistical analysis-based methods such as box plot method, distance-based method, etc. are used to correct or remove outliers.
[0041] Feature extraction:
[0042] Extracting multiple representative and discriminative traffic features from the preprocessed data is the key to accurately identifying abnormal network traffic. The present invention adopts a multi-dimensional and multi-level feature extraction method to fully explore the potential information in network traffic data.
[0043] Statistical features: (m represents the number of samples)
[0044] Mean: Used to describe the average level of the j-th feature in all samples.
[0045] Variance: Reflects the degree of dispersion of the j-th feature.
[0046] Maximum value: Denote the maximum value of the j-th feature among all samples.
[0047] Minimum value: Denote the minimum value of the j-th feature among all samples.
[0048] Skewness: Used to measure the degree of asymmetry of the distribution of the j-th feature.
[0049] Kurtosis: Reflects the kurtosis of the distribution of the j-th feature.
[0050] Time series feature:
[0051] Assume that the traffic data is arranged in chronological order as Describe its changing trend through the autoregressive moving average model (ARMA). The expression of the ARMA(p,q) model is:
[0052]
[0053] Where, is the autoregressive coefficient, θ i is the moving average coefficient, ε t is white noise. In addition, the periodic features of the traffic will also be extracted. The time series data is transformed into the frequency domain through Fourier transform, its spectral characteristics are analyzed, and the main periodic components of the traffic are found. At the same time, calculate the change rate of the traffic, that is, the ratio of the difference in traffic between adjacent time points to the time interval, to reflect the dynamic change speed of the traffic.
[0054] Connection feature:
[0055] Let the source IP address be s, the destination IP address be d, and the number of connections where I sdi is the indicator function. When the source IP of the i-th sample is s and the destination IP is d, I sdi =1, otherwise I sdi =0. The mean value of the connection duration is where T i is the connection duration of the i-th sample. At the same time, calculate the frequency distribution of the connections, that is, the proportion of the number of connections with different connection durations to the total number of connections, to analyze the stability and activity of the connections. In addition, the concurrency of the connections will also be considered, that is, the number of concurrent connections between the source IP and the destination IP at the same time, as well as the establishment and disconnection modes of the connections, such as whether there are frequent short-term connection and disconnection behaviors.
[0056] Protocol feature:
[0057] Analyze the traffic proportion of different protocol types (such as TCP, UDP, HTTP, FTP, etc.) to understand the usage of various protocols in the network. At the same time, extract specific field information from the protocol headers, such as the combination patterns of the flag bits (SYN, ACK, FIN, etc.) in the TCP protocol, the request methods (GET, POST, etc.) and the characteristics of the request paths in the HTTP protocol. For some application-layer protocols, the characteristics of their data content will also be analyzed, such as the size and type of the transferred files in the File Transfer Protocol (FTP).
[0058] Model training:
[0059] In the present invention, an abnormal traffic monitoring model is constructed by combining multiple machine learning algorithms to give full play to the advantages of different algorithms and improve the accuracy and robustness of the model.
[0060] Support Vector Machine (SVM):
[0061] Taking the support vector machine as an example, let the training data be where y i ∈{-1, 1} represents the label of the sample (-1 for abnormal and 1 for normal). The goal of SVM is to solve the following optimization problem:
[0062]
[0063] The constraint condition is: y i (w T x i +b)≥1 - ξ i , ξ i ≥0;
[0064] where w is the weight vector, b is the bias, ξ i is the slack variable, and C is the penalty parameter. To improve the performance of SVM, a kernel function is used to map the data into a high-dimensional space. Commonly used kernel functions include linear kernel, polynomial kernel, radial basis kernel (RBF), etc. The optimal kernel function and penalty parameter C are selected by the method of cross-validation.
[0065] Random Forest:
[0066] Random Forest is an ensemble learning algorithm composed of multiple decision trees. During the training process, a certain number of samples and features are randomly selected from the training data to construct multiple decision trees. Each decision tree is independently trained and predicted, and finally the final prediction result is determined by voting. The advantages of Random Forest are high accuracy and anti-overfitting ability, and it can also handle high-dimensional data. When constructing a Random Forest, parameters such as the number of decision trees, the maximum depth of each decision tree, and the proportion of feature selection need to be adjusted to obtain the best performance.
[0067] Long Short-Term Memory Network (LSTM):
[0068] LSTM is a special type of Recurrent Neural Network (RNN) that can effectively process sequential data and is suitable for analyzing the temporal characteristics of network traffic. The LSTM network solves the problems of vanishing gradients and exploding gradients in traditional RNNs by introducing a gating mechanism, enabling it to better capture long-term dependencies in sequential data. When training an LSTM network, parameters such as the number of network layers, the number of neurons in each layer, the learning rate, and the number of iterations need to be determined. At the same time, to prevent overfitting, the dropout technique is adopted, randomly ignoring a portion of neurons during training to enhance the generalization ability of the model.
[0069] Real-time Monitoring and Anomaly Detection:
[0070] The preprocessed and feature-extracted traffic data collected in real-time is input into the trained model. For the SVM model, if f(x) = w T x + b < 0, then the current traffic is determined to be abnormal traffic; for the Random Forest model, according to the majority voting principle, if more than half of the decision trees determine that the traffic is abnormal, it is determined to be abnormal traffic; for the LSTM model, the traffic is judged to be abnormal by comparing the predicted value output by the model with a preset threshold.
[0071] When the model prediction result is abnormal, the type and characteristics of the abnormal traffic are further analyzed. A method combining rule matching and machine learning is used to classify the abnormal traffic. For example, for traffic suspected of a DDoS attack, analyze features such as whether the source IP addresses of the traffic are scattered, whether the size and frequency of the traffic increase sharply in a short period of time, etc.; for traffic suspected of malware propagation, analyze whether the data content transmitted contains malicious code features, whether the target addresses connected are known malicious addresses, etc. Combining information such as the source IP address, destination IP address, port number, and protocol type of the traffic, trace the source of the abnormal traffic. By querying information such as the network topology structure, IP address allocation table, and log records, determine the source device of the abnormal traffic and the possible scope of influence.
[0072] Alarms and Responses:
[0073] When abnormal network traffic is detected, the system immediately sends an alarm message. The sending methods of alarm messages are diverse, including text messages, emails, system pop-ups, instant messaging tool messages, etc., ensuring that network administrators can receive the alarm in a timely manner. At the same time, the alarm message contains detailed abnormal traffic information, such as the source IP address, destination IP address, traffic size, and abnormal type of the abnormal traffic, so that administrators can quickly understand the abnormal situation.
[0074] The system takes corresponding measures according to the pre-set response strategies. For mild abnormal traffic, such as occasional small traffic anomalies, records and monitoring can be carried out first to observe its development trend; for moderate abnormal traffic, such as abnormal traffic growth lasting for a period of time, the network bandwidth of the source IP address of the abnormal traffic can be automatically restricted to reduce its impact on the network; for severe abnormal traffic, such as large-scale DDoS attacks, the system will automatically block the network connection of the source IP address of the abnormal traffic and start an emergency backup mechanism, such as switching to a standby server, enabling traffic cleaning equipment, etc., to ensure the normal operation of the network. At the same time, the system will record the detailed information of abnormal events, including the time of occurrence of the anomaly, the processing process, the processing results, etc., to provide data support for subsequent security analysis and improvement.
[0075] The above has described in detail an embodiment of the present invention, but the content described is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. An abnormal network traffic monitoring method based on data analysis, characterized in that It includes the following steps: Set up monitoring at key network nodes to collect traffic data in the network in real time; Clean the collected original traffic data to remove noise data, duplicate data, and incomplete data; And perform standardization processing on the cleaned data using the Z-score standardization formula to eliminate the dimension difference; Extract the statistical features of the standardized data. The statistical features include mean, variance, maximum value, minimum value, deviation, and kurtosis; describe the traffic change trend through an autoregressive moving average model, use Fourier transform to extract periodic features, and calculate the traffic change rate; Count the number of connections and the mean connection duration between the source IP and the destination IP, calculate the connection frequency distribution, concurrency, and analyze the connection establishment and disconnection patterns; analyze the traffic proportion of different protocol types, extract specific field information in the protocol header, and analyze the data content characteristics for application layer protocols; Construct an abnormal traffic monitoring model through machine learning algorithms, input the traffic data collected in real time, preprocessed, and feature-extracted into the trained model, and send an alarm message when abnormal traffic is detected.
2. The abnormal network traffic monitoring method based on data analysis according to claim 1, characterized in that, The key network nodes include core switches, border routers, and server clusters; the traffic data includes source IP address, destination IP address, port number, traffic size, number of data packets, traffic duration, data packet protocol type, detailed field information in the protocol header, transmission layer window size, network layer time-to-live value, and data packet timestamp.
3. The anomaly network traffic monitoring method based on data analysis according to claim 1, characterized in that It also includes filling in missing values for the collected original traffic data, where: For a small number of missing values, use mean filling, median filling, or most frequent value filling; for a large number of missing values, use interpolation or machine learning algorithms for prediction filling.
4. A method for monitoring abnormal network traffic based on data analysis according to claim 1, characterized in that Fourier transform is used to analyze the traffic spectrum characteristics to find the main periodic components, and the traffic change rate is calculated by the ratio of the traffic difference between adjacent time points to the time interval.
5. A method for monitoring abnormal network traffic based on data analysis according to claim 1, characterized in that, The construction of the abnormal traffic monitoring model through machine learning algorithms includes support vector machines, random forests, or long short-term memory networks.
6. The anomaly network traffic monitoring method based on data analysis according to claim 5, characterized in that In model training: Support vector machines use kernel functions such as linear kernel, polynomial kernel, and radial basis kernel to map data to a high-dimensional space; Random forests adjust the number of decision trees, maximum depth, and feature selection ratio to optimize performance; Long short-term memory networks improve generalization ability by adjusting the number of layers, number of neurons, learning rate, number of iterations, and dropout rate.
7. A method for monitoring abnormal network traffic based on data analysis according to claim 1, characterized in that The alarm message includes the source IP address, destination IP address, traffic size, and abnormal type of the abnormal traffic.
8. The abnormal network traffic monitoring method based on data analysis according to claim 7, characterized in that, The abnormal type is analyzed by combining rule matching and machine learning, and the source is traced by querying the network topology structure, IP address allocation table, and log records to determine the source device and the affected scope.
Citation Information
Cited By
Personal information security classification and grading method based on SVM (Support Vector Machine) model
CN120541868A
Method, device and system for monitoring abnormal network traffic based on large model
CN120658509A
Big model-based network abnormal traffic monitoring method, device and system
CN120658509B
Terminal data transmission security management method and device and storage medium
CN120750659A
Terminal data transmission security management method and device and storage medium
CN120750659B