Network traffic abnormity monitoring method and system in combination with scene awareness
By monitoring the CPU temperature and load of the server in the data center, combining business demand and voltage fluctuations, and using machine learning models to predict traffic, the problem of insufficient accuracy of traffic prediction in the data center network is solved, and more accurate abnormal monitoring and lowering false alarm rates are achieved.
Patent Information
- Application Number
- CN202510733207.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing technology cannot fully consider the dynamic and complex scenario factors of data centers, resulting in insufficient accuracy of network traffic prediction and high false alarm rates.
By combining the scene-aware network traffic anomaly monitoring method, the server's CPU temperature sequence and load sequence are obtained by using sensing monitoring equipment, the service demand and voltage fluctuation ratio are analyzed, the traffic consumption prediction is predicted using machine learning models, the predicted traffic interval is determined, and real-time monitoring and abnormal warning are performed.
It improves the accuracy of network traffic monitoring and early warning, reduces the false alarm rate, and ensures the stable operation of the data center network.
Smart Images

Figure CN120263702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of network security monitoring, and particularly to a method and system for monitoring network traffic anomalies combined with scenario awareness. Background Art
[0002] Data centers have become the key infrastructure to support the operation of various services. A large number of servers in the data center work together, generating complex and dynamically changing network traffic. Accurately monitoring network traffic and promptly detecting anomalies are crucial for ensuring the stable operation of the data center, improving service quality, and reducing operating costs. However, traditional network traffic monitoring methods mainly focus on the direct measurement and statistical analysis of network traffic, often ignoring the dynamic and complex operating environment and business scenarios within the data center. In actual operation, factors such as the performance status of servers (such as CPU temperature and load), the dynamic changes in business requirements, and the stability of power supply (voltage fluctuations) will have a significant impact on network traffic. For example, when the CPU temperature of a server is too high or the load is too large, it may lead to a decline in the processing capacity of the server, thereby affecting the transmission speed and stability of network traffic; a sudden increase or change in business requirements will directly cause a large fluctuation in network traffic; while voltage fluctuations may damage the server hardware and indirectly affect the normal transmission of network traffic, resulting in a high false alarm rate or missed alarms in network traffic monitoring.
[0003] Therefore, in the current related technologies, there are technical problems that it is impossible to comprehensively consider the dynamic and complex scenario factors in the data center, resulting in insufficient accuracy of network traffic prediction and a high false alarm rate. Summary of the Invention
[0004] This application provides a method and system for monitoring network traffic anomalies combined with scenario awareness, solving the technical problems in the prior art that it is impossible to comprehensively consider the dynamic and complex scenario factors in the data center, resulting in insufficient accuracy of network traffic prediction and a high false alarm rate, and achieving the technical effects of improving the accuracy of network traffic monitoring and early warning and reducing the false alarm rate.
[0005] This application provides a method for monitoring network traffic anomalies in combination with scenario awareness. The method includes: using sensing and monitoring devices to continuously monitor several servers in a data center to obtain temperature sequences and load sequences of several CPUs; analyzing to obtain the business demand and voltage fluctuation ratio within a future time window; performing traffic consumption prediction based on several temperature sequences, several load sequences, business demand, and voltage fluctuation ratio to determine several predicted traffic intervals; using the several predicted traffic intervals to monitor the real-time network traffic of several servers within the future time window and give early warnings of anomalies; where the performing traffic consumption prediction based on several temperature sequences, several load sequences, business demand, and voltage fluctuation ratio to determine several predicted traffic intervals includes: randomly selecting a first temperature sequence and a first load sequence of a first server; performing fluctuation analysis based on the first temperature sequence and the first load sequence to obtain a first temperature fluctuation coefficient and a first load fluctuation coefficient; using the first temperature fluctuation coefficient and the first load fluctuation coefficient to call a traffic prediction plugin to perform traffic consumption prediction based on the first temperature sequence, the first load sequence, business demand, and voltage fluctuation ratio, and output a first predicted traffic; analyzing and determining a first predicted traffic interval based on the first predicted traffic, and adding it to the several predicted traffic intervals.
[0006] In a possible implementation, the method for monitoring network traffic anomalies in combination with scenario awareness further performs the following processing: using sensing and monitoring devices to continuously monitor the CPUs of several servers in a data center to obtain temperature data and load data at consecutive K time nodes, where K is an integer greater than 10; arranging the temperature data and load data at consecutive K time nodes in chronological order to generate a temperature sequence and a load sequence.
[0007] In a possible implementation, the method for monitoring network traffic anomalies in combination with scenario awareness further performs the following processing: retrieving the historical business records of the data center, collecting the historical business data in the same historical time window of the future time window to obtain a sample historical business data set, where the business data is the traffic request volume; calculating the mean value of the business data based on the sample historical business data set, and calculating the maximum deviation ratio based on the mean value of the business data; using the sum of 1 plus the maximum deviation ratio, multiplied by the mean value of the business data, as the business demand.
[0008] In a possible implementation, the method for monitoring network traffic anomalies in combination with scenario awareness further performs the following processing: retrieving the power management records of the data center, collecting the consecutive voltage data in the same historical time window of the future time window to obtain a sample voltage sequence set; performing fluctuation analysis on multiple sample voltage sequences in the sample voltage sequence set respectively, and outputting multiple sample voltage fluctuation ratios, and calculating the mean value to obtain the voltage fluctuation ratio.
[0009] In a possible implementation, the network traffic anomaly monitoring method combining scenario awareness further performs the following processing: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; pre-training a traffic prediction plugin, where the traffic prediction plugin includes Q traffic prediction branches, and the traffic prediction branches are constructed based on machine learning, where Q is an integer greater than 5 and less than 20; analyzing and determining an adaptation selection quantity P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient; randomly selecting P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plugin, respectively predicting the traffic consumption of the first temperature sequence, the first load sequence, the service demand, and the voltage fluctuation ratio, and obtaining a first predicted traffic after calculating the mean value.
[0010] In a possible implementation, the network traffic anomaly monitoring method combining scenario awareness further performs the following processing: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; obtaining the historical maximum fluctuation coefficient of the first server, multiplying the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q and taking the integer to obtain the adaptation selection quantity P.
[0011] In a possible implementation, the network traffic anomaly monitoring method combining scenario awareness further performs the following processing: during the test of the traffic prediction plugin, respectively counting the prediction errors after combining different numbers of traffic prediction branches, and constructing a branch number - prediction error comparison table; using the branch number - prediction error comparison table, matching to obtain a first prediction error according to the adaptation selection quantity P; using the first prediction error to expand the first predicted traffic and outputting a first predicted traffic interval.
[0012] In a possible implementation, the network traffic anomaly monitoring method combining scenario awareness further performs the following processing: using the several predicted traffic intervals to judge the real-time network traffic of several servers in a future time window, and if the corresponding predicted traffic interval is not satisfied, giving a traffic anomaly warning to the server; obtaining the number of abnormal devices and the abnormal device distribution of the abnormally warned server, and calculating the abnormal distribution dispersion degree; if the number of abnormal devices exceeds a preset number threshold and / or the abnormal distribution dispersion degree is less than a preset dispersion threshold, giving a traffic anomaly warning to the data center.
[0013] The present application also provides a network traffic anomaly monitoring system combined with scenario awareness. The system includes: a server monitoring module for continuously monitoring several servers in a data center using sensing monitoring devices to obtain temperature sequences and load sequences of several CPUs; a data analysis module for analyzing and obtaining the business requirements and voltage fluctuation ratios within a future time window; a traffic consumption prediction module for predicting traffic consumption based on several temperature sequences, several load sequences, business requirements, and voltage fluctuation ratios to determine several predicted traffic intervals; and a monitoring and warning module for monitoring and warning of anomalies in the real-time network traffic of several servers within a future time window using the several predicted traffic intervals.
[0014] It is intended to use the network traffic anomaly monitoring method and system combined with scenario awareness proposed in the present application to continuously monitor several servers in a data center using sensing monitoring devices to obtain temperature sequences and load sequences of several CPUs; analyze and obtain the business requirements and voltage fluctuation ratios within a future time window; predict traffic consumption based on several temperature sequences, several load sequences, business requirements, and voltage fluctuation ratios to determine several predicted traffic intervals; and monitor and warn of anomalies in the real-time network traffic of several servers within a future time window using the several predicted traffic intervals. This solves the technical problems in the prior art that it is impossible to comprehensively consider the dynamic and complex scenario factors in the data center, resulting in insufficient accuracy of network traffic prediction and a high false alarm rate, and achieves the technical effects of improving the accuracy of network traffic monitoring and warning and reducing the false alarm rate. Brief Description of the Drawings
[0015] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. Flowcharts are used in the present application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be precisely executed in sequence. On the contrary, according to the need, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0016] Figure 1 It is a schematic flowchart of the network traffic anomaly monitoring method combined with scenario awareness provided by the embodiments of the present application.
[0017] Figure 2 It is a schematic structural diagram of the network traffic anomaly monitoring system combined with scenario awareness provided by the embodiments of the present application.
[0018] Description of the reference numerals: server monitoring module 10, data analysis module 20, traffic consumption prediction module 30, monitoring and warning module 40. Detailed Embodiments
[0019] The above description is only an overview of the technical solution of the present application. In order to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter given.
[0020] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0021] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.
[0022] The embodiment of the present application provides a network traffic anomaly monitoring method combined with scenario awareness, as Figure 1 shown. The method includes: Step S100, using a sensing monitoring device to continuously monitor several servers in the data center, and obtaining temperature sequences and load sequences of several CPUs.
[0023] Preferably, a plurality of servers in the data center are continuously monitored by using a sensing monitoring device, that is, multiple sensors are connected to the servers through hardware interfaces, or the servers are remotely monitored by using network protocols to continuously sense the status information of the hardware of the servers in the data center, including the operation data of the server CPU components, including temperature and load. By continuously monitoring the operation status of the servers, temperature sequences and load sequences of multiple CPUs are obtained. Among them, the CPU temperature sequence contains the CPU temperature data continuously collected at a certain time interval within a period of time and arranged in chronological order. Excessive temperature may cause the CPU performance to decline, errors to occur, or even damage the hardware; the CPU load sequence is a sequence composed of the CPU load data collected at a specific time interval within a period of time. The CPU load represents the busy degree of the CPU in processing tasks and directly reflects the ability and pressure of the server to process services, usually expressed as a percentage. By analyzing the CPU load sequence, the working intensity of the server at different time periods can be understood, the development trend of the service can be predicted, resources can be reasonably allocated, and the resource utilization rate of the entire data center can be improved.
[0024] Further, step S100 further includes step S110 of continuously monitoring the CPUs of a plurality of servers in the data center by using a sensing monitoring device to obtain the temperature data and load data at consecutive K time nodes, where K is an integer greater than 10; step S120 of arranging the temperature data and load data at consecutive K time nodes in chronological order to generate a temperature sequence and a load sequence.
[0025] Preferably, a sensing monitoring device is used to continuously monitor the CPUs of a plurality of servers in the data center, continuously sense the status of the server CPUs, and continuously measure and record the CPU-related data at a certain time interval, including obtaining the temperature data and load data at consecutive K time nodes, where K is an integer greater than 10, such as K can be 15, 20, etc. Assuming that data is collected every 1 minute, the consecutive K time nodes mean that starting from the 1st minute to the Kth minute, the temperature data and load data of the CPU are collected every minute; then the temperature data and load data at consecutive K time nodes are arranged in chronological order to generate a temperature sequence and a load sequence respectively, so as to organize the scattered temperature data and load data collected at different time points into an ordered temperature sequence and load sequence to understand the temperature change trend and load change trend of the server CPU within a period of time.
[0026] Step S200, analyze and obtain the service requirements and voltage fluctuation ratio within the future time window.
[0027] Preferably, historical data related to the operations of the data center over a past period of time are collected, including the number of transactions of various services, data transmission volume, user access volume, etc. These historical data are statistically analyzed to identify the patterns of changes in the volume of operations, including daily peak and trough hours, weekly or monthly periodic change trends, etc., and then the business demand within a future time window is predicted to obtain the business request volume at different times. At the same time, the current operating status of the business is monitored in real time to obtain the real-time traffic, response time, concurrent user number, etc. of the business for dynamically adjusting the business demand. Voltage monitoring devices (such as smart meters) are used to collect the voltage data of the data center, including the incoming voltage and the output voltage, as well as the instantaneous value, effective value, and peak value of the voltage, which are saved as voltage historical record data to accurately record the change of voltage over time. Then, the voltage fluctuation characteristics are analyzed based on the voltage historical record data, such as finding the peak and trough hours of daily voltage fluctuations and analyzing the differences in voltage fluctuations in different seasons, on different weekdays and rest days; the historical voltage record data are further analyzed to predict the voltage fluctuation trend within a future time window and obtain the voltage fluctuation ratio.
[0028] Further, step S200 further includes step S210 of retrieving the historical business records of the data center, collecting the historical business data in the same historical time window of the future time window to obtain a sample historical business data set, where the business data is the traffic request volume; step S220 of calculating the mean value of the business data according to the sample historical business data set and calculating the maximum deviation ratio according to the mean value of the business data; and step S230 of using the sum of 1 plus the maximum deviation ratio and multiplying it by the mean value of the business data as the business demand.
[0029] Preferably, retrieve from the historical business records of the data center to obtain historical business data for the same time period as the future time window. Here, the business data is the traffic request volume. For example, if the future time window is from 9:00 to 10:00, then collect the historical business data from 9:00 to 10:00 yesterday, the same day of last week, or other similar dates in the past to form a sample historical business data set. Conduct statistical analysis on the traffic request volume data in the sample historical business data set and calculate its average value. For example, the sample historical business data set contains 10 traffic request volume data, which are 100, 120, 110, 90, 130, 105, 115, 125, 95, 100 respectively. Add these traffic request volume data and divide by the number of data 10 to get the average value of the business data as (100 + 120 + 110 + 90 + 130 + 105 + 115 + 125 + 95 + 100) ÷ 10 = 109. Then find the minimum and maximum values of the traffic request volume in the sample historical business data set, and calculate their deviation ratios from the average value of the business data respectively. The minimum value is 90 and the maximum value is 130. Then the deviation ratio of the minimum value to the average value is (109 - 90) ÷ 109 ≈ 0.174, and the deviation ratio of the maximum value to the average value is (130 - 109) ÷ 109 ≈ 0.193. Take the larger value as the maximum deviation ratio. Finally, use the sum of 1 plus the maximum deviation ratio, and multiply by the average value of the business data to determine the business demand. The business demand is (1 + 0.193) × 109 ≈ 130.04. By magnifying the average value to estimate the possible business traffic demand within the future time window, a relatively conservative reference value is provided for the resource allocation and management of the data center to cope with the possible business peak.
[0030] Further, step S200 further includes step S240, retrieving the power management records of the data center, collecting the continuous voltage data under the same historical time window as the future time window to obtain a sample voltage sequence set; step S250, respectively performing fluctuation analysis on multiple sample voltage sequences in the sample voltage sequence set, and outputting multiple sample voltage fluctuation ratios, and calculating the average value to obtain the voltage fluctuation ratio.
[0031] Preferably, retrieve from the power management records of the data center to obtain voltage data information on power supply and usage. Similar to the collection of business demand data, collect continuous voltage data for the same time period as the future time window. For example, if the future time window is 8:00 - 9:00, then find multiple continuous voltage data for 8:00 - 9:00 at the same time in the past from the power management records to form multiple sample voltage sequences, which together constitute a sample voltage sequence set. For each sample voltage sequence in the sample voltage sequence set, analyze the change of its voltage value over time and calculate the voltage fluctuation ratio of the sequence. Generally, the voltage fluctuation ratio is obtained by calculating the difference between the maximum and minimum voltage values in the sequence and then dividing by the mean of the sequence. For example, for a sample voltage sequence [220, 222, 218, 225, 215], first calculate the mean as (220 + 222 + 218 + 225 + 215) ÷ 5 = 220, the maximum value is 225, and the minimum value is 215. Then the voltage fluctuation ratio of this sequence is (225 - 215) ÷ 220 ≈ 0.045 (or 4.5%). Finally, average the multiple sample voltage fluctuation ratios to obtain the voltage fluctuation ratio, and then the average fluctuation degree of the voltage within the same past time window can be understood, so as to evaluate the voltage stability and the possible impact on equipment within the future time window.
[0032] Step S300, perform traffic consumption prediction based on several temperature sequences, several load sequences, business demand, and voltage fluctuation ratio, and determine several predicted traffic intervals.
[0033] Preferably, preprocess a number of temperature sequences, a number of load sequences, service demands, and voltage fluctuation ratios, including removing noise and outliers from the temperature sequences, load sequences, service demands, and voltage fluctuation ratio data, and performing normalization to unify different types of data to the same scale. For example, normalize the temperature sequences, load sequences, service demands, and voltage fluctuation ratio data to the [0, 1] interval to avoid affecting the prediction results due to excessive differences in data scales. For instance, a traffic prediction model may be established using machine learning models (such as neural networks, decision trees, support vector machines, etc.). Use the historical temperature sequences, load sequences, service demands, and voltage fluctuation ratio data as input data, and the corresponding historical traffic consumption data as output labels to train the traffic prediction model so that the model can accurately learn the relationship between various factors and traffic consumption. Then, input the predicted values of the temperature sequence, load sequence, service demand, and voltage fluctuation ratio within the future time window into the traffic prediction model, and further calculate and output the predicted result of the traffic consumption within the future time window, which is usually a range, that is, the predicted traffic interval. For example, predict that the network traffic consumption of the data center within one day in the future is between 30TB and 35TB, and each CPU corresponds to a predicted traffic interval. The network traffic usage of multiple CPUs is different, and finally obtain a number of predicted traffic intervals for the network management and resource allocation of the data center.
[0034] Further, step S300 further includes step S310, randomly select the first temperature sequence and the first load sequence of the first server; step S320, perform fluctuation analysis based on the first temperature sequence and the first load sequence to obtain the first temperature fluctuation coefficient and the first load fluctuation coefficient; step S330, use the first temperature fluctuation coefficient and the first load fluctuation coefficient to call the traffic prediction plug-in, and perform traffic consumption prediction based on the first temperature sequence, the first load sequence, the service demand, and the voltage fluctuation ratio, and output the first predicted traffic; step S340, analyze and determine the first predicted traffic interval based on the first predicted traffic, and add it to the number of predicted traffic intervals.
[0035] Preferably, a first server is randomly selected from multiple servers, and the temperature sequence and load sequence of the first server are obtained, named the first temperature sequence and the first load sequence respectively. For example, in a data center with 100 servers, the 23rd server is randomly selected. The sequence composed of the CPU temperature data in the past period (assumed to be 1 hour, with data recorded every minute) is the first temperature sequence, and the sequence composed of the CPU load data is the first load sequence. Then, fluctuation analysis is performed on the first temperature sequence and the first load sequence respectively. Specifically, for the temperature sequence, the average value of the absolute values of the differences between adjacent data points in the temperature sequence is calculated to measure the change range of temperature over time, and then the first temperature fluctuation coefficient is obtained. Similarly, the first load fluctuation coefficient is obtained through analysis and calculation of the first load sequence, which reflects the change degree of the temperature and load of the CPU of the first server over a period of time.
[0036] Preferably, the traffic prediction plugin is a traffic prediction module with a built-in traffic prediction model for predicting traffic consumption based on the input data. Specifically, the traffic prediction plugin is called using the first temperature fluctuation coefficient and the first load fluctuation coefficient, and the first temperature sequence, the first load sequence, the predicted service demand, and the voltage fluctuation ratio are input into the traffic prediction plugin. The traffic prediction model is used for prediction, and the first predicted traffic is output, that is, an estimated value of the network traffic consumption that the server may generate in the future time window based on the relevant data of the first server and the overall service and voltage conditions. Finally, based on the first predicted traffic, a traffic range is further analyzed and determined, that is, the first predicted traffic interval, which represents the possible value range of the traffic. For example, according to the fluctuation situation and experience of historical data, a certain proportion (such as 10%) is fluctuated up and down with the first predicted traffic as the center to determine the interval. If the first predicted traffic is 30TB, the first predicted traffic interval may be [27TB, 33TB]. Then, the first predicted traffic interval is added to several predicted traffic intervals to comprehensively reflect the possible situation of the network traffic consumption in the future time window of the data center.
[0037] Further, step S330 further includes step S331 of obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; step S332 of pre-training a traffic prediction plug-in, where the traffic prediction plug-in includes Q traffic prediction branches, and the traffic prediction branches are constructed based on machine learning, where Q is an integer greater than 5 and less than 20; step S333 of analyzing and determining an adapted selection quantity P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient; and step S334 of randomly selecting P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plug-in, respectively predicting traffic consumption for the first temperature sequence, the first load sequence, the service demand, and the voltage fluctuation ratio, and outputting P predicted traffic volumes, and obtaining a first predicted traffic volume after calculating the mean value.
[0038] Preferably, different weights are assigned to the first temperature fluctuation coefficient and the first load fluctuation coefficient based on historical data, and then the first fluctuation coefficient is calculated by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; pre-training a traffic prediction plug-in, that is, using a large amount of historical data (including temperature sequences, load sequences, service demands, voltage fluctuation ratios, and corresponding actual traffic consumption data, etc.) to train the machine learning model of the traffic prediction plug-in, adjusting the parameters of the model, so that the model can accurately predict traffic consumption according to the input data, where the traffic prediction plug-in includes Q traffic prediction branches, Q is an integer greater than 5 and less than 20, for example, Q can be 10 or 15, etc., and each traffic prediction branch is constructed based on a machine learning model, indicating that each branch is trained based on a machine learning model (such as a neural network, a decision tree, a support vector machine, etc.) for the traffic prediction branch to learn the relationship between the input data and the traffic consumption.
[0039] Preferably, analyze the magnitudes and change trends of the first temperature fluctuation coefficient and the first load fluctuation coefficient to determine the adapted selection quantity P. For example, if both the first temperature fluctuation coefficient and the first load fluctuation coefficient are large, it indicates that the state change of the server is relatively drastic, and more traffic prediction branches are selected for more accurate prediction, and the value of P is large; on the contrary, if both of these fluctuation coefficients are small, the value of P is small; then randomly select P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plug-in, respectively use the first temperature sequence, the first load sequence, the service demand, and the voltage fluctuation ratio as input data, let each branch perform traffic consumption prediction, obtain P predicted traffic volumes, and finally calculate the mean value of these P predicted traffic volumes to obtain a first predicted traffic volume, and the accuracy and reliability of the network traffic prediction of the data center server can be ensured.
[0040] Further, step S333 further includes step A1 of obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; step A2 of obtaining the historical maximum fluctuation coefficient of the first server, multiplying the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q and taking the integer to obtain the adaptation selection quantity P.
[0041] Preferably, the first temperature fluctuation coefficient and the first load fluctuation coefficient are weighted and calculated. For example, a weight α (0 < α < 1) is set to represent the importance degree of the temperature fluctuation coefficient, then 1 - α is the importance degree of the load fluctuation coefficient, and the first fluctuation coefficient is obtained to reflect the overall fluctuation situation of the server; then, in the past operation data records of the first server, the maximum fluctuation coefficient value that has occurred in the past time period is searched for and determined, which reflects the maximum degree of fluctuation of the server; then, the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient is calculated, representing the proportional relationship between the current fluctuation degree of the first server and its historical maximum fluctuation degree, and then multiplied by the number Q of traffic prediction branches in the traffic plug-in and rounded to obtain the adaptation selection quantity P. For example, if the calculation result is 8.3, 8 branches are selected for traffic consumption prediction. If the current fluctuation degree is large (close to or reaching the historical maximum fluctuation degree), all prediction branches are selected for traffic prediction to improve the prediction accuracy; if the current fluctuation degree is small, fewer branches are selected to reduce the calculation amount and resource consumption.
[0042] Further, step S340 further includes step S341 of respectively counting the prediction errors after combining different numbers of traffic prediction branches during the test of the traffic prediction plug-in to construct a branch number - prediction error comparison table; step S342 of using the branch number - prediction error comparison table to match the first prediction error according to the adaptation selection quantity P; step S343 of using the first prediction error to expand the first predicted traffic and output a first predicted traffic interval.
[0043] Preferably, during the test phase of the traffic prediction plugin, since the traffic prediction plugin includes Q (Q is an integer greater than 5 and less than 20) traffic prediction branches constructed based on machine learning, it is necessary to test different combinations of the number of traffic prediction branches. For example, test various combinations of using 1 branch, 2 branches, 3 branches... up to Q branches; for each combination of the number of branches, use this combination to predict the traffic consumption of known historical data (including temperature sequence, load sequence, business requirements, voltage fluctuation ratio, and corresponding actual traffic consumption data), and then compare the predicted traffic value with the actual traffic consumption value. Calculate the prediction error for each branch combination through mean square error, mean absolute error, etc.; finally, organize the different combinations of the number of traffic prediction branches and the corresponding prediction errors into a branch number - prediction error comparison table, which records the corresponding relationship between different branch number combinations and prediction errors. Assume that Q = 10 in the traffic prediction plugin. After testing different combinations of the number of traffic prediction branches using the mean square error (MSE) as the error calculation method, the data shown in Table 1 is obtained: Table 1 Branch Number - Prediction Error Comparison Table Preferably, according to the adaptively selected quantity P, select the prediction error value corresponding to the quantity P in the branch number - prediction error comparison table as the first prediction error. For example, the comparison table records information such as the prediction error is 0.1 when the branch number is 3, and the prediction error is 0.08 when the branch number is 5. If the adaptively selected quantity P = 5, then the first prediction error is 0.08; use the first prediction error to expand the first predicted traffic to more accurately represent the possible range of future traffic consumption. For example, based on the first predicted traffic, float up and down by a certain proportion of the first prediction error value to determine the interval range, and finally obtain the first predicted traffic interval, so as to more comprehensively reflect the possible situation of future traffic consumption, which is convenient for the monitoring and management of network traffic.
[0044] Step S400, use the several predicted traffic intervals to monitor and give early warnings about the real - time network traffic of several servers within the future time window.
[0045] Preferably, for a number of servers within a future time window, their network traffic data is collected in real time, including collecting network traffic information at specific time intervals such as per second, per minute, etc., including the number of bytes uploaded and downloaded, the number of data packets, etc. Then, the network traffic data of each server collected in real time is compared with the corresponding predicted traffic interval to determine whether the real-time network traffic of the server exceeds the corresponding predicted traffic interval. If the real-time network traffic of a certain server exceeds the corresponding predicted traffic interval, it indicates that the network traffic of this server has abnormal fluctuations, such as being attacked by a network, having a software failure resulting in abnormal traffic fluctuations, or a sudden change in business requirements, etc., thereby triggering a corresponding warning mechanism, including displaying prominent warning messages on the monitoring interface, or triggering a sound alarm, etc., so that administrators can pay attention in time and take corresponding measures to handle abnormal situations, such as further checking the server status, troubleshooting the cause of the failure, adjusting the network configuration or increasing server resources, etc., to ensure the stable operation of the network and the normal conduct of business, thereby improving the reliability and stability of the server network.
[0046] Further, step S400 further includes step S410 of using the several predicted traffic intervals to judge the real-time network traffic of several servers within a future time window. If it does not meet the corresponding predicted traffic interval, a traffic anomaly warning is issued for the server; step S420 of obtaining the number of abnormal devices and the distribution of abnormal devices of the server with the anomaly warning, and calculating the anomaly distribution dispersion degree; step S430 of issuing a traffic anomaly warning for the data center if the number of abnormal devices exceeds a preset number threshold and / or the anomaly distribution dispersion degree is less than a preset dispersion threshold.
[0047] Preferably, for a number of servers within a future time window, the network traffic data is obtained in real time, and then the real-time network traffic of each server is compared and judged with its respective corresponding predicted traffic interval. If the real-time network traffic of a certain server does not meet (i.e., exceeds) the corresponding predicted traffic interval, it indicates that the network traffic of this server has abnormal conditions, and a traffic anomaly warning is immediately issued for this server. Then, the number of servers with traffic anomaly warnings is counted, that is, the number of abnormal devices, and the distribution of these servers with anomaly warnings in the data center is determined, such as which cabinets and which network areas they are located in respectively. Then, according to the distribution information of the abnormal devices, statistical quantities such as standard deviation and variance are used to calculate the anomaly distribution dispersion degree. Among them, the dispersion degree is used to measure the dispersion of the abnormal devices in the data center. If the abnormal devices are concentrated in a few cabinets or areas, the dispersion degree is smaller; if the abnormal devices are scattered in various positions in the data center, the dispersion degree is larger.
[0048] Preferably, the preset quantity threshold is a preset value used to measure whether the number of abnormal devices has reached the level that requires an overall early warning for the data center. For example, the preset quantity threshold is 8. When the number of abnormal devices exceeds this threshold (such as 10), it indicates that there are a relatively large number of servers with abnormal traffic in the data center, and there may be relatively big problems. The preset dispersion threshold is a standard value used to measure the dispersion degree of the distribution of abnormal devices. If the abnormal distribution dispersion degree is less than the preset dispersion threshold, it means that the distribution of abnormal devices in the data center is relatively concentrated, and there may be some common factors causing traffic anomalies in these servers. If the number of abnormal devices exceeds the preset quantity threshold and / or the abnormal distribution dispersion degree is less than the preset dispersion threshold, when either or both of these two conditions are met, it indicates that there may be serious problems with the network traffic in the data center, and an early warning for traffic anomalies is issued for the data center to prompt the operation and maintenance personnel to take corresponding measures to investigate and solve the problems to ensure the normal operation of the data center.
[0049] In the foregoing, with reference to Figure 1 a network traffic anomaly monitoring method combined with scenario perception according to an embodiment of the present invention was described in detail. Next, with reference to Figure 2 a network traffic anomaly monitoring system combined with scenario perception according to an embodiment of the present invention will be described.
[0050] The network traffic anomaly monitoring system combined with scenario perception according to an embodiment of the present invention is used to solve the technical problems in the prior art that it is impossible to comprehensively consider the dynamic and complex scenario factors in the data center, resulting in insufficient accuracy of network traffic prediction and a relatively high false alarm rate, and achieves the technical effects of improving the accuracy of network traffic monitoring and early warning and reducing the false alarm rate. As Figure 2 shown, the network traffic anomaly monitoring system combined with scenario perception includes: a server monitoring module 10, a data analysis module 20, a traffic consumption prediction module 30, and a monitoring and early warning module 40.
[0051] The server monitoring module 10 is used to continuously monitor a number of servers in the data center by using sensing monitoring devices to obtain temperature sequences and load sequences of a number of CPUs; the data analysis module 20 is used to analyze and obtain the business requirements and voltage fluctuation ratios within a future time window; the traffic consumption prediction module 30 is used to perform traffic consumption prediction based on a number of temperature sequences, a number of load sequences, business requirements, and voltage fluctuation ratios to determine a number of predicted traffic intervals; the monitoring and early warning module 40 is used to use the number of predicted traffic intervals to monitor and issue early warnings for the real-time network traffic of a number of servers within a future time window.
[0052] Next, the specific configuration of the server monitoring module 10 will be described in detail. The server monitoring module 10 further includes: continuously monitoring the CPUs of several servers in the data center by using sensing monitoring devices, and obtaining temperature data and load data at continuous K time nodes, where K is an integer greater than 10; arranging the temperature data and load data at continuous K time nodes in chronological order to generate a temperature sequence and a load sequence.
[0053] Next, the specific configuration of the data analysis module 20 will be described in detail. The data analysis module 20 further includes: retrieving the historical business records of the data center, collecting the historical business data in the same historical time window of the future time window, and obtaining a sample historical business data set, where the business data is the traffic request volume; calculating the mean value of the business data according to the sample historical business data set, and calculating the maximum deviation ratio according to the mean value of the business data; using the sum of 1 plus the maximum deviation ratio, multiplying by the mean value of the business data as the business demand.
[0054] Next, the specific configuration of the data analysis module 20 will be further described in detail. The data analysis module 20 further includes: retrieving the power management records of the data center, collecting the continuous voltage data in the same historical time window of the future time window, and obtaining a sample voltage sequence set; respectively performing fluctuation analysis on multiple sample voltage sequences in the sample voltage sequence set, and outputting multiple sample voltage fluctuation ratios, and calculating the mean value to obtain the voltage fluctuation ratio.
[0055] Next, the specific configuration of the traffic consumption prediction module 30 will be described in detail. The traffic consumption prediction module 30 further includes: randomly selecting the first temperature sequence and the first load sequence of the first server; performing fluctuation analysis according to the first temperature sequence and the first load sequence to obtain a first temperature fluctuation coefficient and a first load fluctuation coefficient; using the first temperature fluctuation coefficient and the first load fluctuation coefficient, calling a traffic prediction plug-in, and performing traffic consumption prediction according to the first temperature sequence, the first load sequence, the business demand and the voltage fluctuation ratio, and outputting a first predicted traffic volume; analyzing and determining a first predicted traffic volume interval according to the first predicted traffic volume, and adding it to several predicted traffic volume intervals.
[0056] Next, the specific configuration of the traffic consumption prediction module 30 will be further described in detail. The traffic consumption prediction module 30 further includes: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; a pre-trained traffic prediction plug-in, wherein the traffic prediction plug-in includes Q traffic prediction branches, and the traffic prediction branches are constructed based on machine learning, where Q is an integer greater than 5 and less than 20; analyzing and determining an adaptation selection quantity P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient; randomly selecting P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plug-in, respectively performing traffic consumption prediction on the first temperature sequence, the first load sequence, the service demand, and the voltage fluctuation ratio, and outputting P predicted traffic volumes, and obtaining a first predicted traffic volume after calculating the mean value.
[0057] Next, the specific configuration of the traffic consumption prediction module 30 will be further described in detail. The traffic consumption prediction module 30 further includes: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; obtaining the historical maximum fluctuation coefficient of the first server, multiplying the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q and taking the integer to obtain the adaptation selection quantity P.
[0058] Next, the specific configuration of the traffic consumption prediction module 30 will be further described in detail. The traffic consumption prediction module 30 further includes: during the test of the traffic prediction plug-in, respectively counting the prediction errors after combining different numbers of traffic prediction branches, and constructing a branch number - prediction error comparison table; using the branch number - prediction error comparison table, matching the first prediction error according to the adaptation selection quantity P; and expanding the first predicted traffic volume by using the first prediction error to output a first predicted traffic volume interval.
[0059] Next, the specific configuration of the monitoring and warning module 40 will be further described in detail. The monitoring and warning module 40 further includes: using the several predicted traffic volume intervals to judge the real-time network traffic of several servers within a future time window, and if the corresponding predicted traffic volume intervals are not satisfied, giving a traffic anomaly warning to the server; obtaining the number of abnormal devices and the abnormal device distribution of the abnormally warned server, and calculating the abnormal distribution dispersion degree; if the number of abnormal devices exceeds a preset quantity threshold and / or the abnormal distribution dispersion degree is less than a preset dispersion threshold, giving a traffic anomaly warning to the data center.
[0060] The network traffic anomaly monitoring system combining scenario perception provided by the embodiments of the present invention can execute the network traffic anomaly monitoring method combining scenario perception provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0061] Although this application makes various references to certain modules in the system according to the embodiments of this application, however, any number of different modules can be used and run on the user terminal and / or the server. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and are not used to limit the protection scope of the present invention.
[0062] The above specific implementation manners do not constitute a limitation to the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A method for monitoring network traffic anomalies combined with scene perception, characterized in that, The method includes: Using a sensing and monitoring device to continuously monitor several servers in the data center, and obtaining temperature sequences and load sequences of several CPUs; Analyzing and obtaining the business demand and voltage fluctuation ratio within a future time window; Performing traffic consumption prediction based on several temperature sequences, several load sequences, business demand, and voltage fluctuation ratio, and determining several predicted traffic intervals; Using the several predicted traffic intervals to monitor and give early warnings of anomalies for the real-time network traffic of several servers within a future time window; Among them, the performing traffic consumption prediction based on several temperature sequences, several load sequences, business demand, and voltage fluctuation ratio, and determining several predicted traffic intervals includes: Randomly selecting a first temperature sequence and a first load sequence of a first server; Performing fluctuation analysis based on the first temperature sequence and the first load sequence to obtain a first temperature fluctuation coefficient and a first load fluctuation coefficient; Using the first temperature fluctuation coefficient and the first load fluctuation coefficient to call a traffic prediction plugin, and performing traffic consumption prediction based on the first temperature sequence, the first load sequence, business demand, and voltage fluctuation ratio, and outputting a first predicted traffic; Analyzing and determining a first predicted traffic interval according to the first predicted traffic, and adding it to several predicted traffic intervals.
2. The method for monitoring network traffic anomalies in combination with scene perception according to claim 1, wherein Using a sensing and monitoring device to continuously monitor several servers in the data center, and obtaining temperature sequences and load sequences of several CPUs, including: Using a sensing and monitoring device to continuously monitor the CPUs of several servers in the data center, and obtaining temperature data and load data at continuous K time nodes, where K is an integer greater than 10; Arranging the temperature data and load data at continuous K time nodes in chronological order to generate a temperature sequence and a load sequence.
3. The network traffic anomaly monitoring method combined with scenario perception according to claim 1, wherein, Analyzing and obtaining the business demand within a future time window, including: Retrieving the historical business records of the data center, collecting the historical business data in the same historical time window of the future time window, and obtaining a sample historical business data set, where the business data is the traffic request volume; Calculating the mean value of the business data according to the sample historical business data set, and calculating the maximum deviation ratio according to the mean value of the business data; Using the sum of 1 plus the maximum deviation ratio, multiplied by the mean value of the business data, as the business demand.
4. The network traffic anomaly monitoring method combined with scenario awareness according to claim 1, characterized in that, Analyzing and obtaining the voltage fluctuation ratio within a future time window, including: Retrieving the power management records of the data center, collecting the continuous voltage data in the same historical time window of the future time window, and obtaining a sample voltage sequence set; Performing fluctuation analysis on multiple sample voltage sequences in the sample voltage sequence set respectively, outputting multiple sample voltage fluctuation ratios, and calculating the mean value to obtain the voltage fluctuation ratio.
5. The network traffic anomaly monitoring method combined with scene perception according to claim 1, characterized in that, Using the first temperature fluctuation coefficient and the first load fluctuation coefficient to call a traffic prediction plugin, and performing traffic consumption prediction based on the first temperature sequence, the first load sequence, business demand, and voltage fluctuation ratio, including: Weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient to obtain a first fluctuation coefficient; Pre-trained traffic prediction plugin, wherein the traffic prediction plugin includes Q traffic prediction branches, and the traffic prediction branches are constructed based on machine learning, where Q is an integer greater than 5 and less than 20; Analyze and determine the adapted selection quantity P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient; Randomly select P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plugin, and respectively perform traffic consumption prediction on the first temperature sequence, the first load sequence, the service demand, and the voltage fluctuation ratio, and output P predicted traffic volumes, and obtain the first predicted traffic volume after mean calculation.
6. The method for monitoring network traffic anomalies combined with scenario awareness according to claim 5, wherein Analyze and determine the adapted selection quantity P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient, including: Obtain the first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; Obtain the historical maximum fluctuation coefficient of the first server, multiply the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q and round it to obtain the adapted selection quantity P.
7. The method for monitoring network traffic anomalies combining scenario awareness according to claim 5, characterized in that, Analyze and determine the first predicted traffic volume interval according to the first predicted traffic volume, including: During the test of the traffic prediction plugin, respectively count the prediction errors after combining different numbers of traffic prediction branches, and construct a branch number - prediction error comparison table; Use the branch number - prediction error comparison table to match the first prediction error according to the adapted selection quantity P; Expand the first predicted traffic volume by using the first prediction error, and output the first predicted traffic volume interval.
8. The network traffic anomaly monitoring method combined with scene perception according to claim 1, wherein Use the several predicted traffic volume intervals to monitor and give early warnings of anomalies for the real-time network traffic of several servers within a future time window, including: Use the several predicted traffic volume intervals to judge the real-time network traffic of several servers within a future time window. If it does not meet the corresponding predicted traffic volume interval, give an early warning of traffic anomalies for the server; Obtain the number of abnormal devices and the abnormal device distribution of the server with early warning of anomalies, and calculate the abnormal distribution dispersion degree; If the number of abnormal devices exceeds the preset quantity threshold and / or the abnormal distribution dispersion degree is less than the preset dispersion threshold, give an early warning of traffic anomalies for the data center.
9. The network traffic anomaly monitoring system combined with scene perception is characterized in that The system is used to implement the network traffic anomaly monitoring method combined with scenario perception according to any one of claims 1 to 8. The system includes: A server monitoring module, which is used to continuously monitor several servers in the data center by using sensing monitoring devices, and obtain the temperature sequences and load sequences of several CPUs; A data analysis module, which is used to analyze and obtain the service demand and voltage fluctuation ratio within a future time window; A traffic consumption prediction module, which is used to perform traffic consumption prediction according to several temperature sequences, several load sequences, service demand, and voltage fluctuation ratio, and determine several predicted traffic volume intervals; A monitoring and early warning module, which is used to monitor and give early warnings of anomalies for the real-time network traffic of several servers within a future time window by using the several predicted traffic volume intervals.
Citation Information
Patent Citations
Flow control method and device
CN107770088A
Equipment monitoring real-time data acquisition method and system
CN119782711A
Service flow analysis method and device
CN120090957A
A performance control system for wireless access point based on thermal condition and method thereof
TWI667929B
Methods and systems for network traffic forecast and analysis
US20120303413A1