Network traffic anomaly monitoring method and system combined with scene perception
By monitoring the CPU temperature and load of the server in the data center, combining business demand and voltage fluctuations, and using machine learning to predict traffic, the problem of insufficient accuracy and high false alarm rate of data center network traffic prediction is solved, and more accurate traffic monitoring and abnormal warning is achieved.
Patent Information
- Application Number
- CN202510733207.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing technology cannot fully consider the dynamic and complex scenario factors of data centers, resulting in insufficient accuracy of network traffic prediction and high false alarm rates.
By combining the scene-aware network traffic anomaly monitoring method, the server's CPU temperature sequence and load sequence are obtained by using sensing monitoring equipment, the service demand and voltage fluctuation ratio are analyzed, the traffic consumption prediction is predicted using machine learning models, the predicted traffic interval is determined, and real-time monitoring and abnormal warning are performed.
It improves the accuracy of network traffic monitoring, reduces the false alarm rate, and ensures the stable operation and resource utilization of the data center network.
Smart Images

Figure CN120263702B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field related to network security monitoring, and specifically to a method and system for monitoring network traffic anomalies in combination with scenario perception. Background Art
[0002] Data centers have become critical infrastructure supporting the operation of various businesses. The large number of servers within data centers working together generates complex and dynamically changing network traffic. Accurately monitoring network traffic and promptly detecting anomalies are crucial for ensuring stable data center operations, improving service quality, and reducing operating costs. However, traditional network traffic monitoring methods primarily focus on direct measurement and statistical analysis of network traffic, often overlooking the dynamic and complex operating environment and business scenarios within data centers. In actual operation, factors such as server performance (such as CPU temperature and load), dynamic changes in business demand, and power supply stability (voltage fluctuations) can significantly impact network traffic. For example, excessively high server CPU temperatures or excessive loads can reduce server processing power, impacting the speed and stability of network traffic transmission. Sudden increases or changes in business demand can directly lead to significant fluctuations in network traffic. Voltage fluctuations can damage server hardware, indirectly affecting the normal transmission of network traffic, resulting in high false positives or missed reports in network traffic monitoring.
[0003] Therefore, at the current stage, relevant technologies have the problem of being unable to fully consider the dynamic and complex scenario factors of data centers, which leads to insufficient accuracy of network traffic prediction and a high false alarm rate. Summary of the Invention
[0004] This application solves the technical problem in the existing technology that the network traffic anomaly monitoring method and system combined with scenario perception cannot fully consider the dynamic and complex scenario factors of the data center, which leads to insufficient accuracy of network traffic prediction and high false alarm rate. It achieves the technical effect of improving the accuracy of network traffic monitoring and early warning and reducing the false alarm rate.
[0005] The present application provides a network traffic anomaly monitoring method combined with scenario perception, the method comprising: using sensor monitoring equipment to continuously monitor several servers in a data center to obtain temperature sequences and load sequences of several CPUs; analyzing and obtaining business demands and voltage fluctuation ratios within a future time window; predicting traffic consumption based on several temperature sequences, several load sequences, business demands, and voltage fluctuation ratios, and determining several predicted traffic intervals; using the several predicted traffic intervals to monitor and issue anomaly warnings for the real-time network traffic of several servers within the future time window; wherein, predicting traffic consumption based on several temperature sequences, several load sequences, business demands, and voltage fluctuation ratios and determining several predicted traffic intervals comprises: randomly selecting a first temperature sequence and a first load sequence of a first server; performing fluctuation analysis based on the first temperature sequence and the first load sequence to obtain a first temperature fluctuation coefficient and a first load fluctuation coefficient; using the first temperature fluctuation coefficient and the first load fluctuation coefficient to call a traffic prediction plug-in, predicting traffic consumption based on the first temperature sequence, the first load sequence, the business demands, and the voltage fluctuation ratio, and outputting a first predicted traffic; determining a first predicted traffic interval based on the first predicted traffic analysis and adding it to several predicted traffic intervals.
[0006] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: using sensor monitoring equipment to continuously monitor the CPUs of several servers in the data center, and obtain temperature data and load data at K consecutive time nodes, where K is an integer greater than 10; arranging the temperature data and load data at K consecutive time nodes in chronological order to generate a temperature sequence and a load sequence.
[0007] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: retrieving historical business records of the data center, collecting historical business data under the same historical time window of the future time window, and obtaining a sample historical business data set, wherein the business data is the traffic request volume; calculating the mean of the business data based on the sample historical business data set, and calculating the maximum deviation ratio based on the mean of the business data; using 1 plus the sum of the maximum deviation ratio, multiplied by the mean of the business data, as the business demand.
[0008] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: retrieving the power management records of the data center, collecting continuous voltage data under the same historical time window as the future time window, and obtaining a sample voltage sequence set; performing fluctuation analysis on multiple sample voltage sequences in the sample voltage sequence set, outputting multiple sample voltage fluctuation ratios, and calculating the average to obtain the voltage fluctuation ratio.
[0009] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; pre-training a traffic prediction plug-in, wherein the traffic prediction plug-in includes Q traffic prediction branches, and the traffic prediction branch is constructed based on machine learning, wherein Q is an integer greater than 5 and less than 20; determining the number of adaptation selections P based on analysis of the first temperature fluctuation coefficient and the first load fluctuation coefficient; randomly selecting P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plug-in, and performing traffic consumption prediction on the first temperature sequence, the first load sequence, the business demand and the voltage fluctuation ratio respectively, outputting P predicted traffic, and obtaining the first predicted traffic after average calculation.
[0010] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; obtaining the historical maximum fluctuation coefficient of the first server, multiplying the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q and rounding it to obtain the adaptive selection number P.
[0011] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: during the testing of the traffic prediction plug-in, the prediction errors after different numbers of traffic prediction branch combinations are counted respectively, and a branch number-prediction error comparison table is constructed; using the branch number-prediction error comparison table, the first prediction error is obtained according to the adaptation selection number P matching; using the first prediction error, the first predicted traffic is expanded to output the first predicted traffic interval.
[0012] In a possible implementation, the network traffic anomaly monitoring method combined with scenario perception also performs the following processing: using the several predicted traffic intervals, the real-time network traffic of several servers in the future time window is judged, and if the corresponding predicted traffic interval is not met, the server is given a traffic anomaly warning; the number of abnormal devices and the distribution of abnormal devices of the anomaly warning server are obtained, and the discreteness of the abnormal distribution is calculated; if the number of abnormal devices exceeds the preset number threshold and / or the discreteness of the abnormal distribution is less than the preset discrete threshold, the data center is given a traffic anomaly warning.
[0013] The present application also provides a network traffic anomaly monitoring system combined with scene perception, and the system includes: a server monitoring module, which is used to use sensor monitoring equipment to continuously monitor several servers in a data center and obtain temperature sequences and load sequences of several CPUs; a data analysis module, which is used to analyze and obtain business demand and voltage fluctuation ratios within a future time window; a traffic consumption prediction module, which is used to predict traffic consumption based on several temperature sequences, several load sequences, business demand and voltage fluctuation ratios, and determine several predicted traffic intervals; a monitoring and early warning module, which is used to use the several predicted traffic intervals to monitor the real-time network traffic of several servers within the future time window and provide early warning of anomalies.
[0014] The proposed scenario-aware network traffic anomaly monitoring method and system proposed in this application utilizes sensor monitoring equipment to continuously monitor several servers in a data center, obtain temperature and load sequences for several CPUs, analyze and obtain business demand and voltage fluctuation ratios within a future time window, predict traffic consumption based on several temperature sequences, several load sequences, business demand, and voltage fluctuation ratios, and determine several predicted traffic intervals; and utilize several predicted traffic intervals to monitor and issue anomaly warnings for the real-time network traffic of several servers within the future time window. This solves the technical problem in the prior art of failing to fully consider the dynamic and complex scenario factors of the data center, which in turn leads to insufficient accuracy in network traffic prediction and a high false alarm rate, thereby achieving the technical effect of improving the accuracy of network traffic monitoring and warning and reducing the false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0016] Figure 1 A flow chart of a method for monitoring network traffic anomalies in combination with scenario awareness provided in an embodiment of the present application.
[0017] Figure 2 A schematic diagram of the structure of a network traffic anomaly monitoring system combined with scenario awareness provided in an embodiment of the present application.
[0018] Description of the accompanying drawings: server monitoring module 10, data analysis module 20, traffic consumption prediction module 30, monitoring and early warning module 40. DETAILED DESCRIPTION
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0020] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0021] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict, and the terms “first\second” involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.
[0022] The present application embodiment provides a method for monitoring network traffic anomalies in combination with scenario perception, such as Figure 1 As shown, the method includes:
[0023] Step S100: Utilize sensor monitoring equipment to continuously monitor several servers in a data center to obtain temperature sequences and load sequences of several CPUs.
[0024] Preferably, sensor monitoring equipment is used to continuously monitor several servers in the data center, that is, multiple sensors are connected to the server through a hardware interface, or the server is remotely monitored using a network protocol, and the status information of the data center server hardware is perceived in real time, including the operating data of the server CPU component, including temperature and load. By continuously monitoring the server operating status, the temperature series and load series of multiple CPUs are obtained, wherein the CPU temperature series includes CPU temperature data continuously collected at certain time intervals over a period of time and arranged in chronological order. Excessive temperature may cause CPU performance degradation, errors or even damage to hardware; the CPU load series is a series composed of CPU load data collected at specific time intervals over a period of time. The CPU load indicates the busyness of the CPU processing tasks, directly reflects the server's ability and pressure to process business, and is usually expressed as a percentage. By analyzing the CPU load series, the working intensity of the server in different time periods can be understood, the development trend of the business can be predicted, resources can be reasonably allocated, and the resource utilization of the entire data center can be improved.
[0025] Furthermore, step S100 also includes step S110, using sensor monitoring equipment to continuously monitor the CPUs of several servers in the data center to obtain temperature data and load data at K consecutive time nodes, where K is an integer greater than 10; step S120, arranging the temperature data and load data at K consecutive time nodes in chronological order to generate a temperature sequence and a load sequence.
[0026] Preferably, a sensor monitoring device is used to continuously monitor the CPUs of several servers in a data center, perceive the server CPU status in real time, and continuously measure and record CPU-related data at certain time intervals, including obtaining temperature data and load data at K consecutive time nodes, where K is an integer greater than 10, such as K can be 15, 20, etc. Assuming that data collection is performed every 1 minute, then the K consecutive time nodes mean that the CPU temperature data and load data are collected once every minute from the 1st minute to the Kth minute; then the temperature data and load data at the K consecutive time nodes are arranged in chronological order to generate a temperature sequence and a load sequence respectively, and then the scattered temperature data and load data collected at different time points are organized into an ordered temperature sequence and load sequence to understand the temperature change trend and load change trend of the server CPU over a period of time.
[0027] Step S200: Analyze and obtain the service demand and voltage fluctuation ratio in the future time window.
[0028] Preferably, historical data related to the data center's business over a period of time is collected, including the number of transactions, data transmission volume, and user access volume for various types of business. This historical data is statistically analyzed to identify patterns in business volume changes, including daily peak and trough periods, and weekly or monthly cyclical trends. This allows forecasts of business demand within future time windows to be made, obtaining the volume of business requests during different time periods. Simultaneously, the current business's operational status is monitored in real time to obtain real-time business traffic, response time, and the number of concurrent users, which are used to dynamically adjust business demand. Voltage monitoring equipment (such as smart meters) is used to collect voltage data from the data center, including incoming and output voltages, as well as instantaneous, effective, and peak values. This data is stored as voltage history data, accurately recording voltage changes over time. Voltage fluctuation characteristics are then analyzed based on the voltage history data, such as identifying peak and trough periods for daily voltage fluctuations and analyzing differences in voltage fluctuations between seasons, weekdays, and weekends. The historical voltage record data is then analyzed to predict voltage fluctuation trends within future time windows and obtain the voltage fluctuation ratio.
[0029] Furthermore, step S200 also includes step S210, retrieving historical business records of the data center, collecting historical business data under the same historical time window of the future time window, and obtaining a sample historical business data set, wherein the business data is the traffic request volume; step S220, calculating the mean of the business data based on the sample historical business data set, and calculating the maximum deviation ratio based on the mean of the business data; step S230, using 1 plus the sum of the maximum deviation ratio, multiplied by the mean of the business data, as the business demand.
[0030] Preferably, a search is performed in the historical business records of the data center to obtain historical business data of the same time period as the future time window, wherein the business data is the traffic request volume. For example, if the future time window is from 9:00 to 10:00, the historical business data from 9:00 to 10:00 yesterday, the same day last week, or other similar dates in the past are collected to form a sample historical business data set; the traffic request volume data in the sample historical business data set are statistically analyzed to calculate the average value. For example, the sample historical business data set contains 10 traffic request volume data, namely 100, 120, 110, 90, 130, 105, 115, 125, 95, and 100. These traffic request volume data are added together and then divided by the number of data 10 to obtain The business data mean is (100 + 120 + 110 + 90 + 130 + 105 + 115 + 125 + 95 + 100) ÷ 10 = 109. Next, find the minimum and maximum traffic request volumes in the sample historical business data set and calculate their deviation ratios from the business data mean. The minimum is 90, and the maximum is 130. Therefore, the deviation ratio of the minimum to the mean is (109 - 90) ÷ 109 ≈ 0.174, and the deviation ratio of the maximum to the mean is (130 - 109) ÷ 109 ≈ 0.193. The larger value is taken as the maximum deviation ratio. Finally, the sum of 1 and the maximum deviation ratio is multiplied by the business data mean to determine the business demand, which is (1 + 0.193) × 109 ≈ 130.04. By amplifying the mean, we can estimate the possible business traffic demand in the future time window, providing a relatively conservative reference value for data center resource allocation and management to cope with possible business peaks.
[0031] Furthermore, step S200 also includes step S240, retrieving the power management records of the data center, collecting continuous voltage data under the same historical time window of the future time window, and obtaining a sample voltage sequence set; step S250, performing fluctuation analysis on multiple sample voltage sequences in the sample voltage sequence set, outputting multiple sample voltage fluctuation ratios, and calculating the average to obtain the voltage fluctuation ratio.
[0032] Preferably, the power management records of the data center are retrieved to obtain the voltage data information of power supply and use, similar to the business demand data collection, and the continuous voltage data within the same time period as the future time window is collected. For example, if the future time window is 8:00-9:00, then multiple continuous voltage data of 8:00-9:00 in the past are found from the power management records to form multiple sample voltage sequences, which together constitute a sample voltage sequence set; for each sample voltage sequence in the sample voltage sequence set, the change of its voltage value over time is analyzed, and the voltage fluctuation ratio of the sequence is calculated. Generally speaking, the voltage fluctuation ratio is calculated by The difference between the maximum and minimum values is divided by the mean of the sequence. For example, if a sample voltage sequence is [220, 222, 218, 225, 215], the mean is first calculated as (220+222+218+225+215)÷5=220. The maximum value is 225 and the minimum value is 215. The voltage fluctuation ratio of the sequence is (225-215)÷220≈0.045 (or 4.5%). Finally, the voltage fluctuation ratio of multiple samples is averaged to obtain the voltage fluctuation ratio. This can then be used to understand the average voltage fluctuation degree in the same time window in the past, so as to facilitate the assessment of voltage stability in future time windows and the possible impact on equipment.
[0033] Step S300 , performing traffic consumption prediction based on a number of temperature sequences, a number of load sequences, business demands, and voltage fluctuation ratios, and determining a number of predicted traffic intervals.
[0034] Preferably, a plurality of temperature sequences, a plurality of load sequences, business demands and voltage fluctuation ratios are preprocessed, including removing noise and outliers in the temperature sequence, load sequence, business demand and voltage fluctuation ratio data, and performing normalization processing to unify different types of data to the same scale, such as normalizing the temperature sequence, load sequence, business demand and voltage fluctuation ratio data to the interval [0, 1] to avoid the influence of the prediction results due to the large difference in data scale. For example, a machine learning model (such as a neural network, a decision tree, a support vector machine, etc.) may be used to establish a traffic prediction model, and the historical temperature sequence, load sequence, business demand and voltage fluctuation ratio data are used as input data, and the corresponding historical traffic consumption data are used as the input data. The data is used as the output label to train the traffic prediction model so that the model can accurately learn the relationship between various factors and traffic consumption. Then, the temperature series prediction value, load series prediction value, business demand prediction value and voltage fluctuation ratio prediction value in the future time window are input into the traffic prediction model, and the traffic consumption prediction result in the future time window is calculated and output. It is usually a range, that is, the predicted traffic interval. For example, it is predicted that the data center network traffic consumption in the next day will be between 30TB and 35TB, and each CPU corresponds to a predicted traffic interval. The network traffic usage of multiple CPUs is different. Finally, several predicted traffic intervals are obtained to facilitate network management and resource allocation of the data center.
[0035] Furthermore, step S300 also includes step S310, randomly selecting the first temperature sequence and the first load sequence of the first server; step S320, performing fluctuation analysis based on the first temperature sequence and the first load sequence to obtain the first temperature fluctuation coefficient and the first load fluctuation coefficient; step S330, using the first temperature fluctuation coefficient and the first load fluctuation coefficient, calling the traffic prediction plug-in, performing traffic consumption prediction based on the first temperature sequence, the first load sequence, business demand and voltage fluctuation ratio, and outputting the first predicted traffic; step S340, determining the first predicted traffic interval based on the first predicted traffic analysis, and adding it to several predicted traffic intervals.
[0036] Preferably, one is randomly selected from multiple servers as the first server, and the temperature sequence and load sequence of the first server are obtained, which are named the first temperature sequence and the first load sequence respectively. For example, there are 100 servers in the data center, and the 23rd server is randomly selected. The sequence composed of its CPU temperature data over the past period of time (assuming 1 hour, and data is recorded once per minute) is the first temperature sequence, and the sequence composed of its CPU load data is the first load sequence; then, fluctuation analysis is performed on the first temperature sequence and the first load sequence respectively. Specifically, for the temperature sequence, the average value of the absolute value of the difference between adjacent data points in the temperature sequence is calculated to measure the temperature change amplitude over time, and then the first temperature fluctuation coefficient is obtained; similarly, the first load sequence is analyzed and calculated to obtain the first load fluctuation coefficient, which reflects the degree of change of the temperature and load of the CPU of the first server over a period of time.
[0037] Preferably, the traffic prediction plug-in is a traffic prediction module with a built-in traffic prediction model for predicting traffic consumption based on input data. Specifically, the traffic prediction plug-in is called using the first temperature fluctuation coefficient and the first load fluctuation coefficient, and the first temperature sequence, the first load sequence, and the predicted business demand and voltage fluctuation ratio are input into the traffic prediction plug-in. The traffic prediction model is used to perform prediction and output a first predicted traffic, that is, an estimate of the network traffic consumption that may be generated by the server in a future time window based on the relevant data of the first server and the overall business and voltage conditions; finally, based on the first predicted traffic, a traffic range is further analyzed to determine, that is, a first predicted traffic interval, which represents the possible value range of the traffic. For example, based on the fluctuation of historical data and experience, the interval is determined with the first predicted traffic as the center and a certain ratio (such as 10%) fluctuating up and down. If the first predicted traffic is 30TB, then the first predicted traffic interval may be [27TB, 33TB]; the first predicted traffic interval is then added to a number of predicted traffic intervals to comprehensively reflect the possible network traffic consumption in the future time window of the data center.
[0038] Furthermore, step S330 also includes step S331, obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; step S332, pre-training a traffic prediction plug-in, wherein the traffic prediction plug-in includes Q traffic prediction branches, and the traffic prediction branches are constructed based on machine learning, wherein Q is an integer greater than 5 and less than 20; step S333, determining the number of adapted selections P based on analysis of the first temperature fluctuation coefficient and the first load fluctuation coefficient; step S334, randomly selecting P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plug-in, and performing traffic consumption prediction on the first temperature sequence, the first load sequence, the business demand and the voltage fluctuation ratio respectively, outputting P predicted traffic, and obtaining the first predicted traffic after average calculation.
[0039] Preferably, different weights are assigned to the first temperature fluctuation coefficient and the first load fluctuation coefficient based on historical data, and then the first fluctuation coefficient is calculated based on the weighted first temperature fluctuation coefficient and the first load fluctuation coefficient; a traffic prediction plug-in is pre-trained, that is, a large amount of historical data (including temperature series, load series, business demand, voltage fluctuation ratio and corresponding actual traffic consumption data, etc.) is used to train the machine learning model of the traffic prediction plug-in, and the parameters of the model are adjusted so that the model can accurately predict traffic consumption based on the input data, wherein the traffic prediction plug-in includes Q traffic prediction branches, Q is an integer greater than 5 and less than 20, for example, Q can be 10 or 15, etc., each traffic prediction branch is constructed based on a machine learning model, indicating that each branch is based on a machine learning model (such as a neural network, a decision tree, a support vector machine, etc.) to train the traffic prediction branch and learn the relationship between input data and traffic consumption.
[0040] Preferably, the size and change trend of the first temperature fluctuation coefficient and the first load fluctuation coefficient are analyzed to determine the number of adaptive selections P. For example, if the first temperature fluctuation coefficient and the first load fluctuation coefficient are both large, it means that the state change of the server is more drastic, and more traffic prediction branches are selected to make more accurate predictions, and the value of P is larger; conversely, if both fluctuation coefficients are small, the value of P is smaller; then, among the Q traffic prediction branches of the traffic prediction plug-in, P traffic prediction branches are randomly selected, and the first temperature sequence, the first load sequence, the business demand and the voltage fluctuation ratio are respectively used as input data, and each branch is asked to perform traffic consumption prediction to obtain P predicted traffic, and finally the mean of these P predicted traffic is calculated to obtain the first predicted traffic, and the accuracy and reliability of the network traffic prediction of the data center server can be ensured.
[0041] Furthermore, step S333 also includes step A1, obtaining the first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; step A2, obtaining the historical maximum fluctuation coefficient of the first server, multiplying the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q and rounding it to obtain the adaptive selection number P.
[0042] Preferably, a weighted calculation is performed on the first temperature fluctuation coefficient and the first load fluctuation coefficient. For example, a weight α (0 < α < 1) is set to represent the importance of the temperature fluctuation coefficient, and 1-α is the importance of the load fluctuation coefficient. This yields the first fluctuation coefficient, which reflects the overall fluctuation of the server. The first server's past operational data is then searched and determined to find the maximum fluctuation coefficient value that occurred within a previous time period, reflecting the server's maximum fluctuation. The ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient is then calculated, representing the proportional relationship between the current fluctuation level of the first server and its historical maximum fluctuation level. This ratio is then multiplied by the number of traffic prediction branches Q in the traffic plug-in and rounded to the nearest integer to obtain the number of adaptive selections P. For example, if the calculated result is 8.3, then eight branches are selected for traffic consumption prediction. If the current fluctuation level is high (approaching or reaching the historical maximum fluctuation level), all prediction branches are selected for traffic prediction to improve prediction accuracy. If the current fluctuation level is low, fewer branches are selected to reduce computational effort and resource consumption.
[0043] Furthermore, step S340 also includes step S341, during the test process of the traffic prediction plug-in, respectively counting the prediction errors after combining different numbers of traffic prediction branches, and constructing a branch number-prediction error comparison table; step S342, using the branch number-prediction error comparison table, according to the adaptation selection number P matching to obtain the first prediction error; step S343, using the first prediction error to expand the first predicted traffic, and outputting the first predicted traffic interval.
[0044] Preferably, during the testing phase of the traffic prediction plug-in, since the traffic prediction plug-in contains Q (Q is an integer greater than 5 and less than 20) traffic prediction branches built based on machine learning, it is necessary to test different numbers of traffic prediction branch combinations, for example, test various combinations using 1 branch, 2 branches, 3 branches... up to Q branches; for each combination of the number of branches, use the combination to predict traffic consumption for known historical data (including temperature series, load series, business demand, voltage fluctuation ratio and corresponding actual traffic consumption data), and then compare the predicted traffic value with the actual traffic consumption value, and calculate the prediction error for each branch combination through mean square error, mean absolute error, etc.; finally, organize the different numbers of traffic prediction branch combinations and the corresponding prediction errors to construct a branch number-prediction error comparison table, which records the corresponding relationship between different branch number combinations and prediction errors. Assuming that Q=10 in the traffic prediction plug-in, using mean square error (MSE) as the error calculation method, after testing different numbers of traffic prediction branch combinations, the data shown in Table 1 are obtained:
[0045] Table 1 Branch number-prediction error comparison table
[0046]
[0047] Preferably, the prediction error value corresponding to the number P is selected in the branch number-prediction error comparison table according to the adaptation selection number P as the first prediction error. For example, the comparison table records information such as the prediction error is 0.1 when the number of branches is 3, and the prediction error is 0.08 when the number of branches is 5. If the adaptation selection number P=5, the first prediction error is 0.08; the first predicted traffic is expanded using the first prediction error to more accurately represent the possible range of future traffic consumption. For example, on the basis of the first predicted traffic, the first prediction error value is floated up and down by a certain proportion to determine the interval range, and finally the first predicted traffic interval is obtained, thereby more comprehensively reflecting the possible situation of future traffic consumption, so as to facilitate the monitoring and management of network traffic.
[0048] Step S400: using the plurality of predicted traffic intervals, monitor the real-time network traffic of a plurality of servers within a future time window and issue an abnormality warning.
[0049] Preferably, for several servers within the future time window, their network traffic data is collected in real time, including collecting network traffic information within specific time intervals such as every second and every minute, including the number of bytes uploaded and downloaded, the number of data packets, etc., and then the network traffic data of each server collected in real time is compared with the corresponding predicted traffic interval to determine whether the real-time network traffic of the server exceeds the corresponding predicted traffic interval. If the real-time network traffic of a server exceeds the corresponding predicted traffic interval, it means that the network traffic of the server has abnormal fluctuations, such as network attacks, software failures causing abnormal traffic fluctuations, or sudden changes in business needs, etc., and then triggering corresponding early warning mechanisms, including displaying eye-catching warning information on the monitoring interface, or triggering sound alarms, etc., so that administrators can pay attention in time and take corresponding measures to deal with abnormal situations, such as further checking the server status, troubleshooting the cause of the failure, adjusting the network configuration or increasing server resources, etc., to ensure the stable operation of the network and the normal operation of the business, thereby improving the reliability and stability of the server network.
[0050] Furthermore, step S400 also includes step S410, using the several predicted traffic intervals to judge the real-time network traffic of several servers in the future time window. If the corresponding predicted traffic interval is not met, a traffic abnormality warning is issued to the server; step S420, obtaining the number of abnormal devices and the distribution of abnormal devices of the abnormal warning server, and calculating the abnormal distribution discreteness; step S430, if the number of abnormal devices exceeds the preset number threshold and / or the abnormal distribution discreteness is less than the preset discrete threshold, a traffic abnormality warning is issued to the data center.
[0051] Preferably, network traffic data is acquired in real time for several servers within a future time window, and the real-time network traffic of each server is then compared with its corresponding predicted traffic interval. If the real-time network traffic of a server does not meet (i.e., exceeds) the corresponding predicted traffic interval, it indicates that the network traffic of that server has experienced an anomaly, and a traffic anomaly warning is immediately issued to that server. The number of servers that have experienced traffic anomaly warnings, i.e., the number of abnormal devices, is then counted, and the distribution of these abnormal warning servers in the data center is determined, such as which cabinets and network areas they are located in. Based on the distribution information of the abnormal devices, the dispersion of the abnormal distribution is calculated using statistics such as standard deviation and variance. The dispersion is used to measure the dispersion of the abnormal devices in the data center. If the abnormal devices are concentrated in a few cabinets or areas, the dispersion is small; if the abnormal devices are scattered throughout the data center, the dispersion is large.
[0052] Preferably, the preset quantity threshold is a pre-set value used to measure whether the number of abnormal devices has reached a level that requires an early warning for the entire data center. For example, the preset quantity threshold is 8. When the number of abnormal devices exceeds this threshold (such as 10), it indicates that there are a large number of servers with abnormal traffic in the data center, and there may be major problems; the preset discrete threshold is a standard value used to measure the degree of discreteness of the distribution of abnormal devices. If the abnormal distribution discreteness is less than the preset discrete threshold, it means that the distribution of abnormal devices in the data center is relatively concentrated, and there may be some common factors that cause traffic abnormalities on these servers; if the number of abnormal devices exceeds the preset quantity threshold and / or the abnormal distribution discreteness is less than the preset discrete threshold, when any one or both of these two conditions are met, it indicates that there may be serious problems with the network traffic of the data center, and a traffic abnormality early warning is issued to the data center, prompting the operation and maintenance personnel to take corresponding measures to investigate and solve the problem to ensure the normal operation of the data center.
[0053] In the above, refer to Figure 1 The network traffic anomaly monitoring method combined with scene perception according to an embodiment of the present invention is described in detail. Figure 2 A network traffic anomaly monitoring system combined with scenario awareness according to an embodiment of the present invention is described.
[0054] The network traffic anomaly monitoring system combined with scene perception according to the embodiment of the present invention is used to solve the technical problem that the existing technology cannot fully consider the dynamic and complex scene factors of the data center, which leads to insufficient accuracy of network traffic prediction and high false alarm rate, and achieves the technical effect of improving the accuracy of network traffic monitoring and early warning and reducing the false alarm rate. Figure 2 As shown, the network traffic anomaly monitoring system combined with scenario perception includes: a server monitoring module 10, a data analysis module 20, a traffic consumption prediction module 30, and a monitoring and early warning module 40.
[0055] The server monitoring module 10 is used to continuously monitor several servers in the data center using sensor monitoring equipment to obtain temperature sequences and load sequences of several CPUs; the data analysis module 20 is used to analyze and obtain business demands and voltage fluctuation ratios within a future time window; the traffic consumption prediction module 30 is used to predict traffic consumption based on several temperature sequences, several load sequences, business demands and voltage fluctuation ratios, and determine several predicted traffic intervals; the monitoring and early warning module 40 is used to monitor the real-time network traffic of several servers within the future time window using the several predicted traffic intervals and to provide abnormal early warnings.
[0056] The specific configuration of the server monitoring module 10 will be described in detail below. The server monitoring module 10 further includes: utilizing sensor monitoring equipment to continuously monitor the CPUs of several servers in the data center, obtaining temperature and load data at K consecutive time points, where K is an integer greater than 10; and arranging the temperature and load data at the K consecutive time points in chronological order to generate temperature and load sequences.
[0057] The specific configuration of the data analysis module 20 will be described in detail below. The data analysis module 20 further includes: retrieving historical service records of the data center, collecting historical service data within the same historical time window as the future time window, and obtaining a sample historical service data set, where the service data is traffic request volume; calculating the service data mean based on the sample historical service data set, and calculating the maximum deviation ratio based on the service data mean; and multiplying the sum of 1 and the maximum deviation ratio by the service data mean as the service demand.
[0058] The specific configuration of the data analysis module 20 will be described in detail below. The data analysis module 20 further includes: retrieving the data center's power management records, collecting continuous voltage data from the same historical time window as the future time window, and obtaining a sample voltage sequence set; performing fluctuation analysis on each of the multiple sample voltage sequences in the sample voltage sequence set, outputting multiple sample voltage fluctuation ratios, and calculating the average of the voltage fluctuation ratios.
[0059] The specific configuration of the traffic consumption prediction module 30 will be described in detail below. The traffic consumption prediction module 30 further includes: randomly selecting a first temperature sequence and a first load sequence of a first server; performing a fluctuation analysis based on the first temperature sequence and the first load sequence to obtain a first temperature fluctuation coefficient and a first load fluctuation coefficient; using the first temperature fluctuation coefficient and the first load fluctuation coefficient, calling a traffic prediction plug-in to perform traffic consumption prediction based on the first temperature sequence, the first load sequence, business demand, and the voltage fluctuation ratio, and outputting a first predicted traffic; determining a first predicted traffic interval based on the first predicted traffic analysis, and adding the first predicted traffic interval to a plurality of predicted traffic intervals.
[0060] The specific configuration of the traffic consumption prediction module 30 will be described in detail below. The traffic consumption prediction module 30 further includes: obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; a pre-trained traffic prediction plug-in, wherein the traffic prediction plug-in includes Q traffic prediction branches, and the traffic prediction branch is constructed based on machine learning, wherein Q is an integer greater than 5 and less than 20; determining the number of adapted selections P based on the analysis of the first temperature fluctuation coefficient and the first load fluctuation coefficient; randomly selecting P traffic prediction branches from the Q traffic prediction branches of the traffic prediction plug-in, and performing traffic consumption prediction on the first temperature sequence, the first load sequence, the business demand and the voltage fluctuation ratio respectively, outputting P predicted traffic, and obtaining the first predicted traffic after the mean calculation.
[0061] The specific configuration of the traffic consumption prediction module 30 will be described in detail below. The traffic consumption prediction module 30 further includes: obtaining a first fluctuation coefficient based on the weighted first temperature fluctuation coefficient and the first load fluctuation coefficient; obtaining the historical maximum fluctuation coefficient of the first server, multiplying the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient by Q, and rounding to obtain the adaptive selection quantity P.
[0062] The specific configuration of the traffic consumption prediction module 30 will be described in detail below. The traffic consumption prediction module 30 further includes: during the traffic prediction plug-in test process, respectively counting the prediction errors of different numbers of traffic prediction branch combinations to construct a branch number-prediction error comparison table; using the branch number-prediction error comparison table to obtain a first prediction error based on the adaptive selection number P; and using the first prediction error to expand the first predicted traffic flow and output a first predicted traffic flow interval.
[0063] The specific configuration of monitoring and early warning module 40 will be described in detail below. Monitoring and early warning module 40 further includes: utilizing the plurality of predicted traffic intervals to determine the real-time network traffic of a plurality of servers within a future time window; issuing a traffic anomaly warning for the server if the corresponding predicted traffic interval is not met; obtaining the number of abnormal devices and the distribution of abnormal devices on the abnormal warning server, and calculating the dispersion of the abnormal distribution; and issuing a traffic anomaly warning for the data center if the number of abnormal devices exceeds a preset number threshold and / or the dispersion of the abnormal distribution is less than a preset dispersion threshold.
[0064] The network traffic anomaly monitoring system combined with scenario awareness provided by an embodiment of the present invention can execute the network traffic anomaly monitoring method combined with scenario awareness provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0065] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.
[0066] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A network traffic anomaly monitoring method combined with scenario awareness is characterized by: The method comprises: Use sensor monitoring equipment to continuously monitor several servers in the data center and obtain temperature and load sequences of several CPUs; Analyze and obtain business demand and voltage fluctuation ratio within the future time window; Traffic consumption is predicted based on several temperature sequences, several load sequences, business demands, and voltage fluctuation ratios, and several predicted traffic intervals are determined. Using the predicted traffic intervals, monitor the real-time network traffic of several servers within a future time window and provide anomaly warnings; The method of predicting traffic consumption based on several temperature sequences, several load sequences, business demands and voltage fluctuation ratios to determine several predicted traffic intervals includes: Randomly selecting a first temperature sequence and a first load sequence of a first server; Performing a fluctuation analysis based on the first temperature sequence and the first load sequence to obtain a first temperature fluctuation coefficient and a first load fluctuation coefficient; Using the first temperature fluctuation coefficient and the first load fluctuation coefficient, calling a traffic prediction plug-in, performing traffic consumption prediction based on the first temperature sequence, the first load sequence, service demand, and the voltage fluctuation ratio, and outputting a first predicted traffic; A first predicted flow interval is determined based on the first predicted flow analysis and added to the plurality of predicted flow intervals.
2. The method for monitoring network traffic anomalies in combination with scenario perception according to claim 1 is characterized in that: Using sensor monitoring equipment, we continuously monitor several servers in the data center and obtain temperature and load sequences of several CPUs, including: Use sensor monitoring equipment to continuously monitor the CPUs of several servers in the data center and obtain temperature and load data at K consecutive time nodes, where K is an integer greater than 10. Arrange the temperature data and load data at K consecutive time nodes in chronological order to generate a temperature sequence and a load sequence.
3. The method for monitoring network traffic anomalies in combination with scenario perception according to claim 1 is characterized in that: Analyze and obtain business needs within the future time window, including: Retrieving historical business records of the data center, collecting historical business data in the same historical time window as the future time window, and obtaining a sample historical business data set, wherein the business data is traffic request volume; Calculating the business data mean based on the sample historical business data set, and calculating the maximum deviation ratio based on the business data mean; The sum of 1 plus the maximum deviation ratio multiplied by the mean of the business data is used as the business demand.
4. The method for monitoring network traffic anomalies in combination with scenario perception according to claim 1 is characterized in that: Analyze and obtain the voltage fluctuation ratio in the future time window, including: Retrieving the power management records of the data center, collecting continuous voltage data in the same historical time window as the future time window, and obtaining a sample voltage sequence set; Fluctuation analysis is performed on multiple sample voltage sequences in the sample voltage sequence set respectively, multiple sample voltage fluctuation ratios are output, and the voltage fluctuation ratio is obtained by average calculation.
5. The method for monitoring network traffic anomalies in combination with scenario perception according to claim 1 is characterized in that: Using the first temperature fluctuation coefficient and the first load fluctuation coefficient, calling the traffic prediction plug-in, and performing traffic consumption prediction based on the first temperature sequence, the first load sequence, service demand, and voltage fluctuation ratio, including: Obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; A pre-trained traffic prediction plug-in, wherein the traffic prediction plug-in includes Q traffic prediction branches, and the traffic prediction branches are constructed based on machine learning, wherein Q is an integer greater than 5 and less than 20; Determine the number of adaptation selections P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient; Among the Q traffic prediction branches of the traffic prediction plug-in, P traffic prediction branches are randomly selected, and traffic consumption prediction is performed on the first temperature sequence, the first load sequence, business demand and voltage fluctuation ratio respectively, and P predicted traffic flows are output. The first predicted traffic flow is obtained after the mean is calculated.
6. The method for monitoring network traffic anomalies in combination with scenario perception according to claim 5 is characterized in that: Analyzing and determining the adaptive selection quantity P according to the first temperature fluctuation coefficient and the first load fluctuation coefficient includes: Obtaining a first fluctuation coefficient by weighting the first temperature fluctuation coefficient and the first load fluctuation coefficient; The historical maximum fluctuation coefficient of the first server is obtained, and the ratio of the first fluctuation coefficient to the historical maximum fluctuation coefficient is multiplied by Q and rounded to obtain the adaptive selection quantity P.
7. The method for monitoring network traffic anomalies in combination with scenario perception according to claim 5 is characterized in that: Determining a first predicted flow interval according to the first predicted flow analysis includes: During the test of the traffic prediction plug-in, the prediction errors of different numbers of traffic prediction branch combinations were counted, and a branch number-prediction error comparison table was constructed. Using the branch quantity-prediction error comparison table, a first prediction error is obtained according to the adaptation selection quantity P; The first predicted flow rate is expanded using the first prediction error to output a first predicted flow rate interval.
8. The method for monitoring network traffic anomalies in combination with scenario awareness according to claim 1, characterized in that: Using the predicted traffic intervals, the real-time network traffic of several servers within the future time window is monitored and abnormal warnings are issued, including: Using the predicted traffic intervals, the real-time network traffic of several servers within a future time window is judged. If the corresponding predicted traffic intervals are not met, a traffic anomaly warning is issued to the server; Obtain the number of abnormal devices and the distribution of abnormal devices from the abnormal warning server, and calculate the dispersion of the abnormal distribution; If the number of abnormal devices exceeds a preset number threshold and / or the dispersion of abnormal distribution is less than a preset discrete threshold, a traffic abnormality warning is issued to the data center.
9. The network traffic anomaly monitoring system combined with scene perception is characterized by: The system is used to implement the network traffic anomaly monitoring method combined with scenario perception according to any one of claims 1 to 8, and the system includes: The server monitoring module is used to continuously monitor several servers in the data center using sensor monitoring equipment to obtain temperature and load sequences of several CPUs; Data analysis module, used to analyze and obtain business demand and voltage fluctuation ratio within the future time window; The traffic consumption prediction module is used to predict traffic consumption based on several temperature sequences, several load sequences, business demand and voltage fluctuation ratio, and determine several predicted traffic intervals; The monitoring and early warning module is used to monitor and issue abnormal early warnings on the real-time network traffic of several servers within a future time window by using the several predicted traffic intervals.
Citation Information
Patent Citations
Flow control method and device
CN107770088A
Equipment monitoring real-time data acquisition method and system
CN119782711A