Dynamic data flow monitoring system and method for data center

By deploying a dynamic traffic monitoring system in the data center, using the traffic prediction model and anomaly scoring model, combined with multi-level dynamic thresholds, flexible and efficient monitoring and management of data center network traffic is achieved, and the problem of difficulty in responding to burst traffic fluctuations and static thresholds in the existing technology is solved.

CN119945949AActive Publication Date: 2025-05-06SHANGHAI ATHUB CO LTD

Patent Information

Application Number
CN202510424406.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

When existing server load balancing technologies deal with large-scale complex traffic, they are difficult to quickly respond to burst traffic fluctuations, resulting in server overload or network bottlenecks. In addition, traditional traffic monitoring methods rely on static thresholds and cannot adapt to rapidly changing traffic conditions, making them prone to false alarms or missed alarms.

Method used

A dynamic data traffic monitoring system in the data center is proposed, which can realize flexible and efficient dynamic traffic monitoring through the collaborative work of multiple modules. The system includes a traffic acquisition module, a traffic prediction module, a monitoring threshold adjustment and hierarchical alarm module and a network traffic path allocation module. It uses a traffic prediction model, anomaly scoring model and multi-level dynamic threshold to monitor and adjust traffic allocation in real time.

Benefits of technology

Through accurate traffic prediction and abnormal detection, the system can predict traffic change trends in advance, reduce system burden, improve traffic management efficiency and resource allocation accuracy, reduce false alarms and missed alarms, and ensure the stable operation of the data center network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945949A_ABST
    Figure CN119945949A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of server load balancing, in particular to a dynamic data traffic monitoring system and method for a data center, and aims to realize flexible and efficient dynamic traffic monitoring. The system comprises a flow acquisition module, a flow prediction module, a monitoring threshold adjustment and grading alarm module and a network flow path distribution module. The traffic acquisition module acquires traffic data of each node in real time; the flow prediction module predicts a flow value in the future N minutes through a flow prediction model to obtain a flow prediction value; the monitoring threshold adjustment and grading alarm module obtains an abnormal score of each node through an abnormal scoring model, and sets a multi-stage dynamic threshold in combination with a traffic baseline and a traffic predicted value to carry out real-time monitoring and grading alarm on the traffic data; and the network flow path distribution module generates a flow table rule according to the service priority and the real-time flow state, dynamically distributes the flow path of the service instance, and automatically triggers the cross-node migration of the service instance when the secondary alarm is continuous.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server load balancing, and in particular to a data center dynamic data flow monitoring system and method. Background Art

[0002] As the scale of data centers continues to expand and network applications become increasingly complex, traffic management and load balancing have become key links in ensuring the stability of network services and improving resource utilization efficiency. Network traffic in data centers changes rapidly and is highly dynamic. This dynamicity is not only due to the fluctuation of network service demand, but also affected by various factors such as changes in the external environment and changes in user behavior. In order to ensure that the data center network can operate efficiently and stably, load balancing technology based on real-time monitoring and dynamic scheduling must be adopted. Through accurate traffic monitoring and resource scheduling, data centers can achieve higher performance and resource utilization, reduce network bottlenecks, improve service quality, and reduce system maintenance costs.

[0003] In the field of server load balancing technology, dynamic data traffic monitoring and intelligent scheduling play a vital role. By collecting and analyzing the load status, response time and traffic conditions of each server in real time, the system can accurately identify the load distribution of each server and dynamically adjust the load distribution strategy according to traffic fluctuations. This dynamic adjustment mechanism ensures that data traffic can be efficiently distributed among various servers, avoiding the formation of overloaded servers, thereby improving the resource utilization and overall processing capacity of the server cluster. For example, using software-defined networking (SDN) technology, data flow information can be obtained in real time through a centralized controller, and the load and bandwidth of each node can be dynamically scheduled, so that network traffic can be more flexibly optimized and distributed, improving the performance of the entire system.

[0004] In addition, modern load balancing technology continues to develop in the direction of intelligence and automation. In a complex multi-domain network environment, dynamic load balancing technology achieves optimized management of cross-domain traffic through precise traffic monitoring and adaptive scheduling algorithms. For example, F5 load balancer can automatically monitor the load of each server, and perform intelligent traffic distribution based on real-time network performance data and historical traffic patterns. This not only effectively avoids server overload and response delays, but also can cope with the surge in sudden traffic, ensuring the high availability and stability of the system. Through this dynamic adjustment and intelligent decision-making mechanism, data centers can efficiently manage large-scale concurrent traffic and improve service quality and user experience.

[0005] However, although the existing traffic monitoring technology and load balancing methods have solved the problems of traffic scheduling and resource allocation to a certain extent, there are still many challenges. First, the existing load balancing technology often faces the problem of being unable to quickly respond to sudden traffic fluctuations when dealing with large-scale complex traffic. When network traffic fluctuates sharply in a short period of time, the reaction speed of the existing system may not be able to adjust the load in time, resulting in server overload or network bottleneck, affecting the stability of the overall system. Secondly, traditional traffic monitoring methods usually rely on static threshold settings to determine whether the traffic is abnormal based on preset rules. With the diversification of network environments and traffic patterns, static thresholds can no longer adapt to rapidly changing traffic conditions, and are prone to false positives or false negatives, and cannot effectively capture all types of abnormal traffic. Finally, traditional load balancing solutions often lack flexible adaptive mechanisms and intelligent algorithms in multi-level and multi-domain environments, resulting in the system's load distribution capabilities being limited when facing large-scale concurrent requests. Therefore, a more flexible and intelligent dynamic traffic monitoring and load balancing method is needed that can identify traffic changes in real time and adjust intelligently to improve the traffic management efficiency and resource allocation accuracy of data centers.

[0006] To this end, a data center dynamic data flow monitoring system and method are proposed. Summary of the invention

[0007] The purpose of the present invention is to provide a data center dynamic data flow monitoring system and method, which realizes flexible and efficient dynamic flow monitoring through the collaborative work of multiple modules. The system includes a flow collection module, a flow prediction module, a monitoring threshold adjustment and graded alarm module and a network flow path allocation module; the flow collection module collects the flow data of each node in real time; the flow prediction module predicts the flow value of the next N minutes through the flow prediction model to obtain the flow prediction value; the monitoring threshold adjustment and graded alarm module obtains the abnormal score of each node through the abnormal scoring model, and sets a multi-level dynamic threshold in combination with the flow baseline and the flow prediction value to perform real-time monitoring and graded alarm on the flow data; the network flow path allocation module generates flow table rules according to service priority and real-time flow status, dynamically allocates the flow path of the service instance, and automatically triggers the cross-node migration of the service instance when the secondary alarm persists.

[0008] To achieve the above object, the present invention provides the following technical solutions: A data center dynamic data flow monitoring system, comprising: Traffic collection module, which is used to deploy data collection probes on the SDN switches and physical servers in the data center to collect traffic data of each node in real time; The traffic prediction module is used to establish a traffic baseline corresponding to the network service type based on historical traffic data, and predict the traffic value in the next N minutes through a traffic prediction model to obtain a traffic prediction value; The monitoring threshold adjustment and graded alarm module is used to obtain the anomaly score of each node through the anomaly scoring model, and set a multi-level dynamic threshold in combination with the traffic baseline and the traffic prediction value to perform real-time monitoring and graded alarm on the traffic data; the specific process is: if the current traffic value exceeds the multi-level dynamic threshold for M consecutive detection cycles, a first-level alarm is triggered and pushed to the management platform through the low-latency transmission channel of the HTTP / 3 protocol; if abnormal characteristics of the service access mode are detected at the same time, a second-level alarm is triggered and service traffic mirror analysis is started; The network traffic path allocation module generates flow table rules based on service priority and real-time traffic status, dynamically allocates traffic paths for service instances, and automatically triggers cross-node migration of service instances when the second-level alarm persists.

[0009] Preferably, a data preprocessing unit and a data aggregation unit are also included after the traffic collection module; The data preprocessing unit preprocesses the collected traffic data and stores the preprocessed traffic data in a distributed traffic database; the preprocessing includes: deduplication, format conversion and time synchronization; The data aggregation unit aggregates the traffic data according to a preset time window based on the characteristics of the traffic data.

[0010] Preferably, the flow prediction model comprises: a flow data input unit, a flow feature calculation unit, a flow prediction unit and a flow prediction output unit; The flow data input unit inputs the flow data into the flow prediction model; The flow characteristic calculation unit calculates the characteristics of the flow data to obtain flow characteristics; the flow characteristics include: historical average flow, peak time flow characteristics and flow patterns in different time periods; The traffic prediction unit predicts the traffic value in the next N minutes according to the traffic data and the traffic characteristics through an LSTM network to obtain the traffic prediction value; The traffic prediction output unit outputs the traffic prediction value.

[0011] Preferably, the anomaly scoring model comprises: a data input unit, an attention weight calculation unit, an anomaly scoring calculation unit and an anomaly scoring output unit; The data input unit inputs the historical traffic data of each node, the traffic data and the traffic prediction value into the abnormal scoring model; The attention weight calculation unit uses a self-attention mechanism to extract features from the historical traffic data and the traffic prediction value to obtain corresponding historical attention weights and predicted attention weights; The abnormality score calculation unit calculates the abnormality score of each node based on the historical traffic data, the traffic data, the traffic prediction value, the historical attention weight and the predicted attention weight; The anomaly score output unit outputs the anomaly score.

[0012] Preferably, the formula for the abnormality score is: ; in, Score anomalies; is the number of historical data points; is the historical attention weight; is the current traffic data; For the Historical traffic data; To predict the attention weight; is the traffic prediction value.

[0013] Preferably, the formula of the multi-level dynamic threshold is: ; in, It is a multi-level dynamic threshold; is the series; is the flow baseline; is the flow volatility weight; is the flow volatility; is the prediction error weight; is the traffic prediction value; is the anomaly score weight; Score anomalies; is the external environment weight; Environmental impact factors; is the current environment variable value; is the historical average environmental value; is the activation function.

[0014] Preferably, a data center dynamic data flow monitoring method comprises: Deploy data collection probes on the SDN switches and physical servers in the data center to collect traffic data of each node in real time; Based on historical traffic data, a traffic baseline corresponding to the network service type is established, and the traffic value for the next N minutes is predicted through a traffic prediction model to obtain a traffic prediction value; The anomaly score of each node is obtained through the anomaly scoring model, and the multi-level dynamic threshold is set in combination with the traffic baseline and the traffic prediction value to perform real-time monitoring and graded alarm on the traffic data; the specific process is: if the current traffic value exceeds the multi-level dynamic threshold for M consecutive detection cycles, a first-level alarm is triggered and pushed to the management platform through the low-latency transmission channel of the HTTP / 3 protocol; if abnormal characteristics of the service access mode are detected at the same time, a second-level alarm is triggered and service traffic mirror analysis is started; Generate flow table rules based on service priority and real-time traffic status, dynamically allocate traffic paths for service instances, and automatically trigger cross-node migration of service instances when the second-level alarm persists.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention proposes a traffic prediction model that uses an LSTM model to accurately predict the traffic value in the next N minutes by deeply analyzing the traffic data and combining the traffic characteristics of historical average traffic, peak traffic characteristics and traffic patterns in different time periods. The model can predict the traffic change trend in advance, adapt to the dynamic changes of data center traffic, improve the prediction accuracy, and reduce the system burden caused by sudden changes in traffic. Through accurate traffic prediction, it can better provide support for subsequent abnormal traffic detection and network traffic path allocation, ensuring efficient use of network resources.

[0016] 2. The present invention proposes an anomaly scoring model that combines the self-attention mechanism and traffic prediction results. By integrating historical traffic data, current traffic data, traffic prediction values, and environmental factors, the anomaly score of each node is calculated. The model can accurately determine whether there is an abnormal situation when the traffic fluctuates or is abnormal, and quantify the severity of the anomaly. The anomaly scoring model can more accurately identify abnormal traffic, reduce false positives and false negatives, and improve the accuracy and response speed of the monitoring system. At the same time, the use of the self-attention mechanism improves the model's attention to key data points, further enhances the accuracy of anomaly detection, and provides a basis for the subsequent setting of multi-level dynamic thresholds for real-time monitoring of traffic data.

[0017] 3. The present invention proposes a method for setting a multi-level dynamic threshold. By comprehensively considering the traffic baseline, traffic volatility, traffic prediction error, abnormal score and external environmental factors, a multi-level dynamic threshold is set, which can flexibly adjust the monitoring standard according to the actual traffic conditions. When the traffic exceeds the dynamic threshold, the system can trigger the corresponding graded alarm in time to ensure that the abnormal traffic is responded to and processed in time. The multi-level dynamic threshold can not only adapt to traffic fluctuations, but also cope with changes in different types of network services, reducing the errors and inaccuracies caused by static thresholds. The setting of this multi-level dynamic threshold improves the flexibility and reliability of dynamic traffic monitoring, helps to ensure the stable operation of the data center network, and makes dynamic traffic monitoring more flexible and efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A structural diagram of a data center dynamic data flow monitoring system provided by an embodiment of the present invention; Figure 2 A flow chart of a method for monitoring dynamic data flow in a data center provided by an embodiment of the present invention; Figure 3 A flowchart of anomaly scoring provided by an embodiment of the present invention; Figure 4 A flowchart of multi-level dynamic threshold traffic monitoring and graded alarms provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] With the continuous expansion of data center scale and the increasing complexity of network applications, traffic management and optimization have become key links to ensure the stability of network services and improve resource utilization efficiency. Network traffic in data centers changes rapidly and is highly dynamic. This dynamicity is not only due to the fluctuation of network service demand, but also affected by various factors such as changes in the external environment and changes in user behavior. Therefore, in order to ensure that the data center network can operate efficiently and stably, dynamic traffic monitoring methods must be used to identify traffic anomalies in real time and make timely adjustments and optimizations.

[0021] The present invention proposes a data center dynamic data flow monitoring system, which realizes flexible and efficient real-time monitoring of rapidly changing and highly dynamic network dynamic flow. This system is implemented by a data center dynamic data flow monitoring method. For the specific method flowchart and system structure diagram, please refer to Figure 1 and Figure 2 In order to illustrate that the system and method of the present invention can play a role in flexibly and efficiently monitoring dynamic network traffic, the effectiveness of the present invention will be described from two embodiments below.

[0022] See also Figures 1 to 4 ,The present invention provides a data center dynamic data flow monitoring system and method, and the technical solution is as follows.

[0023] Embodiment 1

[0024] In the embodiment of the present application, the system and method proposed by the present invention are used to describe in detail the real-time monitoring process of dynamic network traffic data. In the embodiment of the present application, the real-time monitoring of dynamic network traffic data is aimed at the real-time monitoring of network traffic data in a large cloud data center A. Figure 1 and Figure 2 The content describes in detail the real-time monitoring process of network traffic data in the large cloud data center; Figure 1 The specific structure of the system proposed by the present invention includes a flow collection module, a flow prediction module, a monitoring threshold adjustment and graded alarm module and a network flow path allocation module. Figure 2 The flowchart of the method proposed in the present invention includes: deploying data collection probes on the SDN switches and physical servers in the data center to collect the traffic data of each node in real time; establishing a traffic baseline corresponding to the network service type based on historical traffic data, and predicting the traffic value in the next N minutes through a traffic prediction model to obtain a traffic prediction value; obtaining the anomaly score of each node through an anomaly scoring model; setting multi-level dynamic thresholds in combination with the traffic baseline and traffic prediction value to perform real-time monitoring and graded alarms on traffic data; collaborating with the SDN controller through the OpenFlow protocol to generate flow table rules based on service priority and real-time traffic status, dynamically allocating traffic paths for service instances, and automatically triggering cross-node migration of service instances when the secondary alarm persists. Figure 1 and Figure 2 The following is a description of the contents: A data center dynamic data flow monitoring system, comprising: Traffic collection module, which is used to deploy data collection probes on the SDN switches and physical servers in the data center to collect traffic data of each node in real time; Specifically, in an embodiment of the present application, the data acquisition probe captures network data packets entering and leaving the switch and physical server by sniffing the network interface; and obtains traffic data by parsing network packet header information; the network packet header information includes IP address, port number, protocol type and traffic size.

[0025] The collection frequency of the data collection probe is dynamically adjusted according to the changes in network traffic.

[0026] Preferably, a data preprocessing unit and a data aggregation unit are also included after the traffic collection module; The data preprocessing unit preprocesses the collected traffic data and stores the preprocessed traffic data in a distributed traffic database; the preprocessing includes: deduplication, format conversion and time synchronization; The data aggregation unit aggregates the traffic data according to a preset time window based on the characteristics of the traffic data.

[0027] Specifically, the deduplication deduplicates and cleans the collected traffic data to remove duplicate or erroneous data; The format conversion converts data in different formats into a unified format for subsequent processing; The time synchronization synchronizes the flow data collected by different probes, ensuring that the flow data of each node can be analyzed in an accurate time sequence.

[0028] In the embodiment of the present application, the traffic collection module collects and processes traffic data in real time by deploying data collection probes on the SDN switches and physical servers in the data center, providing basic data support for subsequent traffic analysis and anomaly monitoring. After preprocessing, aggregation and transmission, the collected traffic data can provide efficient and accurate input for traffic prediction and dynamic threshold adjustment, thereby realizing real-time monitoring and optimization of network traffic.

[0029] Preferably, the traffic prediction module is used to establish a traffic baseline corresponding to the network service type based on historical traffic data, and predict the traffic value in the next N minutes through a traffic prediction model to obtain a traffic prediction value; The flow prediction model includes: a flow data input unit, a flow feature calculation unit, a flow prediction unit and a flow prediction output unit; The flow data input unit inputs the flow data into the flow prediction model; The flow characteristic calculation unit calculates the characteristics of the flow data to obtain flow characteristics; the flow characteristics include: historical average flow, peak time flow characteristics and flow patterns in different time periods; The traffic prediction unit predicts the traffic value in the next N minutes according to the traffic data and the traffic characteristics through an LSTM network to obtain the traffic prediction value; The traffic prediction output unit outputs the traffic prediction value.

[0030] Specifically, the traffic data input unit inputs the traffic data into the traffic prediction model; the traffic data includes: real-time traffic data, network service type information and timestamp information; The traffic characteristics include historical average traffic, peak time traffic characteristics and traffic patterns in different time periods; the historical average traffic includes: the average traffic value over a period of time; The peak period traffic characteristics include the traffic peak value of the network traffic during a specific peak period of each day; The traffic patterns in different time periods include the traffic change characteristics of the morning peak, noon and evening peak, and analyze the periodic changes and sudden fluctuations of the traffic.

[0031] Table 1 shows the output results and performance comparison of the traffic prediction model.

[0032]

[0033] The traffic prediction module of the embodiment of the present application can accurately predict the traffic in the next N minutes through the coordinated work of the traffic data input unit, the traffic feature calculation unit, the traffic prediction unit and the traffic prediction output unit. The module is based on the LSTM network and fully considers the timing characteristics and complexity of the traffic data, which helps to improve the prediction accuracy of the data center network traffic and provide support for subsequent traffic monitoring, anomaly detection and network optimization.

[0034] Preferably, the monitoring threshold adjustment and graded alarm module is used to obtain the abnormal score of each node through an abnormal score model, and the abnormal score model includes: a data input unit, an attention weight calculation unit, an abnormal score calculation unit and an abnormal score output unit; refer to Figure 3 ; The data input unit inputs the historical traffic data of each node, the traffic data and the traffic prediction value into the abnormal scoring model; The attention weight calculation unit uses a self-attention mechanism to extract features from the historical traffic data and the traffic prediction value to obtain corresponding historical attention weights and predicted attention weights; The abnormality score calculation unit calculates the abnormality score of each node based on the historical traffic data, the traffic data, the traffic prediction value, the historical attention weight and the predicted attention weight; The anomaly score output unit outputs the anomaly score.

[0035] Specifically, the attention weight calculation unit uses dot product attention to calculate the degree of influence of historical traffic data on current traffic data to obtain the historical attention weight; calculates the credibility of the traffic prediction value to the current traffic data to obtain the predicted attention weight.

[0036] Table 2 is a table showing the results of the anomaly scoring model based on the attention mechanism.

[0037]

[0038] The anomaly scoring model in this embodiment is based on the self-attention mechanism. It extracts features from historical traffic data and traffic prediction values, calculates historical attention weights and predicted attention weights, and accurately evaluates the degree of anomaly of current traffic data. The model can dynamically adapt to different network environments and service types, improve the accuracy of anomaly detection, and effectively reduce false alarm and missed alarm rates by combining predicted data. The model can capture complex time series features, enhance the ability to identify sudden abnormal traffic, and provide a more intelligent and real-time traffic monitoring and alarm mechanism for data centers.

[0039] Preferably, the formula for the abnormality score is: ; in, Score anomalies; is the number of historical data points; is the historical attention weight; is the current traffic data; For the Historical traffic data; To predict the attention weight; is the traffic prediction value.

[0040] The anomaly scoring formula of the embodiment of the present application achieves an accurate assessment of the degree of anomaly by introducing historical attention weights and predicted attention weights, combining current traffic data, historical traffic data and traffic prediction values. Compared with the traditional scoring mechanism based on simple statistical methods, this formula can fully consider historical trends, short-term fluctuations and future forecast information, thereby improving the sensitivity and reliability to abnormal traffic. Through the self-attention mechanism, the system can automatically identify key data points that affect anomaly scoring, reduce false alarm rates and enhance the ability to respond to sudden anomalies, which helps to make data center traffic management intelligent and refined, and provides a reference for the subsequent setting of multi-level dynamic thresholds for real-time monitoring of traffic data.

[0041] Preferably, a multi-level dynamic threshold is set in combination with the traffic baseline and the traffic prediction value to perform real-time monitoring and graded alarm on the traffic data; the specific process is: if the current traffic value exceeds the multi-level dynamic threshold for M consecutive detection cycles, a first-level alarm is triggered and pushed to the management platform through the low-latency transmission channel of the HTTP / 3 protocol; if abnormal characteristics of the service access mode are detected at the same time, a second-level alarm is triggered and the service traffic mirror analysis is started; refer to Figure 4 ;

[0042] The formula of the multi-level dynamic threshold is: ; in, It is a multi-level dynamic threshold; is the series; is the flow baseline; is the flow volatility weight; is the flow volatility; is the prediction error weight; is the traffic prediction value; is the anomaly score weight; Score anomalies; is the external environment weight; Environmental impact factors; is the current environment variable value; is the historical average environmental value; is the activation function.

[0043] Through the setting of multi-level dynamic thresholds and real-time traffic monitoring, the embodiment of the present application enables the system to flexibly respond to changes in the network environment and improve the accuracy of detecting abnormal traffic. The multi-level dynamic threshold formula combines factors from multiple dimensions to ensure that the dynamic threshold can adapt to different traffic patterns, environmental changes, and service requirements. Compared with the traditional static threshold method, the system can dynamically adjust the alarm threshold according to different service requirements, environmental conditions, and traffic prediction values, avoiding false alarms and missed reports. At the same time, traffic mirroring analysis and cross-node migration functions can also respond to emergencies in a timely manner to ensure the efficient and safe operation of the data center network.

[0044] Preferably, the network traffic path allocation module generates flow table rules according to service priority and real-time traffic status, dynamically allocates traffic paths for service instances, and automatically triggers cross-node migration of service instances when the second-level alarm persists.

[0045] Table 3 shows the SDN-based dynamic traffic path allocation performance test data obtained through multi-level dynamic threshold setting and graded alarm.

[0046] Specifically, in this embodiment, the flow table generation response time is measured by simulating different traffic change scenarios and measuring the total time from the SDN controller receiving the traffic status update information to the generation and delivery of the flow table rules; The path switching success rate is measured by monitoring whether the network traffic can be successfully rerouted according to the flow table rules during the switching process. In the experiment, a certain amount of traffic anomalies are intentionally generated, and then the proportion of traffic that successfully completes the path switching is counted; Cross-node migration latency: By simulating migration operations, we measure the time from the start of migration to the completion of migration, and use the internal monitoring tools of the data center to record the latency of the migration process. The system stability score is obtained by recording the network stability under different circumstances in the experiment.

[0047]

[0048] The embodiment of the present application proposes a data center dynamic data traffic monitoring system, which has the comprehensive capabilities of traffic prediction, anomaly detection and dynamic adjustment of thresholds, and can effectively cope with the challenges in data center traffic monitoring. The traffic collection module obtains the traffic data on the SDN switch and physical server in real time, and uses the traffic prediction module to accurately predict future traffic based on historical data and LSTM network to obtain traffic prediction values. Combined with information such as traffic baseline, traffic volatility, external environmental factors, etc., the monitoring threshold adjustment and graded alarm module uses the anomaly scoring model to intelligently monitor and grade alarms for traffic, and uses the self-attention mechanism to extract the features of historical traffic and predicted traffic, effectively improving the detection accuracy of abnormal traffic. Based on multi-level dynamic thresholds, the system can automatically adjust the monitoring threshold according to the real-time traffic situation, trigger a first-level alarm when the traffic exceeds the standard, and push it to the management platform with low latency through the HTTP / 3 protocol; if the service access mode is detected to be abnormal, it will further trigger a second-level alarm and perform service traffic mirroring analysis. Finally, combined with the OpenFlow protocol, the traffic path allocation module works with the SDN controller to automatically adjust the traffic path according to the traffic priority and traffic prediction value, and trigger the cross-node migration of service instances in the case of a secondary alarm, thereby optimizing the data center's traffic management and improving the system's fault tolerance, efficiency, and resource utilization. The entire system ensures the efficiency and accuracy of real-time monitoring and response of data center network traffic, and improves the stability of the data center system.

[0049] Embodiment 2 In Example 1, the system proposed by the present invention successfully realizes real-time monitoring of dynamic traffic of the data center network. To further verify the effectiveness of the present invention, the network traffic data of another large cloud data center B is also monitored in real time in the present application example.

[0050] In order to better perform subsequent real-time monitoring of network traffic data, large cloud data center B applies a data center dynamic data traffic monitoring method, including the following steps: Deploy data collection probes on the SDN switches and physical servers in the data center to collect traffic data of each node in real time; Preferably, a data preprocessing unit and a data aggregation unit are also included after the traffic collection module; The data preprocessing unit preprocesses the collected traffic data and stores the preprocessed traffic data in a distributed traffic database; the preprocessing includes: deduplication, format conversion and time synchronization; The data aggregation unit aggregates the traffic data according to a preset time window based on the characteristics of the traffic data.

[0051] Preferably, a traffic baseline corresponding to the network service type is established based on historical traffic data, and the traffic value in the next N minutes is predicted by a traffic prediction model to obtain a traffic prediction value; The flow prediction model includes: a flow data input unit, a flow feature calculation unit, a flow prediction unit and a flow prediction output unit; The flow data input unit inputs the flow data into the flow prediction model; The flow characteristic calculation unit calculates the characteristics of the flow data to obtain flow characteristics; the flow characteristics include: historical average flow, peak time flow characteristics and flow patterns in different time periods; The traffic prediction unit predicts the traffic value in the next N minutes according to the traffic data and the traffic characteristics through an LSTM network to obtain the traffic prediction value; The traffic prediction output unit outputs the traffic prediction value.

[0052] Preferably, the anomaly score of each node is obtained through an anomaly score model; the anomaly score model includes: a data input unit, an attention weight calculation unit, an anomaly score calculation unit and an anomaly score output unit; The data input unit inputs the historical traffic data of each node, the traffic data and the traffic prediction value into the abnormal scoring model; The attention weight calculation unit uses a self-attention mechanism to extract features from the historical traffic data and the traffic prediction value to obtain corresponding historical attention weights and predicted attention weights; The abnormality score calculation unit calculates the abnormality score of each node based on the historical traffic data, the traffic data, the traffic prediction value, the historical attention weight and the predicted attention weight; The anomaly score output unit outputs the anomaly score.

[0053] Preferably, the formula for the abnormality score is: ; in, Score anomalies; is the number of historical data points; is the historical attention weight; is the current traffic data; For the Historical traffic data; To predict the attention weight; is the traffic prediction value.

[0054] Preferably, a multi-level dynamic threshold is set in combination with the traffic baseline and the traffic prediction value to perform real-time monitoring and graded alarm on the traffic data; the specific process is: if the current traffic value exceeds the multi-level dynamic threshold for M consecutive detection cycles, a first-level alarm is triggered and pushed to the management platform through the low-latency transmission channel of the HTTP / 3 protocol; if abnormal characteristics of the service access mode are detected at the same time, a second-level alarm is triggered and service traffic mirror analysis is started; The formula of the multi-level dynamic threshold is: ; in, It is a multi-level dynamic threshold; is the series; is the flow baseline; is the flow volatility weight; is the flow volatility; is the prediction error weight; is the traffic prediction value; is the anomaly score weight; Score anomalies; is the external environment weight; Environmental impact factors; is the current environment variable value; is the historical average environmental value; is the activation function.

[0055] Preferably, flow table rules are generated according to service priority and real-time traffic status, traffic paths of service instances are dynamically allocated, and cross-node migration of service instances is automatically triggered when the second-level alarm persists.

[0056] Table 4 shows the SDN-based dynamic traffic path allocation performance test data obtained through multi-level dynamic threshold setting and graded alarm.

[0057]

[0058] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A data center dynamic data flow monitoring system, characterized in that: include: Traffic collection module, which is used to deploy data collection probes on SDN switches and physical servers in the data center to collect traffic data of each node in real time; The traffic prediction module is used to establish a traffic baseline corresponding to the network service type based on historical traffic data, and predict the traffic value in the next N minutes through a traffic prediction model to obtain a traffic prediction value; The monitoring threshold adjustment and graded alarm module is used to obtain the anomaly score of each node through the anomaly scoring model, and set a multi-level dynamic threshold in combination with the traffic baseline and the traffic prediction value to perform real-time monitoring and graded alarm on the traffic data; the specific process is: if the current traffic value exceeds the multi-level dynamic threshold for M consecutive detection cycles, a first-level alarm is triggered and pushed to the management platform through the low-latency transmission channel of the HTTP / 3 protocol; If abnormal service access pattern characteristics are detected at the same time, a level 2 alarm is triggered and service traffic mirroring analysis is started; The network traffic path allocation module generates flow table rules based on service priority and real-time traffic status, dynamically allocates traffic paths for service instances, and automatically triggers cross-node migration of service instances when the second-level alarm persists.

2. A data center dynamic data flow monitoring system according to claim 1, characterized in that: A data pre-processing unit and a data aggregation unit are also included after the traffic collection module; The data preprocessing unit preprocesses the collected flow data and stores the preprocessed flow data in a distributed flow database; The preprocessing includes: deduplication, format conversion and time synchronization; The data aggregation unit aggregates the traffic data according to a preset time window based on the characteristics of the traffic data.

3. A data center dynamic data flow monitoring system according to claim 1, characterized in that: The flow prediction model includes: a flow data input unit, a flow feature calculation unit, a flow prediction unit and a flow prediction output unit; The flow data input unit inputs the flow data into the flow prediction model; The flow characteristic calculation unit calculates the characteristics of the flow data to obtain flow characteristics; the flow characteristics include: historical average flow, peak time flow characteristics and flow patterns in different time periods; The traffic prediction unit predicts the traffic value in the next N minutes according to the traffic data and the traffic characteristics through an LSTM network to obtain the traffic prediction value; The traffic prediction output unit outputs the traffic prediction value.

4. A data center dynamic data flow monitoring system according to claim 1, characterized in that: The anomaly scoring model comprises: a data input unit, an attention weight calculation unit, an anomaly scoring calculation unit and an anomaly scoring output unit; The data input unit inputs the historical traffic data of each node, the traffic data and the traffic prediction value into the abnormal scoring model; The attention weight calculation unit uses a self-attention mechanism to extract features from the historical traffic data and the traffic prediction value to obtain corresponding historical attention weights and predicted attention weights; The abnormality score calculation unit calculates the abnormality score of each node based on the historical traffic data, the traffic data, the traffic prediction value, the historical attention weight and the predicted attention weight; The anomaly score output unit outputs the anomaly score.

5. A data center dynamic data flow monitoring system according to claim 4, characterized in that: The formula for the anomaly score is: ; in, Score anomalies; is the number of historical data points; is the historical attention weight; is the current traffic data; For the Historical traffic data; To predict the attention weight; is the traffic prediction value.

6. A data center dynamic data flow monitoring system according to claim 1, characterized in that: The formula of the multi-level dynamic threshold is: ; in, It is a multi-level dynamic threshold; is the series; is the flow baseline; is the flow volatility weight; is the flow volatility; is the prediction error weight; is the traffic prediction value; is the anomaly score weight; Score anomalies; is the external environment weight; Environmental impact factors; is the current environment variable value; is the historical average environmental value; is the activation function.

7. A method for monitoring dynamic data flow in a data center, characterized in that: include: Deploy data collection probes on the SDN switches and physical servers in the data center to collect traffic data of each node in real time; Based on historical traffic data, a traffic baseline corresponding to the network service type is established, and the traffic value for the next N minutes is predicted through a traffic prediction model to obtain a traffic prediction value; The anomaly score of each node is obtained through the anomaly scoring model, and a multi-level dynamic threshold is set in combination with the traffic baseline and the traffic prediction value to perform real-time monitoring and graded alarm on the traffic data; the specific process is: if the current traffic value exceeds the multi-level dynamic threshold for M consecutive detection cycles, a first-level alarm is triggered and pushed to the management platform through the low-latency transmission channel of the HTTP / 3 protocol; If abnormal service access pattern characteristics are detected at the same time, a level 2 alarm is triggered and service traffic mirroring analysis is started; Generate flow table rules based on service priority and real-time traffic status, dynamically allocate traffic paths for service instances, and automatically trigger cross-node migration of service instances when the second-level alarm persists.

Citation Information

Patent Citations

  • Method and system for detecting abnormal traffic data

    CN108737406A

  • Industrial control network flow prediction method and device based on deep learning

    CN112073255A

  • Sensitive data detection and protection method and device, equipment and medium

    CN112787992A

  • Network traffic anomaly detection and analysis method and device

    CN114124492A

  • Network traffic anomaly detection method and device, electronic equipment and storage medium

    CN115150248A

Cited By

  • IPTV edge node real-time quality monitoring and self-repairing system and control method

    CN120186141A

  • A real-time quality monitoring and self-repair system and control method for IPTV edge nodes

    CN120186141B

  • Flow management method, system and equipment of deterministic network and medium

    CN120825462A