Intelligent operation and maintenance data real-time anomaly monitoring method and system

By monitoring operation and maintenance data in real time and dividing time periods, the synchronization of traffic variation coefficients and task volume is obtained, the time points to be analyzed are screened, and the changes in task volume and traffic variation coefficients are combined, and the calculation and cluster classification of corrected variable coefficients are performed, which solves the abnormal monitoring error caused by task volume fluctuations and achieves more accurate real-time abnormal detection.

CN120433971AActive Publication Date: 2025-08-05上海炎凰数据科技有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510504518.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-05
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the real-time abnormal monitoring of intelligent operation and maintenance data, fluctuations in task volume make it difficult to distinguish between normal fluctuations and abnormalities of traffic data, resulting in large errors in monitoring results.

Method used

By monitoring operation and maintenance data in real time and dividing monitoring periods, the synchronization of traffic variation coefficients and tasks is obtained, the time points to be analyzed are screened, and the changes in task volume and traffic variation coefficients are combined, the calculation of corrected variable coefficients is performed, and anomaly determination is performed using the DBSCAN clustering algorithm.

Benefits of technology

It improves the accuracy of abnormal monitoring of operation and maintenance data, avoids normal fluctuations caused by fluctuations in task volume, and realizes more accurate real-time abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433971A_ABST
    Figure CN120433971A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traffic data exception monitoring, in particular to an intelligent operation and maintenance data real-time exception monitoring method and system. The method comprises the steps of firstly monitoring operation and maintenance data in real time and dividing monitoring time periods; further analyzing from the angles of flow change synchronism and fluctuation stability, and preliminarily obtaining a flow variation coefficient; further screening out a to-be-analyzed time point, and correcting the traffic variation coefficient according to the variation synchronism and prominence of the task load and the traffic variation coefficient; and finally, according to the distribution difference between the task load of the to-be-analyzed time point and the task load of the same kind of time point, performing real-time anomaly judgment in combination with the corrected variation coefficient. According to the method, the abnormal fluctuation of the flow is analyzed to determine the abnormal time point, and the flow variation coefficient is corrected by analyzing the synchronism of the abnormal flow variation and the task load variation and the task distribution condition of the similar variation condition, so that the normal fluctuation caused by the task load variation is prevented from being identified as abnormal, and the monitoring result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of flow data anomaly monitoring, and in particular to a method and system for real-time anomaly monitoring of intelligent operation and maintenance data. Background Art

[0002] In order to perform real-time anomaly detection on intelligent operation and maintenance data, traffic data during the intelligent operation and maintenance process is usually collected in real time, and anomaly monitoring is performed based on whether there are abnormal manifestations in the traffic data.

[0003] However, in actual network environments, multiple different types of tasks are often run in parallel, and the demands for uplink and downlink traffic from different task volumes vary greatly. When the network environment is in a state of high task demand, the overall performance of traffic data will deviate significantly from normal levels, making it difficult to accurately distinguish between normal fluctuations caused by task changes and actual anomalies, which can easily lead to "false positive" anomaly judgments and large errors in anomaly monitoring results. Summary of the Invention

[0004] In order to solve the technical problem that the existing operation and maintenance data is affected by task volume fluctuations, resulting in inaccurate anomaly monitoring, the purpose of the present invention is to provide an intelligent operation and maintenance data real-time anomaly monitoring method and system. The technical solutions adopted are as follows:

[0005] A method for real-time anomaly monitoring of intelligent operation and maintenance data, the method comprising:

[0006] Real-time monitoring of operation and maintenance data and division of monitoring time periods; the operation and maintenance data includes uplink traffic, downlink traffic and task volume;

[0007] According to the synchronization of changes in the upstream traffic and the downstream traffic in the preset historical neighborhood of the current monitoring period, combined with the fluctuation stability of the operation and maintenance data, the traffic variation coefficient of each time point in the current monitoring period is obtained; according to the changes in the traffic variation coefficient of the current monitoring period, the time points to be analyzed are screened out; according to the synchronization of changes in the task volume and the traffic variation coefficient, combined with the prominence of the traffic variation coefficient and the task volume at the time point to be analyzed compared with the corresponding data in the preset neighborhood, and the traffic variation coefficient, the corrected variation coefficient of each time point to be analyzed is obtained; the preset neighborhood of each time point to be analyzed is all non-to-be-analyzed time points before the time point to be analyzed in the current monitoring period;

[0008] The time points to be analyzed are clustered and classified according to the flow variation coefficient; and real-time anomaly determination is performed based on the distribution difference between the task volume at the time point to be analyzed and the task volume at similar time points, combined with the corrected variation coefficient.

[0009] Furthermore, the method for obtaining the flow variation coefficient includes:

[0010] Obtaining a traffic imbalance coefficient for a current monitoring period based on changes in synchronization of upstream and downstream traffic within the preset historical neighborhood and changes in data ratio;

[0011] According to the mean slope of the uplink and downlink traffic at each time point and the mean distance between the uplink and downlink traffic and the bandwidth boundary, combined with the traffic imbalance coefficient, the traffic variation coefficient at each time point in the current monitoring period is obtained; the mean slope of the uplink and downlink traffic and the mean slope of the uplink and downlink traffic are both positively correlated with the traffic variation coefficient; the mean distance between the uplink and downlink traffic and the bandwidth boundary is negatively correlated with the traffic variation coefficient.

[0012] Furthermore, the method for obtaining the flow imbalance coefficient includes:

[0013] Obtaining, within a preset historical neighborhood of the current monitoring period, a first correlation coefficient of the uplink and downlink traffic in each monitoring period based on the synchronization of changes in the uplink traffic and the downlink traffic in each monitoring period; and obtaining a data ratio coefficient of the uplink traffic to the downlink traffic in each monitoring period;

[0014] According to the change of the first correlation coefficient and the change of the data proportion coefficient in adjacent monitoring periods, combined with the first correlation coefficient, the flow imbalance coefficient of the current monitoring period is obtained; the change of the first correlation coefficient and the change of the data proportion coefficient are both positively correlated with the flow imbalance coefficient; the first correlation coefficient and the flow imbalance coefficient are negatively correlated.

[0015] Furthermore, the method for obtaining the time point to be analyzed includes:

[0016] The slopes of the flow variation coefficient at each time point in the current monitoring period are arranged from small to large to form a slope sequence, the slope sequence is divided at the maximum value of the element in the first-order difference sequence of the slope sequence, and the time points corresponding to the elements in the largest part of the slope sequence after division are used as the time points to be analyzed.

[0017] Furthermore, the method for obtaining the modified coefficient of variation includes:

[0018] According to the synchronization of changes in the task volume and the flow variation coefficient during the current monitoring period, a second correlation coefficient between the task volume and the flow variation coefficient is obtained; any time point to be analyzed is selected as a target time point, and the overall variation coefficient and the overall task volume corresponding to the target time point are obtained according to the overall characteristics of the flow variation coefficient and the overall characteristics of the task volume in the preset neighborhood of the target time point;

[0019] Obtaining a modified coefficient of variation at the target time point based on a difference between the flow rate variation coefficient and the overall coefficient of variation, a difference between the task volume and the overall task volume, and combining the second correlation coefficient and the flow rate variation coefficient;

[0020] The difference between the flow variation coefficient and the overall variation coefficient, the second correlation coefficient and the flow variation coefficient are all positively correlated with the corrected variation coefficient; the difference between the task volume and the overall task volume is negatively correlated with the corrected variation coefficient.

[0021] Furthermore, the method for obtaining the second correlation coefficient includes:

[0022] The second correlation coefficient is obtained according to the mean square error of the data curve of the task volume and the data curve of the flow variation coefficient in the current monitoring period; the mean square error and the second correlation coefficient are negatively correlated.

[0023] Furthermore, the method for performing real-time abnormality determination includes:

[0024] Obtain the task volume deviation coefficient of each time point to be analyzed based on the difference between the task volume at each time point to be analyzed and the task volume at all time points in the current monitoring period; obtain the overall deviation coefficient of each type of time point to be analyzed based on the overall characteristics of the task volume deviation coefficient of each type of time point to be analyzed;

[0025] According to the difference between the task volume deviation coefficient and the overall deviation coefficient at each time point to be analyzed, combined with the corrected variation coefficient, the actual abnormal volume at each time point to be analyzed is obtained; the difference between the task volume deviation coefficient and the overall deviation coefficient and the corrected variation coefficient are both positively correlated with the actual abnormal volume;

[0026] A real-time abnormality determination is performed based on the actual abnormality amount.

[0027] Furthermore, the method for performing real-time abnormality determination based on the actual abnormality amount includes:

[0028] When the actual abnormality is less than the first preset abnormality threshold, it is judged as a slight abnormality; when the actual abnormality is greater than or equal to the first preset abnormality threshold and less than the second preset abnormality threshold, it is judged as a moderate abnormality; when the actual abnormality is greater than or equal to the second preset abnormality threshold, it is judged as a serious abnormality.

[0029] Furthermore, the clustering adopts the DBSCAN clustering algorithm.

[0030] The present invention also proposes a real-time anomaly monitoring system for intelligent operation and maintenance data, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements any one of the steps of the real-time anomaly monitoring method for intelligent operation and maintenance data.

[0031] The present invention has the following beneficial effects:

[0032] The present invention first monitors operation and maintenance data in real time and divides the monitoring period to facilitate a more detailed analysis of the fluctuation and change characteristics of the operation and maintenance data; further obtains the flow variation coefficient at each time point in the current monitoring period, preliminarily quantifies the degree of flow abnormality change from the perspective of flow change, and prepares for subsequent correction in combination with the task volume; further screens out the time points to be analyzed, preliminarily determines the time points where anomalies may exist, and narrows the analysis scope; further analyzes the influence of the task volume on the flow variation coefficient based on the synchronization of changes in the task volume and the flow variation coefficient, combines the prominence of the flow variation coefficient and the task volume at the time point to be analyzed compared with the corresponding data in the preset neighborhood, analyzes the abnormal prominence characteristics of the flow variation coefficient and the task volume compared with the normal situation at the previous normal time point, corrects the flow variation coefficient, and analyzes the possibility that the flow variation is caused by the task volume; finally, clusters and classifies the time points to be analyzed based on the flow variation coefficient; according to the distribution difference between the task volume at the time point to be analyzed and the task volume at similar time points, the corrected variation coefficient is adjusted with the help of the task volume distribution of similar abnormal situations to perform real-time abnormality judgment. The present invention analyzes the abnormal fluctuation of flow to determine the abnormal time point, and then corrects the flow variation coefficient by analyzing the synchronization of abnormal flow changes and task volume changes, as well as the task distribution of similar abnormal changes, to avoid normal fluctuations caused by task volume fluctuations from being identified as abnormalities, making the monitoring results more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0034] Figure 1 A flowchart of a method for real-time abnormality monitoring of intelligent operation and maintenance data provided by one embodiment of the present invention;

[0035] Figure 2 A waveform diagram of uplink traffic and downlink traffic provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0036] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features and effects of a method and system for real-time anomaly monitoring of intelligent operation and maintenance data proposed by the present invention. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics of one or more embodiments may be combined in any suitable form.

[0037] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0038] The following describes in detail a method and system for real-time abnormality monitoring of intelligent operation and maintenance data provided by the present invention with reference to the accompanying drawings.

[0039] See also Figure 1 , which shows a flow chart of a method for real-time abnormality monitoring of intelligent operation and maintenance data provided by one embodiment of the present invention, specifically including:

[0040] Step S1: monitor operation and maintenance data in real time and divide the monitoring period; the operation and maintenance data includes uplink traffic, downlink traffic and task volume.

[0041] To monitor intelligent operation and maintenance data for real-time anomalies and for more detailed, long-term bandwidth and traffic monitoring, professional network monitoring tools are often required. These tools can regularly capture traffic data, generate reports, and provide more analytical and alerting capabilities.

[0042] In one embodiment of the present invention, NetFlow is used to obtain traffic data generated by network devices and monitor operation and maintenance data in real time. NetFlow is a network protocol specifically used to collect and analyze traffic information in IP networks. When obtaining traffic data from network devices, a unified data format is required to ensure that core fields such as timestamp, traffic direction (uplink / downlink), traffic rate data, and device identification are included. Task logs or reports are retrieved to obtain task quantities.

[0043] In order to facilitate the analysis of local fluctuation characteristics of operation and maintenance data, the acquired operation and maintenance data is divided into monitoring periods. As an example, the time domain length of the monitoring period is 1 minute, each minute is a monitoring period, and the data acquisition frequency is set to 1 Hz.

[0044] It should be noted that the upstream traffic and downstream traffic are specific traffic rate data, and the task volume is the quantity data of the task; in other embodiments of the present invention, the implementer may also use other protocols such as sFlow and IPFIX to monitor operation and maintenance data; other monitoring period lengths and data collection frequencies may also be set, which are all technical means well known to those skilled in the art and will not be elaborated here.

[0045] Step S2: Based on the synchronization of changes in upstream and downstream traffic in the preset historical neighborhood of the current monitoring period, combined with the fluctuation stability of the operation and maintenance data, the traffic variation coefficient of each time point in the current monitoring period is obtained; based on the changes in the traffic variation coefficient of the current monitoring period, the time points to be analyzed are screened out; based on the synchronization of changes in the task volume and the traffic variation coefficient, combined with the prominence of the traffic variation coefficient and the task volume at the time point to be analyzed compared with the corresponding data in the preset neighborhood, and the traffic variation coefficient, the corrected variation coefficient of each time point to be analyzed is obtained; the preset neighborhood of each time point to be analyzed is all non-time points to be analyzed before the corresponding time point to be analyzed in the current monitoring period.

[0046] When the traffic data in intelligent operation and maintenance is relatively normal, the upstream and downstream traffic will be in a relatively stable synchronous change state. When the traffic data is abnormal, the stable state between the upstream and downstream traffic will be broken, and the fluctuation stability of the operation and maintenance data will also be affected; and a longer time range is easier to capture the synchronization of traffic changes, and at the same time, the local change characteristics of the operation and maintenance data at a single time point can be analyzed. Therefore, according to the synchronization of the changes in upstream and downstream traffic in the preset historical neighborhood of the current monitoring period, combined with the fluctuation stability of the operation and maintenance data, the traffic variation coefficient at each time point in the current monitoring period is obtained, and the degree of abnormal traffic change is preliminarily quantified from the perspective of traffic change, in preparation for subsequent corrections based on the task volume.

[0047] Preferably, in one embodiment of the present invention, the greater the degree of change in the synchronization of traffic changes in different monitoring periods within a preset historical neighborhood, the more likely it is that the traffic will fluctuate abnormally, and the stable relationship between upstream and downstream traffic will be unbalanced; considering that the ratio of upstream and downstream traffic in normal operation and maintenance data is relatively stable, such as the downstream traffic in Web services is usually 5-10 times that of the upstream traffic; in live video broadcasting, since the server receives the video stream from the content end (upstream) and distributes the content to many viewers, the downstream traffic in live video broadcasting is usually much higher than the upstream traffic; therefore, the imbalance of traffic can also be measured from the perspective of changes in data ratios;

[0048] Based on this, the traffic imbalance coefficient of the current monitoring period is obtained according to the changes in the synchronization of the upstream and downstream traffic changes and the changes in the data ratio within the preset historical neighborhood;

[0049] Considering that when traffic data in intelligent operation and maintenance is abnormal, the upstream and downstream traffic will drop or surge suddenly, causing traffic congestion or a continuous low state, it is necessary to analyze whether there is a surge or drop in the traffic data; please refer to Figure 2 , which shows a waveform diagram of uplink traffic and downlink traffic provided by an embodiment of the present invention, Figure 2 The horizontal axis is time, the vertical axis is traffic rate, the solid line is upstream traffic, and the dashed line is downstream traffic. The range of 90-105 seconds shows a surge in upstream traffic and a sharp drop in downstream traffic.

[0050] Since communication operators limit bandwidth, uplink and downlink traffic will correspond to a bandwidth range, such as Figure 2 The traffic surge and drop phenomenon shown in , the uplink and downlink traffic will be close to the bandwidth limit; at the same time, the slope of the uplink and downlink traffic is large;

[0051] Based on this, the traffic variation coefficient at each time point in the current monitoring period is obtained according to the mean slope of the uplink and downlink traffic and the mean distance between the uplink and downlink traffic and the bandwidth boundary at each time point, combined with the traffic imbalance coefficient; the mean slope of the uplink and downlink traffic and the mean slope of the uplink and downlink traffic are both positively correlated with the traffic variation coefficient; the mean distance between the uplink and downlink traffic and the bandwidth boundary is negatively correlated with the traffic variation coefficient.

[0052] As an example, the ratio of the average slope of the upstream and downstream traffic and the average distance between the upstream and downstream traffic and the bandwidth boundary at each time point is normalized by multiplying the product by the traffic imbalance coefficient. The normalized result is used as the traffic variation coefficient at each time point. The ratio of the average slope of the upstream and downstream traffic and the average distance between the upstream and downstream traffic and the bandwidth boundary represents the fluctuation stability of the operation and maintenance data. The smaller the ratio, the more stable the operation and maintenance data fluctuation.

[0053] The bandwidth boundary distance refers to the minimum distance between the upstream or downstream traffic and the bandwidth range. For example, if the upstream traffic at a certain point in time is 40 Kbps and the bandwidth range is 0-60 Kbps, the corresponding bandwidth boundary distance for the upstream traffic is 20; if the downstream traffic is 15 Kbps and the bandwidth range is 0-60 Kbps, the corresponding bandwidth boundary distance for the downstream traffic is 15. The average bandwidth boundary distance is 17.5. The average of the upstream and downstream traffic and the bandwidth boundary distance corresponds to the denominator of the ratio.

[0054] It should be noted that the bandwidth ranges of upstream and downstream traffic may be different. When the mean of the bandwidth boundary distance is zero, an extremely small positive parameter such as 0.1 is assigned to it to prevent the denominator from being zero. Alternatively, when calculating the ratio, the denominator can be directly changed to the sum of the mean of the bandwidth boundary distance and a preset positive divisible parameter such as 0.1 to prevent the denominator from being zero.

[0055] In other embodiments of the present invention, the absolute values of the slopes of the upstream and downstream flows can be taken respectively and then the average can be calculated to increase the sensitivity to fluctuations in the upstream and downstream flows. The ratio of the average of the slopes of the upstream and downstream flows and the average of the distances between the upstream and downstream flows and the bandwidth boundary can be normalized first, and then multiplied by the flow imbalance coefficient and then normalized again to avoid a large range of flow fluctuations and the ratio term dominating the flow variation coefficient, thereby enhancing the dimensional consistency and numerical stability among various indicators. When performing normalization, functions such as the hyperbolic tangent function or linear normalization can be used for normalization.

[0056] Preferably, in one embodiment of the present invention, within a preset historical neighborhood of the current monitoring period, a first correlation coefficient of the uplink and downlink traffic in each monitoring period is obtained according to the synchronization of changes in the uplink traffic and the downlink traffic in each monitoring period; a data ratio coefficient of the uplink traffic and the downlink traffic in each monitoring period is obtained;

[0057] According to the changes in the first correlation coefficient and the data proportion coefficient of adjacent monitoring periods, combined with the first correlation coefficient, the flow imbalance coefficient of the current monitoring period is obtained; the changes in the first correlation coefficient and the data proportion coefficient are both positively correlated with the flow imbalance coefficient; the first correlation coefficient and the flow imbalance coefficient are negatively correlated.

[0058] As an example, the preset historical neighborhood is the latest 5 monitoring periods including the current monitoring period. Considering that the smaller the mean square error between the curve of the upstream traffic and the curve of the downstream traffic, the stronger the synchronization of the changes in the upstream and downstream traffic and the stronger the correlation, the inverse of the mean square error of the upstream traffic and the downstream traffic in each monitoring period is used as the first correlation coefficient, representing the synchronization of the changes in the upstream traffic and the downstream traffic; the average value of the ratio of the upstream traffic to the downstream traffic at each time point in the monitoring period is used as the data proportional coefficient;

[0059] The absolute value of the difference between the first correlation coefficient of each monitoring period and the first correlation coefficient of the adjacent previous monitoring period is used as the correlation change coefficient of the uplink and downlink traffic in each monitoring period, representing the change characteristics of the synchronization of the change of the uplink traffic and the downlink traffic; the absolute value of the difference between the data proportion coefficient of each monitoring period and the data proportion coefficient of the adjacent previous monitoring period is used as the proportion change coefficient of the uplink and downlink traffic in each monitoring period, representing the change characteristics of the data proportion;

[0060] The product of the correlation change coefficient and the proportional change coefficient of each monitoring period in the preset historical neighborhood of the current monitoring period is used as the denominator, the first correlation coefficient is used as the numerator, and the fractional ratio is used as the imbalance sub-coefficient of each monitoring period. After normalizing the average value of the imbalance sub-coefficients of all monitoring periods, the normalized result is used as the flow imbalance coefficient of the current monitoring period.

[0061] It should be noted that when calculating the imbalance sub-coefficient, you can start from the second monitoring period in the preset historical neighborhood, or you can call historical data to calculate the imbalance sub-coefficient of the first monitoring period; under normal circumstances, the upstream and downstream traffic will not be completely consistent, so the mean square error is not zero, and the first correlation coefficient is not zero; when normalizing, you can use functions such as the hyperbolic tangent function or linear normalization for normalization.

[0062] As another example, considering that the DTW algorithm can also measure the synchronization of changes between time series data, the smaller the DTW distance between the curve of the upstream traffic and the curve of the downstream traffic, the shorter the optimal path, and the stronger the synchronization of changes between the upstream traffic and the downstream traffic; therefore, the inverse of the DTW distance can also be used as the first correlation coefficient.

[0063] It should be noted that the DTW algorithm and the mean square error algorithm are both existing technologies. The boundaries of uplink traffic and downlink traffic can be obtained by obtaining the bandwidth of traffic through the intelligent operation and maintenance system, and will not be repeated here.

[0064] Taking into account that whether the traffic changes are caused by the task volume or the network factors, they will affect the synchronization of changes in upstream and downstream traffic and the fluctuation stability of operation and maintenance data, thereby causing changes in the traffic variation coefficient, the time points to be analyzed are screened out according to the changes in the traffic variation coefficient during the current monitoring period, and the time points where abnormalities may occur are preliminarily determined, the analysis scope is narrowed, computing resources are concentrated, and the response speed is improved.

[0065] Preferably, in one embodiment of the present invention, considering that the larger the slope of the flow variation coefficient at the time point of the current monitoring period, the more sudden the increase in the flow variation coefficient, and the corresponding abnormal change in the flow; at the same time, the slope of the flow variation coefficient caused by normal flow fluctuations is smaller, and the slope of the flow variation coefficient caused by abnormal fluctuations is larger, so the slope of the flow variation coefficient at each time point of the current monitoring period is arranged from small to large to form a slope sequence, and the slope sequence is divided at the maximum value of the element in the first-order difference sequence of the slope sequence, and the time point corresponding to the element in the largest part of the slope sequence after division is used as the time point to be analyzed.

[0066] The maximum value of the element in the first-order difference sequence represents the most dramatic increase in the slope value before and after the corresponding position, marking the dividing point where the flow variation coefficient changes from normal fluctuation to abnormal fluctuation.

[0067] In another embodiment of the present invention, the implementer may also set a fixed threshold based on abnormal scenario simulation, and when the slope of the flow rate variation coefficient is greater than the fixed threshold, it is marked as a time point to be analyzed.

[0068] Considering that the synchronization of changes in task volume and flow rate variation coefficient represents the correlation strength characteristic between the two, it reflects the impact of task volume on flow rate variation coefficient; the prominence of flow rate variation coefficient and task volume at the time point to be analyzed compared with the corresponding data of all previous non-analyzed time points in the current monitoring period represents the abnormal prominence of flow rate variation coefficient and task volume compared with the normal situation at previous normal time points. Combined with the synchronization of changes in the two, it is possible to analyze the possibility that flow rate variation is caused by task volume, and thus correct the flow rate variation coefficient;

[0069] Therefore, according to the synchronization of the changes in task volume and flow variation coefficient, combined with the prominence of the flow variation coefficient and task volume at the time point to be analyzed compared with the corresponding data in the preset neighborhood, as well as the flow variation coefficient, the corrected variation coefficient of each time point to be analyzed is obtained; the preset neighborhood of each time point to be analyzed is all non-analyzed time points before the corresponding time point to be analyzed in the current monitoring period. From the perspective of task volume change, the perspective of flow abnormality change is supplemented and corrected to improve the accuracy of the corrected variation coefficient, and ultimately improve the accuracy of operation and maintenance data abnormality monitoring.

[0070] The non-to-be-analyzed time points are the remaining time points after the filtered out to-be-analyzed time points, corresponding to the time points when traffic fluctuations are normal.

[0071] Preferably, in one embodiment of the present invention, firstly, a second correlation coefficient between the task volume and the flow variation coefficient is obtained based on the synchronization of changes in the task volume and the flow variation coefficient during the current monitoring period to characterize the correlation strength between the two.

[0072] As an example: considering that the smaller the mean square error of the data curve of the task volume and the data curve of the flow variation coefficient is, the stronger the synchronization of the changes in the task volume and the flow variation coefficient is, and the stronger the correlation is, the second correlation coefficient is obtained based on the mean square error of the data curve of the task volume and the data curve of the flow variation coefficient in the current monitoring period, which represents the synchronization of the changes in the task volume and the flow variation coefficient; the mean square error and the second correlation coefficient are negatively correlated.

[0073] Specifically, the reciprocal of the mean square error between the data curve of the task volume and the data curve of the flow rate variation coefficient is used as the second correlation coefficient.

[0074] Further select any time point to be analyzed as the target time point to analyze one by one; to facilitate the subsequent analysis of the prominence of the flow variation coefficient and the prominence of the task volume, respectively, according to the overall characteristics of the flow variation coefficient and the overall characteristics of the task volume in the current monitoring period of the preset neighborhood of the target time point, obtain the overall variation coefficient and the overall task volume corresponding to the target time point, so that in the subsequent analysis process, compare the task volume and the overall task volume, compare the flow variation coefficient and the overall variation coefficient, and analyze the prominence of the flow variation coefficient and the task volume.

[0075] As an example, the average value of the flow variation coefficients of all time points within a preset neighborhood of the target time point is taken as the overall variation coefficient, and the average value of the task volume is taken as the overall task volume; the overall characteristics of the data are obtained through the average value.

[0076] Further, according to the difference between the flow variation coefficient and the overall variation coefficient at the target time point, and the difference between the task volume and the overall task volume, the second correlation coefficient and the flow variation coefficient are combined to obtain the corrected variation coefficient at the target time point;

[0077] As an example, the difference between the flow variation coefficient at the target time point and the overall variation coefficient is normalized and used as the variation coefficient difference value, which represents the difference between the flow variation coefficient and the overall variation coefficient, reflecting the prominence of the flow variation coefficient compared with the corresponding data in the preset neighborhood; the difference between the task volume at the target time point and the overall task volume is normalized and used as the task volume difference value, which represents the difference between the task volume and the overall task volume, reflecting the prominence of the task volume compared with the corresponding data in the preset neighborhood; the product of the flow variation coefficient, the second correlation coefficient and the variation coefficient difference value is used as the numerator, the sum of the task volume difference value and a preset positive division parameter such as 0.1 is used as the denominator, and the fractional ratio is normalized to serve as the corrected variation coefficient; the normalization can be linear normalization.

[0078] Among them, the larger the difference value of the coefficient of variation, the larger the flow variation coefficient of the flow data at the target time point is compared with the flow data at the previous more normal time point; at the same time, the smaller the difference value of the task volume is, the smaller the task volume at the target time point is compared with the task volume at the previous more normal time point; the larger the second correlation coefficient is, the stronger the synchronous variability of the task volume and the flow variation coefficient is, and the flow variation coefficient should also decrease when the task volume decreases. Combining these three, it means that when the flow changes abnormally, the task volume is smaller, the abnormal situation of the flow data at the target time point is more likely to be caused by network reasons, and the larger the correction coefficient of variation is; since the correction is made on the basis of the flow variation coefficient, the larger the flow variation coefficient is, the larger the correction coefficient of variation is.

[0079] Therefore, the difference between the flow variation coefficient and the overall variation coefficient, the second correlation coefficient and the flow variation coefficient are all positively correlated with the corrected variation coefficient; the difference between the task volume and the overall task volume is negatively correlated with the corrected variation coefficient.

[0080] In other embodiments of the present invention, the DTW algorithm may also be used, and the reciprocal of the DTW distance between the data curve of the task volume and the data curve of the flow variation coefficient may be taken as the second correlation coefficient; negative correlation mapping may also be performed by replacing the reciprocal with a negative correlation mapping function such as the exp(-x) function, where x represents the independent variable; when analyzing the overall characteristics of the data, the average value, mode, and median of the flow variation coefficient of all time points within a preset neighborhood of the target time point may also be weighted and summed to obtain the overall variation coefficient, for example, with weighted weights of 0.5, 0.3, and 0.2, and the overall task volume may be obtained similarly;

[0081] Step S3: Cluster and classify the time points to be analyzed according to the flow variation coefficient; perform real-time anomaly determination based on the distribution difference between the task volume at the time point to be analyzed and the task volume at similar time points, combined with the modified variation coefficient.

[0082] Taking into account the driving effect of task volume on traffic changes, if the abnormal traffic changes are caused by task volume, the distribution of task volume corresponding to similar traffic variation coefficients will be a reference for the task volume at the time point to be analyzed. Therefore, the time points to be analyzed are clustered and classified according to the traffic variation coefficients. According to the distribution difference between the task volume at the time point to be analyzed and the task volume at similar time points, real-time anomaly judgment is performed in combination with the modified variation coefficient, which helps to identify sudden traffic anomalies caused by non-business behaviors and improve the accuracy of anomaly detection and the adaptability of system judgment.

[0083] Preferably, in one embodiment of the present invention, when clustering and classifying the time points to be analyzed according to the flow variation coefficient, a DBSCAN clustering algorithm is used.

[0084] Considering that the overall characteristics of the task volume at all time points in the current monitoring period represent the average state of the normal task volume in the current period and provide a benchmark for the distribution of the task volume at the time point to be analyzed, the task volume deviation coefficient of each time point to be analyzed is obtained based on the difference between the task volume at each time point to be analyzed and the overall task volume at all time points in the current monitoring period;

[0085] As an example, the average value of the task volume at all time points in the current monitoring period is taken as the benchmark task volume, and the difference between the task volume at the time point to be analyzed and the benchmark task volume is normalized to serve as the task volume deviation coefficient at the time point to be analyzed, representing the distribution characteristics of the task volume at the time point to be analyzed.

[0086] Furthermore, based on the overall characteristics of the task volume deviation coefficient of each type of time point to be analyzed, the overall deviation coefficient of each type of time point to be analyzed is obtained, which represents the overall distribution characteristics of a type of time points to be analyzed with similar flow variation coefficients. This provides a basis for subsequent analysis of whether the flow anomaly at the time point to be analyzed is driven by task behavior and for further adjustment of the modified variation coefficient.

[0087] As an example, the average value of the task volume deviation coefficient at each type of time point to be analyzed is used as the corresponding overall deviation coefficient.

[0088] Considering that the smaller the difference between the task volume deviation coefficient at the time point to be analyzed and the overall deviation coefficient, the more similar the task volume deviation is to the overall task volume deviation, it means that the distribution characteristics of the task volume in the traffic data of similar abnormal situations are similar. The more likely the abnormality of the traffic data is caused by the task volume factor, the smaller the actual abnormal amount.

[0089] Based on this, the actual abnormal amount at each time point to be analyzed is obtained according to the difference between the task deviation coefficient and the overall deviation coefficient at each time point to be analyzed, combined with the corrected variation coefficient; the difference between the task deviation coefficient and the overall deviation coefficient and the corrected variation coefficient are both positively correlated with the actual abnormal amount.

[0090] As an example, the absolute value of the difference between the overall deviation coefficient and the task volume deviation coefficient is linearly normalized, and then multiplied by the corrected variation coefficient. The product is linearly normalized and used as the actual abnormality. The absolute value of the difference between the overall deviation coefficient and the task volume deviation coefficient represents the distribution difference between the task volume at the time point to be analyzed and the task volume at similar time points.

[0091] Finally, real-time abnormality determination is performed based on the actual abnormality amount.

[0092] In one embodiment of the present invention, considering that the larger the actual abnormality is, the more abnormal the fluctuation in the operation and maintenance data is, and the more likely it is caused by network factors, it is set that when the actual abnormality is less than the first preset abnormality threshold, it is judged as a slight abnormality; when the actual abnormality is greater than or equal to the first preset abnormality threshold and less than the second preset abnormality threshold, it is judged as a moderate abnormality; when the actual abnormality is greater than or equal to the second preset abnormality threshold, it is judged as a serious abnormality.

[0093] As an example, the first preset abnormality threshold is 0.3, and the second preset abnormality threshold is 0.7;

[0094] In another embodiment of the present invention, the implementer may also use the task volume difference value as the task volume deviation coefficient; or perform weighted summation on the mean, mode and median of the task volume deviation coefficient to obtain the corresponding overall deviation coefficient; or use a larger time domain range when performing clustering, such as the most recent 20 monitoring periods, to increase the sample capacity. The DBSCAN clustering algorithm is already an existing technology and will not be described in detail.

[0095] It should be noted that, in one embodiment of the present invention, after performing real-time abnormality determination, the following steps are further performed:

[0096] When minor anomalies occur, the system can automatically adjust resources (e.g., load balancing, auto-scaling, etc.) without manual intervention; log abnormal data and use machine learning models to further analyze traffic patterns and identify potential trends;

[0097] When moderate anomalies occur, operations personnel need to manually intervene to check the source of the abnormal traffic and determine whether it is a potential attack (such as a DDoS attack) or a configuration error. Furthermore, they can take restrictive measures (such as limiting traffic speed and blocking specific IP addresses) to prevent excessive traffic from causing system crashes. Finally, based on the current anomaly pattern, they adjust alarm rules to ensure that subsequent similar events are detected in a timely manner.

[0098] When a serious anomaly occurs, the operation and maintenance team must be notified immediately for processing, which may involve shutting down some services, restricting user access, blocking malicious traffic, and other measures; and responding according to the preset emergency plan, such as adjusting network routing, switching to a backup data center, etc. Ultimately, after the anomaly is under control, the system recovery and repair process is immediately initiated to ensure that the service returns to normal, and the cause analysis is conducted to prevent similar problems from happening again.

[0099] An embodiment of the present invention also provides a real-time anomaly monitoring system for intelligent operation and maintenance data, which includes a memory, a processor and a computer program, wherein the memory is used to store the corresponding computer program, and the processor is used to run the corresponding computer program. When the computer program runs in the processor, it can implement a real-time anomaly monitoring method for intelligent operation and maintenance data described in steps S1-S3.

[0100] In summary, in response to the existing technical problem that operation and maintenance data is affected by task volume fluctuations, resulting in inaccurate abnormal monitoring, the present invention proposes a real-time abnormal monitoring method and system for intelligent operation and maintenance data. The present invention first monitors the operation and maintenance data in real time and divides the monitoring period; further, based on the synchronization of changes in upstream and downstream traffic and the fluctuation stability of operation and maintenance data, the flow variation coefficient is preliminarily obtained; further, the time points to be analyzed are screened out, and the flow variation coefficient is corrected according to the synchronization of changes in task volume and flow variation coefficient, combined with the prominence of flow variation coefficient and task volume; finally, the time points to be analyzed are classified, and real-time abnormality judgment is performed based on the distribution difference between the task volume at the time point to be analyzed and the task volume at the same time point, combined with the corrected variation coefficient. The present invention analyzes the abnormal fluctuation of traffic to determine the abnormal time point, and then corrects the flow variation coefficient by analyzing the synchronization of abnormal traffic changes and task volume changes, as well as the task distribution of the same abnormal situation, to avoid normal fluctuations caused by task volume fluctuations being identified as abnormal changes, so that the monitoring results are more accurate.

[0101] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0102] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for real-time abnormality monitoring of intelligent operation and maintenance data, characterized in that: The method comprises: Real-time monitoring of operation and maintenance data and division of monitoring time periods; the operation and maintenance data includes uplink traffic, downlink traffic and task volume; According to the synchronization of changes in the upstream traffic and the downstream traffic in the preset historical neighborhood of the current monitoring period, combined with the fluctuation stability of the operation and maintenance data, the traffic variation coefficient of each time point in the current monitoring period is obtained; according to the changes in the traffic variation coefficient of the current monitoring period, the time points to be analyzed are screened out; according to the synchronization of changes in the task volume and the traffic variation coefficient, combined with the prominence of the traffic variation coefficient and the task volume at the time point to be analyzed compared with the corresponding data in the preset neighborhood, and the traffic variation coefficient, the corrected variation coefficient of each time point to be analyzed is obtained; the preset neighborhood of each time point to be analyzed is all non-to-be-analyzed time points before the time point to be analyzed in the current monitoring period; The time points to be analyzed are clustered and classified according to the flow variation coefficient; and real-time anomaly determination is performed based on the distribution difference between the task volume at the time point to be analyzed and the task volume at similar time points, combined with the corrected variation coefficient.

2. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 1, characterized in that: The method for obtaining the flow variation coefficient includes: Obtaining a traffic imbalance coefficient for a current monitoring period based on changes in synchronization of upstream and downstream traffic within the preset historical neighborhood and changes in data ratio; According to the mean slope of the uplink and downlink traffic at each time point and the mean distance between the uplink and downlink traffic and the bandwidth boundary, combined with the traffic imbalance coefficient, the traffic variation coefficient at each time point in the current monitoring period is obtained; the mean slope of the uplink and downlink traffic and the mean slope of the uplink and downlink traffic are both positively correlated with the traffic variation coefficient; the mean distance between the uplink and downlink traffic and the bandwidth boundary is negatively correlated with the traffic variation coefficient.

3. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 2, characterized in that: The method for obtaining the flow imbalance coefficient includes: Within a preset historical neighborhood of the current monitoring period, obtaining a first correlation coefficient of the uplink and downlink traffic in each monitoring period based on the synchronization of changes in the uplink traffic and the downlink traffic in each monitoring period; obtaining a data ratio coefficient of the uplink traffic to the downlink traffic in each monitoring period; According to the change of the first correlation coefficient and the change of the data proportion coefficient in adjacent monitoring periods, combined with the first correlation coefficient, the flow imbalance coefficient of the current monitoring period is obtained; the change of the first correlation coefficient and the change of the data proportion coefficient are both positively correlated with the flow imbalance coefficient; the first correlation coefficient and the flow imbalance coefficient are negatively correlated.

4. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 1, characterized in that: The method for obtaining the time point to be analyzed includes: The slopes of the flow variation coefficient at each time point in the current monitoring period are arranged from small to large to form a slope sequence, the slope sequence is divided at the maximum value of the element in the first-order difference sequence of the slope sequence, and the time points corresponding to the elements in the largest part of the slope sequence after division are used as the time points to be analyzed.

5. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 1, characterized in that: The method for obtaining the modified coefficient of variation includes: According to the synchronization of changes in the task volume and the flow variation coefficient during the current monitoring period, a second correlation coefficient between the task volume and the flow variation coefficient is obtained; any time point to be analyzed is selected as a target time point, and the overall variation coefficient and the overall task volume corresponding to the target time point are obtained according to the overall characteristics of the flow variation coefficient and the overall characteristics of the task volume in the preset neighborhood of the target time point; Obtaining a modified coefficient of variation at the target time point based on a difference between the flow rate variation coefficient and the overall coefficient of variation, a difference between the task volume and the overall task volume, and combining the second correlation coefficient and the flow rate variation coefficient; The difference between the flow variation coefficient and the overall variation coefficient, the second correlation coefficient and the flow variation coefficient are all positively correlated with the corrected variation coefficient; the difference between the task volume and the overall task volume is negatively correlated with the corrected variation coefficient.

6. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 5, characterized in that: The method for obtaining the second correlation coefficient includes: The second correlation coefficient is obtained according to the mean square error of the data curve of the task volume and the data curve of the flow variation coefficient in the current monitoring period; the mean square error and the second correlation coefficient are negatively correlated.

7. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 1, characterized in that: The method for performing real-time abnormality determination includes: Obtain the task volume deviation coefficient of each time point to be analyzed based on the difference between the task volume at each time point to be analyzed and the task volume at all time points in the current monitoring period; obtain the overall deviation coefficient of each type of time point to be analyzed based on the overall characteristics of the task volume deviation coefficient of each type of time point to be analyzed; According to the difference between the task volume deviation coefficient and the overall deviation coefficient at each time point to be analyzed, combined with the corrected variation coefficient, the actual abnormal volume at each time point to be analyzed is obtained; the difference between the task volume deviation coefficient and the overall deviation coefficient and the corrected variation coefficient are both positively correlated with the actual abnormal volume; A real-time abnormality determination is performed based on the actual abnormality amount.

8. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 7, characterized in that: The method for performing real-time abnormality determination based on the actual abnormality amount includes: When the actual abnormality is less than the first preset abnormality threshold, it is judged as a slight abnormality; when the actual abnormality is greater than or equal to the first preset abnormality threshold and less than the second preset abnormality threshold, it is judged as a moderate abnormality; when the actual abnormality is greater than or equal to the second preset abnormality threshold, it is judged as a serious abnormality.

9. The method for real-time abnormality monitoring of intelligent operation and maintenance data according to claim 1, characterized in that: The clustering adopts the DBSCAN clustering algorithm.

10. A real-time anomaly monitoring system for intelligent operation and maintenance data, the system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for real-time anomaly monitoring of intelligent operation and maintenance data as described in any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Network traffic big data analysis method

    CN117439827A

  • Intelligent service information analysis method for comprehensive operation and maintenance platform

    CN117828371A

  • Intelligent analysis method for operation condition of anaerobic system based on machine learning

    CN118094446A

  • Data analysis method, computer device and storage medium

    US20210365421A1

  • Battery abnormality recognition method and apparatus, electronic device, and storage medium

    WO2024082104A1