A data monitoring method and device
By dynamically determining the granularity of monitoring time in the monitoring system using classification algorithms, the problems of weakening monitoring data characteristics and large workload of operation and maintenance personnel in the existing technology are solved, and higher monitoring and analysis accuracy and operation and maintenance efficiency are achieved.
Patent Information
- Application Number
- CN202111368616.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-11-18
AI Technical Summary
In the prior art, multiple monitoring objects acquire monitoring data through the same monitoring time granularity, resulting in weakening of data characteristics and prone to abnormal misreports and misreports. The time granularity selection of monitoring algorithms depends on the experience of operation and maintenance personnel, and the workload is large.
The optimal monitoring time granularity is determined through the classification algorithm, and the data that is most in line with the monitoring data changes of the monitoring object are obtained, so as to improve the accuracy of the monitoring analysis results. The specific method includes: obtaining the data to be detected for the monitoring item according to the first monitoring time granularity, detecting and determining the second monitoring time granularity through the monitoring algorithm, and performing feature extraction and classification algorithm updates.
It improves the accuracy of monitoring and analysis results, reduces the workload of operation and maintenance personnel and the monitoring algorithm determination cycle, reduces abnormal missed and false alarms, and improves operation and maintenance efficiency.
Smart Images

Figure CN114090377B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network technologies, and in particular, to a data monitoring method and apparatus. Background Art
[0002] With the rapid development of computer technologies and the increasing demand for services by people, the system structures and functions of enterprise business systems have become increasingly complex. Correspondingly, in order to ensure the timeliness and accuracy of the operation and maintenance of business systems, the number of monitoring objects to be monitored and the amount of monitoring data of monitoring objects in business systems have increased exponentially. For example, the information central centers of some enterprises have realized real-time monitoring functions based on various types of data sources such as database transactions, network traffic, and application logs, covering approximately 200 systems including core key systems, key application components, and co-operating systems in business systems, with a total of 37,000 monitoring objects. By monitoring and analyzing the monitoring data of 37,000 monitoring objects, the operation status and business processing status of business systems are obtained to ensure the operation and maintenance of business systems and meet the transaction monitoring requirements of internal and external departments and branches of the information central center.
[0003] In the prior art, due to the large number of monitoring objects, generally the same monitoring algorithm is set for multiple monitoring objects, and the monitoring data of multiple monitoring objects with the same set monitoring time granularity is directly obtained. Based on the analysis of each monitoring data obtained with the same monitoring time granularity, the operation and maintenance of business systems are realized. Although this method can reflect some operation status and business processing status of business systems, however, obtaining monitoring data for multiple monitoring objects with the same monitoring time granularity weakens the data characteristics of the monitoring data of each monitoring object, and it is easy to have abnormal missed reports and misreports. In addition, as time goes by, the monitoring time granularity for obtaining monitoring data remains unchanged while the monitoring data of monitoring objects changes, and it is easy to have a mismatch between the monitoring algorithm and the monitoring data, reducing the accuracy of monitoring analysis results and causing missed reports or false alarms of abnormalities. Moreover, the selection of the time granularity of the monitoring algorithm depends on the experience of operation and maintenance personnel, and a large amount of calculation and analysis work needs to be done by operation and maintenance personnel, with a long determination period and a large workload.
[0004] Therefore, there is an urgent need for a data monitoring method and apparatus that can determine the optimal monitoring time granularity to obtain monitoring data that best conforms to the changes in the monitoring data of monitoring objects, improve the accuracy of monitoring analysis results, and reduce the workload of operation and maintenance personnel and the determination period of the monitoring algorithm. Summary of the Invention
[0005] The embodiments of this application provide a data monitoring method and apparatus that can determine the optimal monitoring time granularity to obtain monitoring data that best conforms to the changes in the monitoring data of monitoring objects, improve the accuracy of monitoring analysis results, and reduce the workload of operation and maintenance personnel and the determination period of the monitoring algorithm.
[0006] In a first aspect, an embodiment of the present application provides a data monitoring method, which includes:
[0007] Obtain the data to be detected of the monitoring item within a preset period according to a first monitoring time granularity; the first monitoring time granularity is determined by a classification algorithm for the historical data of the monitoring item; detect the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitoring item has an abnormality; determine a second monitoring time granularity according to the detection result, and perform feature extraction on the data to be detected to obtain a feature vector; use the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
[0008] In the above method, the classification algorithm determines the first monitoring time granularity through the historical data of the monitoring item, so that the data to be detected obtained by the first monitoring time granularity is the time granularity that can best reflect the operating condition of the component corresponding to the monitoring item. In this way, the accuracy of the detection result is improved. Further, the data to be detected is detected through a monitoring algorithm to obtain a detection result on whether the monitoring item has an abnormality, and the second monitoring time granularity is determined according to the detection result. In this way, a correction effect is achieved on the first monitoring time granularity according to the detection result, and the second monitoring time granularity is obtained, that is, a more suitable monitoring time granularity is obtained. In this way, the monitoring time granularity is optimized, that is, the data to be detected obtained by the second monitoring time granularity is the time granularity that can best reflect the operating condition of the component corresponding to the monitoring item. It is realized that the monitoring time granularity changes according to the change of the monitoring data of the monitoring item. Compared with the inaccurate data to be detected obtained due to the unchanged monitoring time granularity in the prior art, which further leads to inaccurate detection results and cannot perform operation and maintenance accurately and quickly, the present application can improve the operation and maintenance efficiency and accuracy. Further still, feature extraction is performed on the data to be detected to obtain a feature vector, and the feature vector and the second monitoring time granularity are used as sample data to train and update the classification algorithm. In this way, the accuracy of the classification algorithm is improved, and the accuracy of determining the monitoring time granularity next time is improved. In this way, the feature extraction of the monitoring data is performed specifically for each monitoring item, and further the "customized" monitoring time granularity for the monitoring item is obtained, and the original monitoring time granularity is updated to obtain the data to be detected, ensuring that each monitoring item can obtain the monitoring time granularity that can best obtain the monitoring data features of the monitoring item. Compared with the manual analysis of the monitoring time granularity in the prior art, the present application can automatically analyze the optimal monitoring time granularity, obtain the monitoring data to be detected according to the optimal monitoring time granularity, detect the monitoring data to be detected through a monitoring algorithm, improve the authenticity of obtaining the monitoring data features of the monitoring item, reduce the workload of operation and maintenance personnel, reduce the cycle of obtaining the monitoring time granularity, and further reduce the operation and maintenance cost.
[0009] Optionally, detecting the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitored item is abnormal, including: determining a monitoring model applicable to the monitored item from the monitoring algorithm; the monitoring model applicable to the monitored item is determined based on the curve characteristics obtained from the historical data of the monitored item; analyzing the data to be detected through the monitoring model applicable to the monitored item to obtain a detection result on whether the monitored item is abnormal.
[0010] In the above method, the monitoring algorithm determines the monitoring model used for the data to be detected according to the input data to be detected, and analyzes the data to be detected according to the monitoring model to obtain a detection result. Thus, compared with the prior art where the monitoring model corresponding to the monitored item remains unchanged, in this application, for the detection data of different monitored items, a suitable monitoring model is adaptively selected to detect the data to be detected, improving the detection accuracy.
[0011] Optionally, the monitoring algorithm includes a machine learning monitoring model and a fixed threshold monitoring model; if the curve characteristics obtained from the historical data of the monitored item are strongly periodic and weakly fluctuating, the monitoring model applicable to the monitored item is the machine learning monitoring model; if the curve characteristics obtained from the historical data of the monitored item are weakly periodic and strongly fluctuating, the monitoring model applicable to the monitored item is the fixed threshold monitoring model.
[0012] In the above method, for the historical data curve characteristics of the monitored item that are strongly periodic and weakly fluctuating, the monitoring model of the monitored item is selected as the machine learning monitoring model. Thus, the machine learning monitoring model can well analyze features such as the change law of the curve, improving the detection accuracy. For the historical data curve characteristics of the monitored item that are weakly periodic and strongly fluctuating, the monitoring model of the monitored item is selected as the fixed threshold monitoring model. Thus, for a curve with weak regularity, but the boundary values, such as the upper baseline and the lower baseline, can be obtained, then it can be monitored according to the fixed threshold monitoring model, reducing the false negative rate and the false positive rate.
[0013] Optionally, extracting features from the data to be detected to obtain a feature vector, including: determining the fluctuation type index, the glitch type index for characterizing glitches, the period type index for characterizing the period, and the statistical type index of the data to be detected under different preset monitoring time granularities.
[0014] In the above method, the feature vector includes fluctuation type indicators of the data to be detected under different preset monitoring time granularities, burr type indicators for characterizing burrs, periodic type indicators for characterizing periods, and statistical type indicators of the data to be detected. In this way, the fluctuation characteristics, burr characteristics, periodic characteristics of the data to be detected, and statistical characteristics such as information entropy and preset quantiles under different time granularities can be covered, improving the comprehensiveness of the curve feature extraction of the data to be detected, and further increasing the accuracy of determining the monitoring time granularity of this monitoring item.
[0015] Optionally, the fluctuation type indicators are the information entropy and preset quantiles of the data to be detected under different preset monitoring time granularities; the burr type indicators are the maximum normalized derivatives determined based on the data to be detected; the periodic type indicators are the frequencies of the sine waves with the largest signal amplitudes obtained by performing Fourier transform on the data to be detected.
[0016] Optionally, using the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm includes: normalizing the feature vector; using the normalized feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
[0017] In the above method, normalizing the feature vector facilitates training the classification algorithm.
[0018] Optionally, obtaining the data to be detected of the monitoring item within a preset period according to the first monitoring time granularity includes:
[0019] Obtaining the monitoring data of the monitoring item within a preset period according to the first monitoring time granularity;
[0020] Complementing zero to the data items of the missing data in the monitoring data of the monitoring item;
[0021] Performing local smoothing on the historical data after complementing zero through a linear interpolation algorithm to obtain the data to be detected.
[0022] In the above method, complementing zero to the data items of the missing data of the monitoring data according to the monitoring data storage logic corresponding to the monitoring item ensures the integrity of the monitoring data and facilitates subsequent feature extraction to determine the monitoring time granularity. Performing local smoothing on the monitoring data group. In this way, abnormal data caused by normal operations such as stress testing and daily cutover can be smoothed out, improving the accuracy of the change characteristics of the monitoring data group.
[0023] In a second aspect, an embodiment of the present application provides a data monitoring device, and the device includes:
[0024] An acquisition module, configured to acquire the data to be detected of a monitored item within a preset time period according to a first monitoring time granularity; the first monitoring time granularity is determined by a classification algorithm for the historical data of the monitored item;
[0025] A processing module, configured to detect the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitored item has an anomaly;
[0026] The processing module is further configured to determine a second monitoring time granularity according to the detection result, and perform feature extraction on the data to be detected to obtain a feature vector;
[0027] The processing module is further configured to use the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
[0028] In a third aspect, an embodiment of the present application further provides a computing device, including: a memory, configured to store a program; a processor, configured to call the program stored in the memory and execute the method described in the various possible designs of the first aspect according to the obtained program.
[0029] In a fourth aspect, an embodiment of the present application further provides a computer-readable non-volatile storage medium, including a computer-readable program, and when a computer reads and executes the computer-readable program, the computer is caused to execute the method described in the various possible designs of the first aspect.
[0030] These implementation manners or other implementation manners of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0032] Figure 1 It is a schematic diagram of the architecture of a data monitoring system provided by an embodiment of the present application;
[0033] Figure 2 It is a schematic diagram of the architecture of a data monitoring system provided by an embodiment of the present application;
[0034] Figure 3 It is a schematic diagram of the architecture of a data monitoring system provided by an embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of the flowchart of a data monitoring method provided by an embodiment of the present application;
[0036] Figure 5 Schematic flowchart of a data monitoring method provided by an embodiment of the present application;
[0037] Figure 6 Schematic diagram of a data monitoring device provided by an embodiment of the present application. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0039] Figure 1 Schematic architecture diagram of a data monitoring system provided by an embodiment of the present application. Among them, the external system can be any component with monitoring items, such as a business system, a subsystem, etc., that needs to monitor the operation data of the monitoring items. For example, the business system, subsystem or components in the business system or subsystem of an enterprise, or the equipment system or equipment system components of workshop equipment, etc. Here, the type of the monitoring item and the attribution of the monitoring item are not specifically limited.
[0040] The data acquisition module in the data monitoring system is used to collect the operation data of the monitoring items in the external system according to the first monitoring time granularity to obtain the monitoring data of the monitoring items. Among them, the first monitoring time granularity is determined by the classifier according to the historical data of the monitoring items. The data acquisition module sends the monitoring data of the monitoring items to the data cleaning module.
[0041] The data cleaning module contains the storage logic corresponding to the monitoring data of the monitoring items. Each data item corresponding to the monitoring data is included in the monitoring data storage logic. According to the monitoring data storage logic of the monitoring items, zeros are filled in the data items with missing data in the monitoring data of the monitoring items, and local smoothing is performed on the historical data after filling zeros through a linear interpolation algorithm to obtain the data to be detected. The data cleaning module sends the data to be detected to the monitoring algorithm module.
[0042] The monitoring algorithms in the monitoring algorithm module can include multiple monitoring models, and these multiple monitoring models can be respectively divided into machine learning monitoring models and fixed threshold monitoring models. After receiving the data to be detected, each monitoring model in the monitoring algorithm determines which monitoring model is used to detect the data to be detected according to the curve characteristics of the data to be detected through voting to obtain the detection result. The monitoring algorithm module sends the detection result to the analysis module.
[0043] The analysis module analyzes the detection results, adjusts the first monitoring time granularity to obtain the second monitoring time granularity, obtains the second monitoring time granularity corresponding to the data to be detected, and sends the data to be detected to the feature extraction module.
[0044] The feature extraction module obtains the feature vector of the data to be detected according to the data to be detected, and sends the feature vector and the corresponding second monitoring time granularity to the classifier, and trains the classifier to obtain a more accurate classifier.
[0045] Based on the above system architecture, a schematic diagram of the architecture of another data monitoring system provided by an embodiment of the present application is as Figure 2 shown. Similarly, the external system can be any component with monitoring items, such as a business system, a subsystem, etc., that needs to monitor the operation data of the monitoring items. For example, the business system, subsystem, or components in the business system or subsystem of an enterprise, or the equipment system or equipment system components of workshop equipment, etc. Here, the type of the monitoring item and the attribution of the monitoring item are not specifically limited.
[0046] The data acquisition module in the data monitoring system is used to collect the operation data of the monitoring items in the external system according to the first monitoring time granularity to obtain the monitoring data of the monitoring items. The first monitoring time granularity is determined by the classifier according to the historical data of the monitoring items. The data acquisition module sends the monitoring data of the monitoring items to the data cleaning module.
[0047] The data cleaning module contains the storage logic corresponding to the monitoring data of the monitoring items. The monitoring data storage logic contains each data item corresponding to the monitoring data. According to the monitoring data storage logic of the monitoring items, zeros are filled in the data items with missing data in the monitoring data of the monitoring items, and local smoothing is performed on the historical data after filling zeros through a linear interpolation algorithm to obtain the data to be detected. The data cleaning module sends the data to be detected to the feature extraction module.
[0048] The feature extraction module obtains the feature vector of the data to be detected according to the data to be detected, and sends the feature vector to the classifier. The classifier determines the third monitoring time granularity and monitoring algorithm corresponding to the monitoring items of the data to be detected according to the feature vector of the data to be detected.
[0049] The classifier sends the third monitoring time granularity and the monitoring algorithm corresponding to the monitored item of the data to be detected to the detection and update module, so that the detection and update module updates the first monitoring time granularity and the monitoring algorithm for the monitored item of the data to be detected to the third monitoring time granularity and the corresponding monitoring algorithm, and detects the data to be detected according to the monitoring algorithm. This enables the detection and update module to subsequently obtain the subsequent data to be detected for the monitored item obtained by the data cleaning module at the third monitoring time granularity, and detect the subsequent data to be detected for the monitored item with the updated monitoring algorithm to obtain a detection result. The classifier also sends the third monitoring time granularity corresponding to the monitored item and the feature vector of the data to be detected for the monitored item to the classifier training module to train and optimize the classifier, and replaces the unoptimized classifier with the optimized classifier. In this way, the above two data monitoring systems can achieve targeted detection of monitored items, enabling the monitoring algorithm and the monitoring time granularity to change in a timely manner according to changes in the monitored items, increasing the accuracy of the detection results, and improving the operation and maintenance efficiency.
[0050] In addition, it should be noted that for the above two system architectures, Figure 1 in the corresponding system architecture, the monitoring time granularity can be corrected according to the detection result, and the classifier is trained with the corrected monitoring time granularity and the feature vector of the data to be detected for the monitored item. Figure 2 in the corresponding system architecture, the monitoring time granularity is the monitoring time granularity directly determined by the classifier, and the classifier is trained with the monitoring time granularity directly determined by the classifier and the feature vector of the data to be detected for the monitored item.
[0051] These two methods can be used in combination as needed. For example, to ensure the accuracy of the classifier and the speed of the data monitoring system, the classifier can be trained with the monitoring time granularity directly determined by the classifier and the feature vector of the data to be detected for the monitored item according to the system architecture in Figure 2 After a preset number of training rounds, the classifier can be trained with the corrected monitoring time granularity and the feature vector of the data to be detected for the monitored item according to the system architecture in Figure 1 For example, in Figure 3Schematic diagram of the architecture of the data monitoring system shown, where the data cleaning module can determine to send the data to be detected of the monitoring item to the feature extraction module according to the set cycle type (for example, the cycle for performing a detection once, or the cycle for other time periods, which can be specifically set according to needs and is not limited here), that is, the monitoring time granularity is the monitoring time granularity directly determined by the classifier, and the classifier is trained with the monitoring time granularity directly determined by the classifier and the feature vector of the data to be detected of the monitoring item; or send the data to be detected of the monitoring item to the detection and update module / monitoring algorithm module, that is, the monitoring time granularity can be corrected according to the detection result, and the classifier is trained with the corrected monitoring time granularity and the feature vector of the data to be detected of the monitoring item.
[0052] Here, determining whether to send the data to be detected of the monitoring item to the feature extraction module or to the detection and update module / monitoring algorithm module can be set according to specific needs. For example, according to the operating status of the monitoring item, determine the frequency of correcting the monitoring time granularity according to the detection result. If the monitoring item needs to frequently correct the monitoring time granularity, then it can be set that the proportion of the number of cycles of sending the data to be detected of the monitoring item to the detection and update module / monitoring algorithm module is larger, and the proportion of sending the data to be detected of the monitoring item to the feature extraction module is smaller. Correspondingly, if the monitoring item does not need to frequently correct the monitoring time granularity, then it can be set that the proportion of the number of cycles of sending the data to be detected of the monitoring item to the detection and update module / monitoring algorithm module is smaller, and the proportion of sending the data to be detected of the monitoring item to the feature extraction module is larger. In one example, when it is an odd cycle currently, the data cleaning module sends the data to be detected to the feature extraction module. The feature extraction module extracts the feature vector from the data to be detected of the monitoring item, further obtains the monitoring time granularity through the classifier, and sends the feature vector and the monitoring time granularity to the classifier training module to train the classifier and obtain an optimized classifier. When it is an even cycle currently, the data cleaning module sends the data to be detected to the detection and update module / monitoring algorithm module. The detection and update module / monitoring algorithm module detects the data to be detected of the monitoring item to obtain a detection result. Further, the analysis module obtains the corrected monitoring time granularity and monitoring algorithm according to the detection result. According to the feature extraction module extracting the feature vector of the data to be detected of the monitoring item, the feature vector, the corrected monitoring time granularity are sent to the classifier training module to train the classifier and obtain an optimized classifier. In this way, through the above method, the detection speed can be accelerated and the accuracy of the detection result can be ensured.
[0053] Based on the above system architecture, an embodiment of the present application provides a process of a data monitoring method, as Figure 4 shown, including:
[0054] Step 401: Obtain the data to be detected of the monitoring item within a preset time period according to the first monitoring time granularity; the first monitoring time granularity is determined by the classification algorithm for the historical data of the monitoring item.
[0055] Here, the classification algorithm can be a random forest algorithm, a support vector machine classification algorithm, a neural network, etc., and specific limitations are not made here. Among them, the random forest algorithm can be an algorithm that integrates the results of multiple decision trees based on the idea of Bagging (Bootstrap Aggregating) in ensemble learning, and improves the stability and accuracy of the classification results in a way of random features and voting judgment.
[0056] Step 402: Detect the data to be detected through the monitoring algorithm to obtain the detection result of whether the monitoring item has an abnormality.
[0057] Step 403: Determine the second monitoring time granularity according to the detection result, and extract features from the data to be detected to obtain a feature vector.
[0058] Step 404: Use the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
[0059] In the above method, the classification algorithm determines the first monitoring time granularity based on the historical data of the monitoring item, so that the data to be detected obtained with the first monitoring time granularity is the time granularity that can best reflect the operating conditions of the component corresponding to the monitoring item. In this way, the accuracy of the detection result is improved. Further, the data to be detected is detected by the monitoring algorithm to obtain the detection result of whether the monitoring item has an abnormality, and the second monitoring time granularity is determined according to the detection result. In this way, a correction effect is achieved on the first monitoring time granularity according to the detection result to obtain the second monitoring time granularity, that is, a more suitable monitoring time granularity is obtained. In this way, the monitoring time granularity is optimized, that is, the data to be detected obtained with the second monitoring time granularity is the time granularity that can best reflect the operating conditions of the component corresponding to the monitoring item. Implementing the monitoring time granularity to change according to the change of the monitoring data of the monitoring item, compared with the inaccurate data to be detected obtained due to the unchanged monitoring time granularity in the prior art, which further leads to inaccurate detection results and cannot perform operation and maintenance accurately and quickly, the present application can improve the operation and maintenance efficiency and accuracy. Further still, feature extraction is performed on the data to be detected to obtain a feature vector, and the feature vector and the second monitoring time granularity are used as sample data to train and update the classification algorithm. In this way, the accuracy of the classification algorithm is improved, and the accuracy of determining the monitoring time granularity next time is improved. In this way, the monitoring data features are extracted specifically for each monitoring item, and further the "customized" monitoring time granularity for the monitoring item is obtained, and the original monitoring time granularity is updated to obtain the data to be detected, ensuring that each monitoring item can obtain the monitoring time granularity that can best obtain the monitoring data features of the monitoring item. Compared with the manual analysis of the monitoring time granularity in the prior art, the present application can automatically analyze the optimal monitoring time granularity, obtain the monitoring data to be detected according to the optimal monitoring time granularity, detect the monitoring data to be detected through the monitoring algorithm, improve the authenticity of obtaining the monitoring data features of the monitoring item, reduce the workload of the operation and maintenance personnel, reduce the cycle of obtaining the monitoring time granularity, and further reduce the operation and maintenance cost.
[0060] An embodiment of the present application provides a method for obtaining a detection result. The method detects the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitored item is abnormal, including: determining a monitoring model applicable to the monitored item from the monitoring algorithm; the monitoring model applicable to the monitored item is determined based on the curve features obtained from the historical data of the monitored item; analyzing the data to be detected through the monitoring model applicable to the monitored item to obtain a detection result on whether the monitored item is abnormal. That is to say, the monitoring algorithm may include multiple monitoring models (for example, a monitoring algorithm may include a decision tree residual algorithm, an exponentially weighted moving average algorithm, a polynomial fitting algorithm, etc.). The multiple monitoring models can also be divided into machine learning monitoring models and fixed threshold monitoring models. The multiple monitoring models can determine the monitoring model for detecting the data to be detected through a voting decision method. Among them, the basis for the voting decision of the multiple monitoring models can be determined according to the curve features of the historical data of the monitored item. Further, the data to be detected is detected according to the monitoring model to obtain a detection result. In this way, the monitoring algorithm determined according to the curve features of the historical data of the monitored item and the monitoring models in the monitoring algorithm can obtain a more accurate detection result for the detection of the data to be detected for the monitored item. It should be noted here that there can also be multiple monitoring algorithms. For example, the monitoring algorithm can be a machine learning monitoring algorithm or a fixed threshold monitoring algorithm. Among them, the type of the monitoring algorithm can be determined according to the curve features of the historical data of the monitored item. For example, data with weak periodicity and strong volatility in curve features is applied with a fixed threshold monitoring algorithm. Data with strong periodicity and weak volatility in curve features is applied with a machine learning monitoring algorithm.
[0061] An embodiment of the present application provides a method for determining the type of a monitoring model. The monitoring algorithm includes a machine learning monitoring model and a fixed threshold monitoring model; if the curve features obtained from the historical data of the monitored item are strongly periodic and weakly volatile, the monitoring model applicable to the monitored item is a machine learning monitoring model; if the curve features obtained from the historical data of the monitored item are weakly periodic and strongly volatile, the monitoring model applicable to the monitored item is a fixed threshold monitoring model. That is to say, the monitoring model can be a machine learning monitoring model or a fixed threshold monitoring model. The type of the monitoring model can also be determined according to the curve features of the historical data of the monitored item. For example, data with weak periodicity and strong volatility in curve features is applied with a fixed threshold monitoring model. Data with strong periodicity and weak volatility in curve features is applied with a machine learning monitoring model. It should be noted that the determination of the type of the monitoring algorithm and the type of the monitoring model here is not limited to the analysis of the periodicity and volatility of the curve features, but can also be the analysis of features such as periodicity, volatility, trend, non-trend, statistics, etc. of the curve features.
[0062] An embodiment of the present application provides a method for obtaining a feature vector. Feature extraction is performed on the data to be detected to obtain a feature vector, including: determining fluctuation type indicators of the data to be detected under different preset monitoring time granularities, glitch type indicators for characterizing glitches, periodic type indicators for characterizing periods, and statistical type indicators of the data to be detected. That is to say, the feature vector may include features of fluctuation type indicators, glitch type indicators, periodic type indicators, and statistical type indicators under different preset monitoring time granularities, etc.
[0063] An embodiment of the present application provides another method for obtaining a feature vector. The fluctuation type index is the information entropy and preset quantiles of the data to be detected under different preset monitoring time granularities; the burr type index is the maximum normalized derivative determined based on the data to be detected; the periodic type index is the frequency of the sine wave with the largest signal amplitude obtained by performing a Fourier transform on the data to be detected. That is to say, the burr type index can be the maximum normalized derivative determined based on the data to be detected, and the periodic type index is the frequency of the sine wave with the largest signal amplitude obtained by performing a Fourier transform on the data to be detected. An embodiment of the present application also provides a 23-dimensional feature vector, including: the monitoring time granularity of the data to be detected, the information entropy features of the three monitoring time granularities used in the past (it can also be three or more monitoring time granularities close to the monitoring time granularity of the data to be detected, etc. The determination of different monitoring time granularities here can be analyzed and obtained as needed, and no specific limitation is made. In addition, it should be noted that based on the data to be detected, through an interpolation algorithm (when the time granularity in the three monitoring time granularities is smaller than the monitoring time granularity of the data to be detected, the data to be detected at the smaller time granularity is interpolated according to the data to be detected. For example, if the monitoring time granularity of the data to be detected is 10s, and the time granularity in the three monitoring time granularities that is smaller than the monitoring time granularity of the data to be detected is 5s, and the data to be detected obtained according to the 10s monitoring time granularity is 80, 80, 80, the interpolated data to be detected after interpolation is 80, 80, 80, 80, 80, 80) or an average algorithm (when the time granularity in the three monitoring time granularities is larger than the monitoring time granularity of the data to be detected, the data to be detected at the larger time granularity is obtained according to the average value of the data to be detected. For example, if the monitoring time granularity of the data to be detected is 5s, and the time granularity in the three monitoring time granularities that is larger than the monitoring time granularity of the data to be detected is 10s, and the data to be detected obtained according to the 5s monitoring time granularity is 80, 80, 80, 80, 80, 80, the interpolated data to be detected after interpolation is 80, 80, 80)), the features of the preset quantiles corresponding to the above three monitoring time granularities in the data to be detected of the monitoring item, the features of the preset quantiles of the data compliance rate in the above three monitoring time granularities, the features of the coefficient of variation of the data to be detected of the monitoring item, the features of the zero value ratio of the data to be detected of the monitoring item (the ratio of the number of data with a value of 0 in the data to be detected to the total number of the data to be detected), the features of the fluctuation amplitude of the data to be detected of the monitoring item, the features of the peak value of the data to be detected of the monitoring item, the features of the peak value variability of the data to be detected of the monitoring item, the features of the valley value of the data to be detected of the monitoring item, the features of the valley value variability of the data to be detected of the monitoring item, the features of the period of the data to be detected of the monitoring item, and the features of the daily average value of the data to be detected of the monitoring item.Among them, the fluctuation indicators include: the information entropy and monitoring time granularity of the data to be detected under different monitoring time granularities; the statistical indicators include: the preset quantiles of the data to be detected under different monitoring time granularities, as well as the daily average trading volume, peak value, peak variation degree, trough value, and trough variation degree of the data to be detected; the mixed indicators include: the coefficient of variation, zero value ratio, maximum normalized derivative (burr indicator), and fluctuation amplitude of the data to be detected; the periodic indicators include: the period of the data to be detected.
[0064] Based on the above 23-dimensional feature vector, the embodiments of the present application also provide a 23-dimensional feature vector of the data to be detected for a trading monitoring item, as shown in the following table:
[0065]
[0066]
[0067] Table 1
[0068] Here, the maximum normalized derivative is used to describe the burr situation of the curve and can be defined as:
[0069]
[0070] Among them, X t is the value corresponding to the preset quantile (such as 0.999 in the formula), and X t -X k is any data in the monitoring data except X t .
[0071] The information entropy is used to describe the volatility of the curve. By calculating the information entropy after aggregating the curve at different time granularities, the selection of the monitoring time granularity is realized:
[0072] entropy = max{p i *(1 - p i )}
[0073] Among them, p i represents the proportion of extremely small value data within a natural hour.
[0074] Periodicity. Based on the fast Fourier transform, the signal in the time domain is mapped to the frequency domain, and the frequency of the sine wave with the largest signal amplitude is the period. And other statistical features can include features such as daily average trading volume, peak value, trough value, kurtosis, skewness, trading volume within a natural hour, and standard deviation. In this way, the present application can obtain the optimal monitoring time granularity according to multiple features.
[0075] Based on the feature vector obtained according to the above indicators to determine the monitoring time granularity of the monitoring item, the data to be detected for the monitoring item can be obtained according to the determined monitoring time granularity, solving the problem of false alarms and missed alarms that are likely to occur when the monitoring algorithm detects the data to be detected for the monitoring item. Here, in an example, for a monitoring item of payment preposition, taking the information entropy of the indicator as an example, the monitoring data of this monitoring item has significant periodic characteristics. However, the local volatility of the 10-second granularity curve is relatively large, the burrs are obvious, and the maximum information entropy is 0.25. When directly using a dynamic real-time (monitoring time granularity is very small) monitoring model, it is found that the burr points frequently trigger the lower baseline, thus triggering false alarms. However, according to the data monitoring method of the present application, the data to be detected can be obtained by processing the monitoring data of this monitoring item, the feature vector of the data to be detected is obtained by feature extraction of the data to be detected, the feature vector is input into the classification algorithm to obtain a monitoring time granularity of 5 minutes. After aggregating the data to be detected - time series data at 5-minute granularity, the maximum information entropy quickly drops to 0. The upper and lower baselines constructed by XGBoost (optimized distributed gradient boosting library) can stably detect abnormal scenarios.
[0076] In another example, the total information center monitors that the transaction volume of barcode Q has dropped abnormally. Since the daily transaction volume of barcode Q is relatively low, the monitoring curve fluctuates violently and frequently touches zero. When using the traditional monitoring time granularity of 10 seconds for machine learning algorithms to monitor, because the algorithm takes into account that the transaction volume also frequently touches zero in the 10-second interval of the same period in history, it is impossible to accurately identify this abnormality. However, after obtaining the monitoring time granularity of 1 minute through the classification algorithm of the present application, the information entropy drops rapidly from 0.25 to 0.07 after aggregating at 1-minute granularity. The scheduling result is a machine learning model with a 1-minute granularity, which can accurately identify this abnormality.
[0077] The embodiment of the present application provides a data monitoring method, which uses the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm, including: normalizing the feature vector; using the normalized feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm. That is to say, before updating the classification algorithm with the feature vector, it needs to be normalized. Here, a normalization method is provided. The feature vector is standardized by the Z-score method to eliminate the differential influence of the dimensionality of numerical type features. The transformation function is:
[0078]
[0079] where σ is the standard deviation of each eigenvalue in the feature vector, u is the mean of each eigenvalue in the feature vector, and X represents any eigenvalue in the feature vector.
[0080] An embodiment of the present application provides a data processing method, which obtains the data to be detected of a monitored item within a preset period according to a first monitoring time granularity, including: obtaining the monitoring data of the monitored item within the preset period according to the first monitoring time granularity; filling zero for the data items of the missing data in the monitoring data of the monitored item; performing local smoothing on the historical data after filling zero through a linear interpolation algorithm to obtain the data to be detected. In one example, first, the data cleaning module extracts the historical data of the past 30 days of each monitored item from impala (a new query system). There should be 86,400 data items in the daily complete monitoring data (with a 10-second monitoring time granularity). According to the kafka data storage logic, missing data is directly filled with zero. Then, a linear interpolation algorithm is used to perform local smoothing on the historical data to eliminate the abnormal points caused by scenarios such as stress testing and daily cutover. In this way, it is ensured that the obtained data to be detected or feature vectors are convenient for subsequent acquisition of detection results or monitoring time granularity.
[0081] Based on the above system architecture and method, the flow of a data monitoring method provided by an embodiment of the present application is as Figure 5 shown, including:
[0082] Step 501: Train a classification algorithm based on the historical data of the monitored item to obtain the trained classification algorithm.
[0083] Step 502: Obtain the monitoring data of the monitored item according to the first monitoring time granularity.
[0084] Step 503: Preprocess the monitoring data of the monitored item to obtain the data to be detected.
[0085] Step 504: Detect the data to be detected through a monitoring algorithm to obtain a detection result.
[0086] Step 505: Determine a second monitoring time granularity according to the detection result and the first monitoring time granularity.
[0087] Step 506: Extract features from the data to be detected to obtain a feature vector.
[0088] Step 507: Input the second monitoring time granularity and the feature vector of the data to be detected into the classification algorithm to train the classification algorithm.
[0089] Step 508: Update the first monitoring time granularity of the monitored item to the second monitoring time granularity.
[0090] Step 509: Obtain the monitoring data of the monitored item according to the second monitoring time granularity.
[0091] Step 510: Preprocess the monitoring data of the monitored item to obtain the data to be detected.
[0092] Step 511: Extract features from the data to be detected to obtain a feature vector.
[0093] Step 512: Input the data to be detected and the feature vector into the trained classification algorithm to obtain the third monitoring time granularity and the monitoring algorithm.
[0094] Step 513: Update the second monitoring time granularity of the monitoring item to the third monitoring time granularity.
[0095] Step 514: Detect the data to be detected through the monitoring algorithm to obtain a detection result.
[0096] It should be noted here that the above method flow is not unique. For example, the above steps 502 to 508 are the method flow of correcting the monitoring time granularity according to the detection result and training the classification algorithm according to the corrected monitoring time granularity. This process can be executed independently. The above steps 509 to 514 are the method flow of obtaining the monitoring time granularity according to the feature vector of the data to be detected and training the classification algorithm according to the monitoring time granularity. This process can be executed independently. And step 514 can be executed before or after any step after step 510. Moreover, the two method flows of the above steps 502 to 508 and steps 509 to 514 can be combined and executed organically. For example, they can be executed once or multiple times respectively during the detection process.
[0097] Based on the same concept, an embodiment of the present invention provides a data monitoring device. Figure 6 As shown in the schematic diagram of a data monitoring device provided by an embodiment of the present application, Figure 6 it includes:
[0098] An acquisition module 601, configured to acquire the data to be detected of the monitoring item within a preset time period according to the first monitoring time granularity; the first monitoring time granularity is determined by the classification algorithm for the historical data of the monitoring item;
[0099] A processing module 602, configured to detect the data to be detected through the monitoring algorithm to obtain a detection result on whether the monitoring item has an abnormality;
[0100] The processing module is further configured to determine the second monitoring time granularity according to the detection result, and extract features from the data to be detected to obtain a feature vector;
[0101] The processing module is further configured to use the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
[0102] Optionally, the processing module 602 is specifically configured to determine a monitoring model applicable to the monitoring item from the monitoring algorithms; the monitoring model applicable to the monitoring item is determined based on the curve features obtained from the historical data of the monitoring item; analyze the data to be detected through the monitoring model applicable to the monitoring item, and obtain a detection result on whether the monitoring item has an anomaly.
[0103] Optionally, the monitoring algorithms include a machine learning monitoring model and a fixed threshold monitoring model; if the curve features obtained from the historical data of the monitoring item are strong in periodicity and weak in volatility, the monitoring model applicable to the monitoring item is the machine learning monitoring model; if the curve features obtained from the historical data of the monitoring item are weak in periodicity and strong in volatility, the monitoring model applicable to the monitoring item is the fixed threshold monitoring model.
[0104] Optionally, the processing module 602 is specifically configured to determine the fluctuation type index, the burr type index for characterizing burrs, the period type index for characterizing periods, and the statistical type index of the data to be detected under different preset monitoring time granularities.
[0105] Optionally, the fluctuation type index is the information entropy and the preset quantile of the data to be detected under different preset monitoring time granularities;
[0106] The burr type index is the maximum normalized derivative determined based on the data to be detected;
[0107] The period type index is the frequency of the sine wave with the largest signal amplitude obtained by performing Fourier transform on the data to be detected.
[0108] Optionally, the processing module 602 is specifically configured to perform normalization processing on the feature vector; use the normalized feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
[0109] Optionally, the processing module 602 is specifically configured to obtain the monitoring data of the monitoring item within a preset period according to the first monitoring time granularity;
[0110] Fill the data items of the missing data in the monitoring data of the monitoring item with zeros;
[0111] Perform local smoothing on the historical data filled with zeros through a linear interpolation algorithm to obtain the data to be detected.
[0112] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0116] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A data monitoring method, characterized in that, The method described above includes: Obtaining the data to be detected of a monitored item within a preset time period according to a first monitoring time granularity; the first monitoring time granularity is determined by a classification algorithm for the historical data of the monitored item; Detecting the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitored item has an anomaly; Determining a second monitoring time granularity according to the detection result, and performing feature extraction on the data to be detected to obtain a feature vector; Using the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm; Performing feature extraction on the data to be detected to obtain a feature vector, including: Determining the fluctuation type index, the burr type index for characterizing burrs, the period type index for characterizing periods, and the statistical type index of the data to be detected under different preset monitoring time granularities; The fluctuation type index is the information entropy and preset quantiles of the data to be detected under different preset monitoring time granularities; The burr type index is the maximum normalized derivative determined based on the data to be detected.
2. The method according to claim 1, characterized in that, Detecting the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitored item has an anomaly, including: Determining a monitoring model applicable to the monitored item from the monitoring algorithms; the monitoring model applicable to the monitored item is determined based on the curve features of the historical data of the monitored item; Analyzing the data to be detected through the monitoring model applicable to the monitored item to obtain a detection result on whether the monitored item has an anomaly.
3. The method according to claim 2, characterized in that, The monitoring algorithms include a machine learning monitoring model and a fixed threshold monitoring model; If the curve features obtained from the historical data of the monitored item are strongly periodic and weakly fluctuating, the monitoring model applicable to the monitored item is a machine learning monitoring model; If the curve features obtained from the historical data of the monitored item are weakly periodic and strongly fluctuating, the monitoring model applicable to the monitored item is a fixed threshold monitoring model.
4. The method according to claim 1, characterized in that, The period type index is the frequency of the sine wave with the largest signal amplitude obtained by performing Fourier transform on the data to be detected.
5. The method according to claim 1, characterized in that, Using the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm, including: Performing normalization processing on the feature vector; Using the normalized feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm.
6. The method according to claim 1, characterized in that, Obtaining the data to be detected of a monitored item within a preset time period according to a first monitoring time granularity, including: Obtaining the monitoring data of the monitored item within a preset time period according to the first monitoring time granularity; Complementing zero for the data items of the missing data in the monitoring data of the monitored item; Performing local smoothing on the historical data after zero complementing through a linear interpolation algorithm to obtain the data to be detected.
7. A data monitoring device, characterized in that, The device described above includes: An obtaining module, configured to obtain the data to be detected of a monitored item within a preset time period according to a first monitoring time granularity; the first monitoring time granularity is determined by a classification algorithm for the historical data of the monitored item; A processing module, configured to detect the data to be detected through a monitoring algorithm to obtain a detection result on whether the monitored item has an anomaly; The processing module is further configured to determine a second monitoring time granularity according to the detection result, and extract features from the data to be detected to obtain a feature vector; The processing module is further configured to use the feature vector and the second monitoring time granularity as sample data to train and update the classification algorithm; The processing module is further configured to determine a fluctuation type index, a glitch type index for characterizing glitches, a periodic type index for characterizing periods, and a statistical type index of the data to be detected under different preset monitoring time granularities; the fluctuation type index is the information entropy and preset quantiles of the data to be detected under different preset monitoring time granularities; the glitch type index is the maximum normalized derivative determined based on the data to be detected.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and when the program runs on a computer, the computer is caused to implement the method according to any one of claims 1 to 6.
9. A computer device, characterized in that, Including: A memory for storing a computer program; A processor for calling the computer program stored in the memory and executing the method according to any one of claims 1 to 6 according to the obtained program.
Citation Information
Patent Citations
Anomaly monitoring method and device
CN111708678A
Video data-based fraud detection method and apparatus, computer device, and storage medium
WO2021051607A1