Data detection method and device, electronic equipment and program product
By obtaining network monitoring data in 5G scenarios and using regression models and codecs to detect abnormalities on periodic and non-periodic single-dimensional timing data, the automated detection problems caused by the numerous KPI indicators in 5G scenarios are solved, and an efficient and low-cost abnormal detection effect is achieved.
Patent Information
- Application Number
- CN202311798643.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
In 5G scenarios, there are many KPI indicators, and it is difficult for the existing technology to realize automated data anomaly detection. Traditional methods have shortcomings in applicability and computing complexity.
By acquiring the network monitoring data set, using the trained regression model and codec, abnormal detection is performed on periodic and non-periodic single-dimensional timing data. For periodic data, predict the normal network monitoring index value and determine the error threshold; for non-periodic data, clustering is performed to judge abnormalities.
It realizes automated abnormal detection of a large number of KPI indicators in 5G scenarios, reducing the computational complexity and hardware costs, and improving detection efficiency.
Smart Images

Figure CN120223564A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular, to a data detection method, apparatus, electronic device, and program product. Background Art
[0002] In daily network operation and maintenance, network monitoring is an essential maintenance means. And an important direction in network monitoring is to monitor various key performance indicator (KPI) metrics.
[0003] Currently, the method of manually setting a constant threshold is mainly used to determine whether a KPI metric is abnormal. This method is determined through the experimental experience of network administrators and the statistical analysis of historical data by operation and maintenance experts. In the case of a large number of KPIs, it is difficult to use manually set thresholds to determine whether a KPI metric is abnormal.
[0004] To solve the above problems, traditional anomaly detection methods are also used in related technologies to determine whether a KPI metric is abnormal, such as statistics and clustering. However, these methods can only be applied to scenarios where there are significant differences in the distribution of abnormal data and normal data, and their applicability is weak. In addition, there are various problems in feature extraction of KPI metrics in related technologies. For example, the self-attention mechanism of the neural network model (Transformer) often lacks attention to the most relevant information in the search area and has poor performance in extracting local information; the rectified linear unit (ReLU) in the fully convolutional network (FCN) will cause most components to be unable to be updated, etc.
[0005] Due to the large number of KPI metrics in the fifth generation (5G) mobile communication technology scenario, the current traditional methods and mainstream methods are not applicable in the 5G scenario.
[0006] In summary, it is urgent to design a general and automated data anomaly detection method for the 5G scenario. Summary of the Invention
[0007] Embodiments of this application provide a data detection method, apparatus, electronic device, and program product, for designing a general and automated data anomaly detection method.
[0008] A data detection method provided by an embodiment of this application includes:
[0009] Obtaining a data set to be analyzed through network monitoring; the data set includes at least one single-dimensional time series data; each single-dimensional time series data corresponds to a network monitoring metric;
[0010] For each actual network monitoring metric value at each moment in each single-dimensional time series data, perform the following operations respectively:
[0011] If a single-dimensional time series data is periodic data, then predict the normal network monitoring metric value at the moment according to the trained regression model and the neighbor normal data corresponding to the moment; reconstruct the error of the historical normal sample data according to the trained codec to determine the error threshold of the single-dimensional time series data; determine whether the actual network monitoring metric value at the moment is abnormal according to the error threshold and the predicted normal network monitoring metric value at the moment; the neighbor normal data is obtained by removing abnormal data from the single-dimensional time series data within a preset time period before the moment in the single-dimensional time series data; the historical normal sample data is obtained by removing abnormal data from the historical ordinary sample data;
[0012] If the single-dimensional time series data is non-periodic data, then determine whether the actual network monitoring metric value at a moment in the single-dimensional time series data is abnormal according to the clustering result by clustering the single-dimensional time series data.
[0013] A data detection device provided by an embodiment of the present application includes:
[0014] An acquisition module, configured to acquire a data set to be analyzed through network monitoring; the data set includes at least one single-dimensional time series data; each single-dimensional time series data corresponds to a network monitoring metric;
[0015] An execution module, configured to perform the following operations respectively for each actual network monitoring metric value at each moment in each single-dimensional time series data:
[0016] If a single-dimensional time series data is periodic data, then predict the normal network monitoring metric value at the moment according to the trained regression model and the neighbor normal data corresponding to the moment; reconstruct the error of the historical normal sample data according to the trained codec to determine the error threshold of the single-dimensional time series data; determine whether the actual network monitoring metric value at the moment is abnormal according to the error threshold and the predicted normal network monitoring metric value at the moment; the adjacent normal data is obtained by removing abnormal data from the single-dimensional time series data within a preset time period before the moment in the single-dimensional time series data; the historical normal sample data is obtained by removing abnormal data from the historical ordinary sample data;
[0017] If the single-dimensional time series data is non-periodic data, then by clustering the single-dimensional time series data, according to the clustering result, it is determined whether the actual network monitoring metric value at a moment in the single-dimensional time series data is abnormal.
[0018] In the above embodiment, for the actual network monitoring metric value at each moment in the periodic single-dimensional time series data, first, the normal network monitoring metric value at this moment is predicted. Then, according to the trained encoder-decoder, the historical normal sample data is reconstructed for error to obtain an error threshold. A machine learning method with a simple structure is used to replace the method of manually setting the error threshold, avoiding the problem of difficulty in abnormal judgment when the number of KPI metrics is huge. Finally, through this normal network monitoring metric value and the error threshold, it can be determined the degree to which the actual network monitoring metric value at this moment deviates from the normal network monitoring metric value, and it is determined whether the actual network monitoring metric value is abnormal according to the deviation degree.
[0019] For the actual network monitoring metric value at each moment in the non-periodic single-dimensional time series data, the abnormal clusters are determined through clustering, and then it is determined whether the actual network monitoring metric value at each moment in the single-dimensional time series data is abnormal.
[0020] In summary, the present application proposes a method for detecting anomalies in single-dimensional time series data. Since historical normal sample data is easy to collect, an error threshold can be obtained through a machine learning method with a simple structure, thereby overcoming the difficulty caused by manually setting the error threshold when the number of KPIs in the 5G scenario is huge; and it is determined whether the actual network monitoring metric value is abnormal by the degree to which the actual network monitoring metric value deviates from the normal network monitoring metric value. Compared with the anomaly detection methods used in the related art, it may face problems such as a sharp increase in computational complexity, gradient disappearance or explosion, and is extremely dependent on the hardware resources of the Graphics Processing Unit (GPU) and the requirements for periodic data. The present application obtains the error threshold through a machine learning method with a simple structure, making the anomaly detection method of the present application have a lower computational complexity, reducing the hardware cost, and improving the detection efficiency.
[0021] Optionally, the execution module is used to determine whether the single-dimensional time series data is periodic data through the following operations:
[0022] Based on a preset correlation rule, perform a correlation analysis on the single-dimensional time series data to determine the correlation value of the single-dimensional time series data; the correlation value characterizes the periodicity in the single-dimensional time series data;
[0023] If the correlation value is greater than a preset correlation threshold, it is determined that the single-dimensional time series data is periodic data;
[0024] If the correlation value is not greater than a preset correlation threshold, it is determined that the one-dimensional time series data is aperiodic data.
[0025] In the above embodiment, by performing correlation analysis on the one-dimensional time series data to determine whether the one-dimensional time series data is periodic, it is convenient to perform anomaly detection on the one-dimensional time series data with periodicity and the one-dimensional time series data without periodicity respectively in the subsequent process.
[0026] Optionally, the execution module is specifically configured to:
[0027] By clustering historical normal sample data, obtain a first preset number of clustering clusters; by comparing the sizes of the first preset number of clustering clusters, determine a first abnormal cluster; if the actual network monitoring index value at a certain moment in the one-dimensional time series data belongs to the first abnormal cluster, it is determined that the actual network monitoring index value at this moment is abnormal; or
[0028] By clustering the one-dimensional time series data and the historical normal sample data, obtain a second preset number of clustering clusters; by comparing the sizes of the second preset number of clustering clusters, determine a second abnormal cluster; if the actual network monitoring index value at a certain moment in the one-dimensional time series data is in the second abnormal cluster, it is determined that the actual network monitoring index value at this moment is abnormal.
[0029] In the above embodiment, by clustering the one-dimensional time series data without periodicity, after determining the abnormal cluster, anomaly detection is performed on the one-dimensional time series data.
[0030] Optionally, if there are fixed index requirements for the one-dimensional time series data, the execution module is further configured to:
[0031] If the actual network monitoring index value at a certain moment in the one-dimensional time series data does not meet the fixed index requirements, it is determined that the actual network monitoring index value is abnormal.
[0032] In the above embodiment, the one-dimensional time series data is screened by solid index requirements, and static thresholds are used for anomaly detection of the one-dimensional time series data with solid index requirements.
[0033] Optionally, before the execution module determines whether the one-dimensional time series data is periodic data, the execution module is further configured to:
[0034] If there are missing data values in the one-dimensional time series data, determine a supplementary value according to a preset supplementary rule and the one-dimensional time series data;
[0035] Fill the missing data value with the supplementary value.
[0036] In the above embodiment, the missing data values in the one-dimensional time series data are filled by using supplementary values, which facilitates subsequent periodic judgment and anomaly detection of the one-dimensional time series data.
[0037] Optionally, the execution module is used to train a regression model through the following operations:
[0038] If a piece of one-dimensional time series data is periodic data, feature extraction is performed on the historical normal sample data to obtain a third preset number of statistical features;
[0039] The third preset number of statistical features are combined to obtain input features;
[0040] Based on the input features and the regression model, the sample normal network monitoring index value at the historical prediction time is predicted;
[0041] Based on the difference between the sample normal network monitoring index value at the historical prediction time and the sample actual network monitoring index value at the historical prediction time, the parameters of the regression model are adjusted.
[0042] In the above embodiment, since the historical normal sample data is easy to collect, a regression model suitable for semi-supervised learning is used to predict the normal network monitoring index value. Since a machine learning method with a simple structure, that is, the XGBoost regression model, is used for regression prediction, it can run directly on the CPU without relying on the GPU, thereby reducing the required hardware cost and improving the prediction efficiency.
[0043] Optionally, the execution module specifically is used for:
[0044] Determine the absolute value of the difference between the normal network monitoring index value and the actual network monitoring index value;
[0045] If the absolute value is greater than the error threshold, it is determined that the actual network monitoring index value is abnormal.
[0046] In the above embodiment, whether the actual network monitoring index value is abnormal is judged by the degree to which the actual network monitoring index value deviates from the normal network monitoring index value.
[0047] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of any one of the above data detection methods.
[0048] A computer program product provided by an embodiment of the present application includes a computer program. When the computer program is executed by a processor, it implements any one of the above data detection methods.
[0049] Optionally, the computer-readable storage medium may be implemented as a computer program product. That is, the embodiments of the present application further provide a computer-readable storage medium, which includes a computer program that, when executed by a processor, implements any of the above data detection methods.
[0050] Other features and advantages of the present application will be described in the subsequent description. Moreover, some of them will become apparent from the description or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0052] Figure 1 FIG. is a schematic diagram of an application scenario for data detection provided by an embodiment of the present application;
[0053] Figure 2 FIG. is a flowchart of an implementation of a data detection method provided by an embodiment of the present application;
[0054] Figure 3 FIG. is a schematic diagram of the structure of an codec AE provided by an embodiment of the present application;
[0055] Figure 4 FIG. is a flowchart of abnormal judgment of a dynamic threshold provided by an embodiment of the present application;
[0056] Figure 5 FIG. is a flowchart of abnormal detection of one-dimensional time series data provided by an embodiment of the present application;
[0057] Figure 6A FIG. is a schematic diagram of an abnormal detection result provided by an embodiment of the present application;
[0058] Figure 6B FIG. is another schematic diagram of an abnormal detection result provided by an embodiment of the present application;
[0059] Figure 7 FIG. is yet another schematic diagram of an abnormal detection result provided by an embodiment of the present application;
[0060] Figure 8 FIG. is a schematic diagram of a hardware composition structure of an electronic device applying an embodiment of the present application;
[0061] Figure 9 FIG. is a schematic diagram of the structure of a data detection device in an embodiment of the present application;
[0062] Figure 10 It is a schematic diagram of a hardware component structure of a computing device for applying an embodiment of the present application. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.
[0064] Some concepts involved in the embodiments of the present application are introduced below.
[0065] Network monitoring metrics: refer to the KPI metrics monitored in network monitoring, including but not limited to bandwidth utilization, network latency, network packet loss rate, network throughput, response time, CPU utilization, memory utilization, and network traffic distribution. Among them, network monitoring refers to the real-time monitoring and management of a computer network to ensure the normal operation and performance optimization of the network.
[0066] Fast Fourier Transform (FFT): is an efficient algorithm for calculating the Discrete Fourier Transform (DFT), which can convert a signal in the time domain into a frequency domain representation and analyze the frequency components and phase information of the signal. It is widely used in fields such as signal processing, spectrum analysis, filter design, image compression, encoding and decoding, etc. The present application calculates the FFT period value of one-dimensional time series data through the Fast Fourier Transform.
[0067] eXtreme Gradient Boost (XGBoost) regression model: is an ensemble machine learning algorithm based on decision trees. Its principle is based on the Gradient Boosting algorithm. By iteratively training multiple decision trees and continuously optimizing the loss function to improve the model performance, it is suitable for semi-supervised learning problems and performs well in data modeling and prediction tasks. The present application predicts the normal network monitoring metric value at a moment in one-dimensional time series data through the XGBoost regression model.
[0068] Gaussian Mixture Model (GMM): It is a commonly used probability model. Based on the linear combination of multiple Gaussian distributions, it approximately represents the distribution of data and can be used to model and estimate complex data distributions. The main goal of GMM is to learn appropriate mean, covariance, and weight parameters from the given data through maximum likelihood estimation or the expectation-maximization algorithm. It is widely used in tasks such as generating new samples, data clustering, and anomaly detection. In this application, clustering is performed through GMM. After obtaining a preset number of clustering clusters, the abnormal clusters are determined, and then the abnormal judgment is made on the actual network monitoring index values at each moment in the non-periodic one-dimensional time series data.
[0069] The preferred embodiments of this application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain this application and are not used to limit this application. And without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0070] As Figure 1 shown, it is a schematic diagram of an application scenario for data detection provided by an embodiment of this application. This application scenario diagram includes a terminal device 110 and a server 120.
[0071] In the embodiments of this application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablets, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, and vehicle-mounted terminals; a client related to data detection can be installed on the terminal device, and this client can be software (such as a browser, etc.), or a web page, a small program, etc. The server 120 is the background server corresponding to the software, web page, small program, etc., and this application does not make specific limitations. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0072] It should be noted that the data detection method in the embodiments of the present application can be executed by an electronic device, which can be the server 120 or the terminal device 110. That is, this method can be executed independently by the server 120 or the terminal device 110, or can be jointly executed by the server 120 and the terminal device 110. For example, when jointly executed by the server 120 and the terminal device 110, the object inputs the collected data set to be analyzed into the terminal device 110, the terminal device 110 sends the data set to the server 120, and the server 120 predicts the normal network monitoring index value at each moment for each item of one-dimensional time series data in the data set periodically, and determines the error threshold of this item of one-dimensional time series data. The server 120 determines whether the actual network monitoring index value at this moment is abnormal according to the normal network monitoring index value at this moment and the error threshold of this item of one-dimensional time series data. The server 120 clusters each item of one-dimensional time series data in the data set that is not periodic, and determines whether the actual network monitoring index value at each moment in this item of one-dimensional time series data is abnormal according to the clustering result. Furthermore, the server 120 sends the result of anomaly detection to the terminal device 110, and the terminal device 110 feeds back the anomaly detection result to the object.
[0073] In an alternative embodiment, the terminal device 110 and the server 120 can communicate through a communication network.
[0074] In an alternative embodiment, the communication network is a wired network or a wireless network.
[0075] It should be noted that Figure 1 The above is only an example, and actually the number of terminal devices and servers is not limited, and no specific limitation is made in the embodiments of the present application.
[0076] Next, in combination with the application scenarios described above, the data detection method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0077] As Figure 2 shown, it is a flowchart of the implementation of a data detection method provided by the embodiments of the present application. Taking the server as the execution entity as an example, the implementation process includes the following steps S21 to S22:
[0078] S21: Obtain the data set to be analyzed through network monitoring.
[0079] Among them, the data set refers to the KPI set stream, which contains at least one single-dimensional time series data; each single-dimensional time series data corresponds to a network monitoring metric, that is, a KPI metric. In this application, the KPI metrics include but are not limited to bandwidth utilization, network latency, network packet loss rate, network throughput, response time, CPU utilization, memory utilization, and network traffic distribution.
[0080] In this application, the "single dimension" in the single-dimensional time series data refers to considering the relationship between the KPI metric and the dimension of time.
[0081] Specifically, when obtaining the KPI set stream to be analyzed, each KPI metric is collected at a preset time interval, and the preset time intervals corresponding to different KPI metrics can be the same or different. For example, the time intervals for collecting the handover success rate between stations and the bandwidth utilization are both 15 minutes; for another example, a network throughput KPI metric can be collected every 10 minutes, and a network packet loss rate KPI metric can be collected every 5 minutes.
[0082] After obtaining the data set to be analyzed, anomaly detection can be performed on each single-dimensional time series data in the data set. This application uses different anomaly detection methods for periodic single-dimensional time series data and non-periodic single-dimensional time series data. The specific process is as follows:
[0083] S22: For the actual network monitoring metric value at each moment in each single-dimensional time series data, perform the following operations S221 - S222 respectively:
[0084] Since this application uses different anomaly detection methods for periodic single-dimensional time series data and non-periodic single-dimensional time series data, before performing anomaly detection on the single-dimensional time series data, it is necessary to first determine whether the single-dimensional time series data is periodic.
[0085] In an alternative embodiment, determine whether a single-dimensional time series data is periodic through the following operations:
[0086] Based on a preset correlation rule, perform a correlation analysis on a single-dimensional time series data to determine the correlation value of the single-dimensional time series data; the correlation value characterizes the periodicity in the single-dimensional time series data.
[0087] If the correlation value is greater than the preset correlation threshold, determine that a single-dimensional time series data is periodic data;
[0088] If the correlation value is not greater than the preset correlation threshold, determine that a single-dimensional time series data is non-periodic data.
[0089] Specifically, the preset correlation rule is the Autocorrelation Function (ACF), the correlation value is the ACF value, and the preset correlation threshold is 0.5 preset according to experience.
[0090] Calculate the ACF value of a single-dimensional time series data according to the ACF. If the ACF value of this single-dimensional time series data is greater than 0.5, then this single-dimensional time series data is periodic data. If the ACF value of this single-dimensional time series data is not greater than 0.5, then this single-dimensional time series data is non-periodic data.
[0091] For example, when judging the periodicity of single-item time series data once, the ACF value of this single-dimensional time series data calculated according to the ACF is 0.7, which is greater than the preset correlation threshold of 0.5, so this single-dimensional time series data is periodic data.
[0092] In the above embodiment, by performing correlation analysis on the single-dimensional time series data, it is judged whether the single-dimensional time series data is periodic, which is convenient for subsequent anomaly detection of the single-dimensional time series data with periodicity and the single-dimensional time series data without periodicity respectively.
[0093] Before determining whether the single-dimensional time series data is periodic data, it is also necessary to first fill in the missing data values in the single-dimensional time series data.
[0094] In an alternative embodiment, the missing data values in the single-dimensional time series data are filled by the following operations:
[0095] If there are missing data values in a single-dimensional time series data, then determine the supplementary value according to the preset supplementary rule and a single-dimensional time series data; fill in the missing data values according to the supplementary value.
[0096] Among them, the preset supplementary rule can be to fill in the missing data values with the mean, mode, and median of this single-dimensional time series data. At this time, the supplementary values are the mean, mode, and median respectively; it can also be nearest neighbor interpolation. At this time, the supplementary value is the value of the adjacent observation point; it can also be regression interpolation, random sampling filling and other rules. The present application does not make specific limitations on this.
[0097] For example, if the preset supplementary rule is to fill in the missing data values according to the mean, then calculate the average value of the data of this single-dimensional time series data except for the missing data values, and use this average value as the supplementary value to fill in each missing data value.
[0098] Considering the mainstream anomaly detection methods adopted in the related art, such as the anomaly detection method based on deep learning, such as the Long Short-Term Memory (LSTM) algorithm, may face problems such as a sharp increase in computational complexity, gradient disappearance or explosion, and is extremely dependent on the Graphics Processing Unit (GPU) hardware resources and periodic data requirements. This application proposes to combine a semi-supervised regression model and a self-developed encoder-decoder Autoencoder (AE) for one-dimensional periodic time series data to perform anomaly detection on the one-dimensional time series data. For details, refer to the steps shown in S221.
[0099] S221: If a one-dimensional time series data is periodic data, then according to the trained regression model and the normal data of the neighbors corresponding to a moment, predict the normal network monitoring index value of a moment; reconstruct the error of the historical normal sample data according to the trained encoder-decoder to determine the error threshold of a one-dimensional time series data; determine whether the actual network monitoring index value of a moment is abnormal according to the error threshold and the normal network monitoring index value of a moment.
[0100] Among them, the neighbor normal data is obtained by removing abnormal data from the one-dimensional time series data within a preset time period before a moment in a one-dimensional time series data; the historical normal sample data is obtained by removing abnormal data from the historical ordinary sample data.
[0101] In this application, the regression model can use the XGBoost regression model to predict the normal network monitoring index value of a moment, or can use neural network regression models such as Multilayer Perceptron (MLP) and Convolutional Neural Network (CNN) to predict the normal network monitoring index value of a moment, which is not specifically limited in this article.
[0102] When predicting the normal network monitoring index value of a moment through the regression model, it is necessary to train the regression model according to the historical normal sample data.
[0103] In an optional implementation manner, the regression model is trained by the following operations:
[0104] If a one-dimensional time series data is periodic data, extract features from the historical normal sample data to obtain a third preset number of statistical features;
[0105] Combine the third preset number of statistical features to obtain input features;
[0106] Predict the sample normal network monitoring index value at the historical prediction moment based on the input features and the regression model;
[0107] Adjust the parameters of the regression model based on the difference between the sample normal network monitoring index value and the sample actual network monitoring index value at the historical prediction moment.
[0108] Among them, the third preset quantity can be set according to experience for historical normal sample data, or can be automatically selected by the algorithm for training the regression model, and the present application does not make specific limitations on this.
[0109] Since there are also various problems in the feature extraction of KPI indicators in the related art. For example, the self-attention mechanism of Transformer often lacks attention to the most relevant information in the search area and has poor effect in extracting local information, and the ReLU function in FCN will cause most components to be unable to be updated, etc. The present application proposes a method of first extracting the sliding window features of KPI indicators and then combining the sliding window features, which avoids the above problems.
[0110] Specifically, when training the regression model according to historical normal sample data, the historical normal sample data is automatically divided into multiple historical normal sample sub-data, and then the sliding window features within the period of each historical normal sample sub-data are extracted, such as statistical features such as difference, average value, variance, sum, median, quartile, maximum value, minimum value, etc. The statistical features of each historical normal sample sub-data are combined. For example, if the historical normal sample data is divided into n historical normal sample sub-data, and for each historical normal sample sub-data, k statistical features are extracted, then these statistical features are combined into an input feature with n rows and k columns.
[0111] After inputting the input feature into the regression model, predict the sample normal network monitoring index value at the historical prediction moment. If the input feature is n rows and k columns, then the predicted sample normal network monitoring index value is n rows and 1 column. Furthermore, adjust the parameters of the regression model based on the difference between the sample normal network monitoring index value and the sample actual network monitoring index value at the historical prediction moment.
[0112] After the regression model is trained, input the neighbor normal data corresponding to a moment to be analyzed into the trained regression model, and predict the normal network monitoring index value at this moment.
[0113] Among them, the neighbor normal data is obtained by removing abnormal data from the one-dimensional time series data within a preset time period before a moment in a single-dimensional time series data.
[0114] Specifically, the preset duration is the periodic duration corresponding to the one-dimensional time series data. For example, if the period of the one-dimensional time series data is 50 minutes, the normal neighbor data at each moment in this one-dimensional time series data is 50 minutes long. When removing abnormal data from the one-dimensional time series data within the preset duration before a moment, first remove the abnormal data in the one-dimensional time series data within the preset duration. At this time, the original abnormal data is set to empty, and then fill the emptied data according to the mean, mode, and other data of the one-dimensional time series data to obtain adjacent normal data.
[0115] For example, if the period of a unit time series data is 50 minutes and the sampling interval of this one-dimensional time series data is 5 minutes, a total of 500 minutes of this one-dimensional time series data is collected. If it is necessary to predict the normal network monitoring index value at the 100th minute of this one-dimensional time series data, then remove the abnormal data from the 10 neighbor data in the 50 minutes before this moment and fill the emptied abnormal data as the neighbor normal data, and input it into the trained regression model to obtain the normal network monitoring index value at the 100th minute.
[0116] In the above embodiment, since historical normal sample data is easy to collect, an XGBoost regression model suitable for semi-supervised learning is used to predict the normal network monitoring index value. Since a machine learning method with a simple structure, that is, the XGBoost regression model, is used for regression prediction, it can run directly on the CPU without relying on the GPU, thereby reducing the required hardware cost and improving the prediction efficiency.
[0117] Furthermore, a self-developed encoder-decoder AE is used to obtain the error threshold of the one-dimensional time series data.
[0118] Specifically, it is necessary to input the historical normal sample data corresponding to the one-dimensional time series data into the encoder-decoder AE, and determine the error threshold of this one-dimensional time series data according to the difference between the input historical normal sample data and the output data. The structure of the encoder-decoder AE can be referred to Figure 3 as shown.
[0119] As Figure 3 shown, it is a schematic structural diagram of an encoder-decoder AE provided in the embodiment of the present application. Taking the historical normal sample data as the input (INPUT), first encode the input data through the encoder (Encoded), input the encoder output (Encoded-output) into the decoder (Dncoded), and then obtain the output (OUTPUT).
[0120] Each structure in the encoder-decoder AE is implemented by a dense layer (Dense), with a total of three dense layers. Since its structure is simple, the training speed is fast.
[0121] In the above embodiment, an encoder-decoder AE with a simple structure is used to obtain an error threshold to replace the method of manually setting the error threshold, so that in the 5G scenario, when there are many KPI indicators, a more accurate error threshold can be obtained.
[0122] After obtaining the normal network monitoring index value and the error threshold, the abnormality of the actual network monitoring index value can be judged.
[0123] An optional implementation is to determine the absolute value of the difference between the normal network monitoring index value and the actual network monitoring index value; if the absolute value is greater than the error threshold, it is determined that the actual network monitoring index value is abnormal. An optional calculation formula is as follows:
[0124]
[0125] Among them, f(i) represents whether the actual network monitoring index value at time i is abnormal. If f(i)=0, it means that the actual network monitoring index value at time i is normal. If f(i)=1, it means that the actual network monitoring index value at time i is abnormal. Real_i represents the actual network monitoring index value at time i; Predict_i represents the normal network monitoring index value at time i; Thresh_ae represents the error threshold of the single-dimensional time series data where time i is located.
[0126] In the above embodiment, whether the actual network monitoring index value is abnormal is judged by the degree to which the actual network monitoring index value deviates from the normal network monitoring index value.
[0127] For non-periodic single-dimensional time series data, an unsupervised outlier anomaly detection method is adopted.
[0128] S222: If a single-dimensional time series data is non-periodic data, then by clustering the single-dimensional time series data, according to the clustering result, it is determined whether the actual network monitoring index value at a certain moment in the single-dimensional time series data is abnormal.
[0129] In this application, GMM is used to cluster the single-dimensional time series data.
[0130] When clustering, it can be to cluster the historical normal sample data, and according to the clustering result, judge whether the actual network monitoring index value at a certain moment in the single-dimensional time series data is abnormal; it can also be to cluster the historical normal sample data and the single-dimensional time series data together, and then according to the clustering result, judge whether the actual network monitoring index value at a certain moment in the single-dimensional time series data is abnormal.
[0131] Specifically, by clustering historical normal sample data, a first preset number of clustering clusters are obtained; by comparing the sizes of the first preset number of clustering clusters, a first abnormal cluster is determined; if the actual network monitoring index value at a moment in a one-dimensional time series data belongs to the first abnormal cluster, it is determined that the actual network monitoring index value at a moment is abnormal; or
[0132] By clustering a one-dimensional time series data and historical normal sample data, a second preset number of clustering clusters are obtained; by comparing the sizes of the second preset number of clustering clusters, a second abnormal cluster is determined; if the actual network monitoring index value at a moment in a one-dimensional time series data is in the second abnormal cluster, it is determined that the actual network monitoring index value at a moment is abnormal.
[0133] Both the first preset number and the second preset number are 2, and the first abnormal cluster and the second abnormal cluster refer to the clusters with less data distribution among the 2 clustering clusters.
[0134] For example, for a non-periodic one-dimensional time series data, by clustering its historical normal sample data, 2 clustering clusters are obtained. By counting the sizes of these 2 clusters, the cluster with less data distribution is used as the first abnormal cluster. If the actual network monitoring index value at a moment in this one-dimensional time series data belongs to the first abnormal cluster, the actual network monitoring index value at this moment is abnormal.
[0135] Or, for a non-periodic one-dimensional time series data, by clustering its historical normal sample data and this one-dimensional time series data, 2 clustering clusters are obtained. By counting the sizes of these 2 clusters, the cluster with less data distribution is used as the second abnormal cluster, and then the actual network monitoring index value at the corresponding moment in this one-dimensional time series data that is in the second abnormal cluster is abnormal.
[0136] The above is the process of abnormal judgment for periodic and non-periodic one-dimensional time series data without fixed index requirements through dynamic thresholds. This process can refer to Figure 4 The abnormal judgment flow chart of the dynamic threshold shown.
[0137] As Figure 4 shown, it is an abnormal judgment flow chart of a dynamic threshold provided by an embodiment of the present application. For each one-dimensional time series data in the dataset to be analyzed, first, the missing data values in the one-dimensional time series data are filled, and then the one-dimensional time series data is automatically identified for periodicity to determine whether it has periodicity.
[0138] If the one-dimensional time series data has periodicity, extract the period value. By removing noise from the historical ordinary sample data, obtain the historical normal sample data, perform feature extraction on the historical normal sample data, realize the training of the extreme gradient boosting regression model, and then obtain the normal network monitoring index value at a certain moment in this one-dimensional time series data. Reconstruct the error of the historical normal sample data through the self-developed encoder-decoder AE, so as to obtain the error threshold of this one-dimensional time series data. Among them, obtaining the normal network monitoring index value and the error threshold can be obtained in an offline state. In an online state, determine whether the actual network monitoring index value at the corresponding moment in this one-dimensional time series data is abnormal according to the normal network monitoring index value and the error threshold.
[0139] If the one-dimensional time series data is non-periodic, obtain the historical ordinary sample data, and use the Gaussian mixture model to cluster the historical normal sample data, or cluster the historical normal sample data and the one-dimensional time series data to obtain a preset number of clustering clusters, determine the abnormal clusters among them, and realize outlier anomaly detection by judging whether the actual network monitoring index values at each moment in the one-dimensional time series data belong to the abnormal clusters.
[0140] In addition, for one-dimensional time series data with solid state index requirements, use static thresholds for anomaly detection.
[0141] Specifically, if there are fixed index requirements for the one-dimensional time series data, then if the actual network monitoring index value at a certain moment in a one-dimensional time series data does not meet the fixed index requirements, determine that the actual network monitoring index value is abnormal.
[0142] For example, for the KPI index of CPU utilization rate, if there are solid state index requirements, for example, its solid state index requirement is that the CPU utilization rate reaches 99%, then according to this fixed index requirement, judge the abnormality of the actual network monitoring index values at each moment in the one-dimensional time series data. If the actual network monitoring index value at a certain moment does not meet the fixed index requirement that the CPU utilization rate reaches 99%, then the actual network monitoring index value at that moment is abnormal. If the actual network monitoring index value at a certain moment reaches the fixed index requirement that the CPU utilization rate reaches 99%, then the actual network monitoring index value at that moment is normal.
[0143] As Figure 5 shown, it is a flowchart of anomaly detection for one-dimensional time series data provided by the embodiments of the present application. In the embodiments of the present application, for the dataset to be analyzed, first select a certain KPI index and judge whether there are fixed requirement indexes. If there are fixed index requirements for this KPI index, use static thresholds to judge the abnormality of the KPI index; if there are no fixed index requirements for this KPI index, use the dynamic threshold detection algorithm to judge the abnormality of the KPI index.
[0144] In the above embodiments, one-dimensional time series data is screened according to solid state index requirements, and static thresholds are used for anomaly detection of one-dimensional time series data with solid state index requirements.
[0145] As Figure 6A shown, it is a schematic diagram of an anomaly detection result provided by an embodiment of the present application. Figure 6A The middle is the result of anomaly detection for a certain one-dimensional time series data, where the horizontal axis of the data is time and the vertical axis is the data value. Figure 6A The upper-side data in it is the actual network monitoring index value of a certain one-dimensional time series data, and there are 7 actual abnormal data among them. Figure 6A The lower-side data in it is the normal network monitoring index value of the predicted one-dimensional time series data, and 3 detected abnormal data.
[0146] As Figure 6B shown, it is another schematic diagram of an anomaly detection result provided by an embodiment of the present application. Figure 6B The middle is the result of anomaly detection for a certain one-dimensional time series data, where the horizontal axis of the data is time and the vertical axis is the data value. Figure 6B The upper-side data in it is the actual network monitoring index value of the one-dimensional time series data, and the white area is the actual abnormal data. Figure 6B The lower-side data in it is the normal network monitoring index value of the predicted one-dimensional time series data, and the white area is the predicted abnormal data.
[0147] As Figure 7 shown, it is yet another schematic diagram of an anomaly detection result provided by an embodiment of the present application. Figure 7 It is the result of outlier anomaly detection for the one-dimensional time series data of the inter-station handover success rate. Figure 7 In it, the horizontal axis of the data is time. This one-dimensional time series data is collected every 15 minutes, and the vertical axis is the inter-station handover success rate. The inter-station handover success rate of this one-dimensional time series data is distributed between 92% and 100%. It can be seen that the fixed index requirement for this one-dimensional time series data is an inter-station handover success rate of 99%. For outlier anomaly detection of this one-dimensional time series data, Figure 7 the black data points in it are abnormal detection results, and the white data points are normal detection results. The actual abnormal data are the data points below the dotted line, and the actual normal data are the data points above the dotted line. From Figure 7 it can be seen that the accuracy rate of this outlier anomaly detection is very high.
[0148] In the embodiments of the present application, first, it is determined whether there are fixed index requirements for a certain one-dimensional time series data in the dataset to be analyzed. If there are fixed index requirements, static thresholds are used for anomaly detection; if there are no fixed index requirements, dynamic thresholds are used for anomaly detection.
[0149] For one-dimensional time series data without fixed index requirements, it is further determined whether it has periodicity. For the actual network monitoring index value at each moment in the periodic one-dimensional time series data, first predict the normal network monitoring index value at this moment, and then reconstruct the error of the historical normal sample data according to the trained encoder-decoder to obtain an error threshold. A machine learning method with a simple structure is used to replace the method of manually setting the error threshold, and this method runs directly on the Central Processing Unit (CPU) without relying on the GPU. Finally, through this normal network monitoring index value and the error threshold, it is possible to judge the degree to which the actual network monitoring index value at this moment deviates from the normal network monitoring index value, and judge whether the actual network monitoring index value is abnormal according to the degree of deviation.
[0150] For the actual network monitoring index value at each moment in the non-periodic one-dimensional time series data, determine the abnormal cluster by clustering, and then judge whether the actual network monitoring index value at each moment in the one-dimensional time series data is abnormal.
[0151] In summary, this application mainly aims at the anomaly detection problem in the actual 5G scenario, where there are numerous KPI indicators, possibly reaching tens of thousands, and they have characteristics such as large differences, unknown periodicity, changes, and growth. A data detection method with low computational complexity that can run directly on the CPU without relying on the GPU is proposed, which reduces the hardware cost and improves the detection efficiency. It can automatically monitor and give early warnings of anomalies for numerous KPI indicators in the 5G scenario, thereby achieving the purpose of improving the efficiency of fault location and reducing the actual operation and maintenance costs at the same time.
[0152] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in an embodiment of this application. In one embodiment, the electronic device can be Figure 1 the terminal device 110 shown. In this embodiment, the structure of the electronic device can be as Figure 8 shown, including a memory 801, a communication module 803, and one or more processors 802.
[0153] The memory 801 is used to store the computer program executed by the processor 802. The memory 801 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.
[0154] The memory 801 may be a volatile memory, such as a random-access memory (RAM); the memory 801 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 801 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 801 may be a combination of the above memories.
[0155] The processor 802 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 802 is used to implement the above data detection method when calling the computer program stored in the memory 801.
[0156] The communication module 803 is used to communicate with the terminal device and other servers.
[0157] In the embodiments of the present application, the specific connection medium between the above memory 801, communication module 803, and processor 802 is not limited. In the embodiments of the present application Figure 8 it is connected between the memory 801 and the processor 802 through a bus 804, and the bus 804 is described in thick lines in Figure 8 The connection manners between other components are only for illustrative purposes and are not to be taken as limiting. The bus 804 may be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 8 only a thick line is used to describe it in
[0158] The memory 801 stores a computer storage medium, and the computer storage medium stores computer-executable instructions, and the computer-executable instructions are used to implement the data detection method of the embodiments of the present application. The processor 802 is used to execute the above data detection method.
[0159] In some implementation manners, after the method executed by the above processor forms a program, the hardware execution modules corresponding to each program function module may include: an acquisition module and an execution module, and the acquisition module and the execution module are connected.
[0160] As Figure 9 shown, it is a schematic structural diagram of a data detection device 900 in the embodiments of the present application. The data detection device includes an acquisition module 901 and an execution module 902, where:
[0161] An acquisition module 901, configured to acquire a data set to be analyzed through network monitoring; the data set includes at least one item of one-dimensional time-series data; each item of one-dimensional time-series data corresponds to a network monitoring metric;
[0162] An execution module 902, configured to perform the following operations respectively on the actual network monitoring metric value at each moment in each item of one-dimensional time-series data:
[0163] If an item of one-dimensional time-series data is periodic data, then according to the trained regression model and the neighbor normal data corresponding to a moment, predict the normal network monitoring metric value at a moment; according to the trained encoder-decoder, perform error reconstruction on the historical normal sample data to determine the error threshold of an item of one-dimensional time-series data; according to the error threshold and the predicted normal network monitoring metric value at a moment, determine whether the actual network monitoring metric value at a moment is abnormal; the neighbor normal data is obtained by removing abnormal data from the one-dimensional time-series data within a preset time period before a moment in an item of one-dimensional time-series data; the historical normal sample data is obtained by removing abnormal data from the historical ordinary sample data;
[0164] If an item of one-dimensional time-series data is non-periodic data, then perform clustering on the item of one-dimensional time-series data, and determine whether the actual network monitoring metric value at a moment in the item of one-dimensional time-series data is abnormal according to the clustering result.
[0165] Optionally, the execution module 902 is configured to determine whether an item of one-dimensional time-series data is periodic data through the following operations:
[0166] Based on a preset correlation rule, perform correlation analysis on an item of one-dimensional time-series data to determine the correlation value of the item of one-dimensional time-series data; the correlation value characterizes the periodicity in the one-dimensional time-series data;
[0167] If the correlation value is greater than a preset correlation threshold, then determine that an item of one-dimensional time-series data is periodic data;
[0168] If the correlation value is not greater than the preset correlation threshold, then determine that an item of one-dimensional time-series data is non-periodic data.
[0169] Optionally, the execution module 902 is specifically configured to:
[0170] Perform clustering on the historical ordinary sample data to obtain a first preset number of clustering clusters; determine the first abnormal cluster by comparing the sizes of the first preset number of clustering clusters; if the actual network monitoring metric value at a moment in an item of one-dimensional time-series data belongs to the first abnormal cluster, then determine that the actual network monitoring metric value at a moment is abnormal; or
[0171] By clustering a single-dimensional time series data and historical normal sample data, obtain a second preset number of clustering clusters; by comparing the sizes of the second preset number of clustering clusters, determine the second abnormal cluster; if the actual network monitoring index value at a moment in a single-dimensional time series data is in the second abnormal cluster, then determine that the actual network monitoring index value at the moment is abnormal.
[0172] Optionally, if there are fixed index requirements for the single-dimensional time series data, the execution module 902 is further configured to:
[0173] If the actual network monitoring index value at a moment in a single-dimensional time series data does not meet the fixed index requirements, then determine that the actual network monitoring index value is abnormal.
[0174] Optionally, before the execution module 902 determines whether the single-dimensional time series data is periodic data, the execution module 902 is further configured to:
[0175] If there are missing data values in a single-dimensional time series data, determine a supplementary value according to a preset supplementary rule and the single-dimensional time series data;
[0176] Fill the missing data values with the supplementary value.
[0177] Optionally, the execution module 902 is configured to train a regression model through the following operations:
[0178] If a single-dimensional time series data is periodic data, perform feature extraction on the historical normal sample data to obtain a third preset number of statistical features;
[0179] Combine the third preset number of statistical features to obtain input features;
[0180] Based on the input features and the regression model, predict the sample normal network monitoring index value at the historical prediction moment;
[0181] Based on the difference between the sample normal network monitoring index value at the historical prediction moment and the sample actual network monitoring index value at the historical prediction moment, adjust the parameters of the regression model.
[0182] Optionally, the execution module 902 is specifically configured to:
[0183] Determine the absolute value of the difference between the normal network monitoring index value and the actual network monitoring index value;
[0184] If the absolute value is greater than the error threshold, then determine that the actual network monitoring index value is abnormal.
[0185] For the convenience of description, the above - mentioned parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware. Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0186] Next, refer to Figure 10 to describe the computing device 1000 according to this embodiment of the present application. Figure 10 The computing device 1000 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0187] As Figure 10 , the computing device 1000 is presented in the form of a general - purpose computing device. The components of the computing device 1000 may include, but are not limited to: at least one of the above - mentioned processing units 1001, at least one of the above - mentioned storage units 1002, and a bus 1003 connecting different system components (including the storage unit 1002 and the processing unit 1001).
[0188] The bus 1003 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.
[0189] The storage unit 1002 may include a readable medium in the form of volatile memory, such as a random - access memory (RAM) 1021 and / or a cache memory 1022, and may further include a read - only memory (ROM) 1023.
[0190] The storage unit 1002 may further include a program / utility 1025 having a set (at least one) of program modules 1024. Such program modules 1024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0191] The computing device 1000 may also communicate with one or more external devices 1004 (such as a keyboard, a pointing device, etc.), and may also communicate with one or more devices that enable a user to interact with the computing device 1000, and / or communicate with any device that enables the computing device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 1005. Moreover, the computing device 1000 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1006. As Figure 10 shown, the network adapter 1006 communicates with other modules for the computing device 1000 through the bus 1003. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computing device 1000, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0192] Embodiments of the present application also provide a computer program product. The methods in the present application may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the present application are executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM, or other programmable devices.
[0193] The computer-readable storage medium may be an implementation of the computer program product. That is, embodiments of the present application also provide a computer-readable storage medium, which includes a computer program, and when the computer program is executed by a processor, the data detection method as described above is implemented.
[0194] The computer program or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data detection device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0195] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0196] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0197] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0198] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 of the functions specified in one process or a plurality of processes and / or blocks.
[0199] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A data detection method, characterized in that, The method includes: Obtaining a data set to be analyzed through network monitoring; the data set includes at least one item of one-dimensional time series data; each item of one-dimensional time series data corresponds to a network monitoring metric; For the actual network monitoring metric value at each moment in each item of one-dimensional time series data, the following operations are respectively performed: If an item of one-dimensional time series data is periodic data, then according to the trained regression model and the neighbor normal data corresponding to a moment, predict the normal network monitoring metric value at the moment; reconstruct the error of the historical normal sample data according to the trained encoder-decoder, and determine the error threshold of the item of one-dimensional time series data; according to the error threshold and the predicted normal network monitoring metric value at the moment, determine whether the actual network monitoring metric value at the moment is abnormal; the neighbor normal data is obtained by removing abnormal data from the one-dimensional time series data within a preset time period before the moment in the item of one-dimensional time series data; the historical normal sample data is obtained by removing abnormal data from the historical ordinary sample data; If the item of one-dimensional time series data is non-periodic data, then determine whether the actual network monitoring metric value at a moment in the item of one-dimensional time series data is abnormal according to the clustering result by clustering the item of one-dimensional time series data.
2. The method according to claim 1, wherein Determine whether the item of one-dimensional time series data is periodic data through the following operations: Based on a preset correlation rule, perform a correlation analysis on the item of one-dimensional time series data to determine the correlation value of the item of one-dimensional time series data; the correlation value characterizes the periodicity in the one-dimensional time series data; If the correlation value is greater than the preset correlation threshold, then determine that the item of one-dimensional time series data is periodic data; If the correlation value is not greater than the preset correlation threshold, then determine that the item of one-dimensional time series data is non-periodic data.
3. The method according to claim 1, wherein The determining whether the actual network monitoring metric value at a moment in the item of one-dimensional time series data is abnormal according to the clustering result by clustering the item of one-dimensional time series data includes: Obtain a first preset number of clustering clusters by clustering the historical ordinary sample data; determine the first abnormal cluster by comparing the sizes of the first preset number of clustering clusters; if the actual network monitoring metric value at a moment in the item of one-dimensional time series data belongs to the first abnormal cluster, then determine that the actual network monitoring metric value at the moment is abnormal; or Obtain a second preset number of clustering clusters by clustering the item of one-dimensional time series data and the historical ordinary sample data; determine the second abnormal cluster by comparing the sizes of the second preset number of clustering clusters; if the actual network monitoring metric value at a moment in the item of one-dimensional time series data is in the second abnormal cluster, then determine that the actual network monitoring metric value at the moment is abnormal.
4. The method according to claim 1, characterized in that If there are fixed metric requirements for the one-dimensional time series data, then the method further includes: If the actual network monitoring metric value at a moment in the item of one-dimensional time series data does not meet the fixed metric requirements, then determine that the actual network monitoring metric value is abnormal.
5. The method according to claim 1, characterized in that Before determining whether the one-dimensional time series data is periodic data, the method further includes: If there is a missing data value in the one-dimensional time series data, determine a supplementary value according to a preset supplementary rule and the one-dimensional time series data; Fill the missing data value with the supplementary value.
6. The method according to claim 1, wherein Train a regression model through the following operations: If one-dimensional time series data is periodic data, perform feature extraction on the historical normal sample data to obtain a third preset number of statistical features; Combine the third preset number of statistical features to obtain input features; Based on the input features and the regression model, predict the sample normal network monitoring index value at the historical prediction time; Based on the difference between the sample normal network monitoring index value at the historical prediction time and the sample actual network monitoring index value at the historical prediction time, adjust the parameters of the regression model.
7. The method according to any one of claims 1 to 6, characterized in that, The determining whether the actual network monitoring index value at a certain time is abnormal according to the error threshold and the normal network monitoring index value at the certain time includes: Determine the absolute value of the difference between the normal network monitoring index value and the actual network monitoring index value; If the absolute value is greater than the error threshold, determine that the actual network monitoring index value is abnormal.
8. A data detection device, characterized in that, Includes: An acquisition module, configured to acquire a data set to be analyzed through network monitoring; the data set includes at least one one-dimensional time series data; Each one-dimensional time series data corresponds to a network monitoring index; An execution module, configured to perform the following operations respectively on the actual network monitoring index value at each time in each one-dimensional time series data: If one-dimensional time series data is periodic data, predict the normal network monitoring index value at the certain time according to the trained regression model and the neighboring normal data corresponding to the certain time; determine the error threshold of the one-dimensional time series data according to the error reconstruction of the historical normal sample data by the trained encoder-decoder; determine whether the actual network monitoring index value at the certain time is abnormal according to the error threshold and the predicted normal network monitoring index value at the certain time; the neighboring normal data is obtained by removing abnormal data from the one-dimensional time series data within a preset time period before the certain time in the one-dimensional time series data; the historical normal sample data is obtained by removing abnormal data from the historical ordinary sample data; If the one-dimensional time series data is non-periodic data, determine whether the actual network monitoring index value at a certain time in the one-dimensional time series data is abnormal according to the clustering result by clustering the one-dimensional time series data.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.