A KPI time series detection method and related device

By obtaining the data characteristics of the KPI time series and using the random forest algorithm to train the model, combining the similarity matrix and ensemble learning, the generalization ability problem of KPI time series detection is solved, achieving higher detection accuracy and diversity.

CN115130606BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210857783.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-07-11
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

The existing KPI time series anomaly detection scheme lacks generalization capabilities, which requires long-term experience accumulation of operators and set different alarm thresholds for each KPI time series, resulting in poor detection results.

Method used

By obtaining the KPI time series of the time points to be detected and the sequence set of historical time periods, the data features are extracted to generate a training set, and the anomaly detection model is trained using a random forest algorithm, the similarity matrix is used for detection, and the final result is output in combination with the integrated learning idea.

Benefits of technology

It improves the generalization performance and detection accuracy of KPI time series anomaly detection, reduces dependence on operational experience, and enhances the diversity and accuracy of detection algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115130606B_ABST
    Figure CN115130606B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a KPI time series detection method and related devices, which are used to improve the general generalization ability of the KPI time series detection method. The method includes: obtaining a first KPI time series at a time point to be detected and a set of KPI time series within a preset time period before the time point to be detected; extracting data features of the first KPI time series to generate a test feature vector, and extracting data features of the set of KPI time series to generate a first training set; obtaining a positive sample set and a negative sample set in the first training set; sampling according to the positive sample set and the negative sample set to obtain a second training set; training a random forest algorithm using the second training set to obtain an anomaly detection model; inputting the second training set and the test feature vector into the anomaly detection model to calculate a similarity matrix; and determining a first detection result of the first KPI time series according to the similarity matrix, positive samples, and negative samples. The present application can be applied to the fields of cloud technology, artificial intelligence, and big data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of operation and maintenance monitoring, and particularly to a KPI time series detection method and related devices. Background Art

[0002] For a long time, the anomaly detection and warning of key performance indicators (KPIs) have been a hot topic in academic and industrial research. Especially in the Internet field, it is necessary to monitor the fluctuations of data such as product behavior logs, application information, Polaris metrics, and function point usage in a timely manner every day. Once an abnormal situation occurs in the KPI time series, relevant personnel should be alerted and adjusted immediately. The KPI time series is also an indicator that product, operation, and development personnel need to pay attention to at all times. Through the fluctuations of the KPI time series, it is possible to analyze in a timely manner whether there are abnormalities in the product system, operation strategy, etc.

[0003] The current process for anomaly detection of KPI time series is as follows: Combining the data characteristics of different KPI time series to determine whether to use other methods such as month-on-month or year-on-year comparison. This solution requires determining the threshold for anomaly warning in advance, and only when the fluctuation exceeds the threshold will it be recognized as an anomaly.

[0004] The determination of the threshold requires the long-term experience accumulation of operators. The thresholds for each KPI time series are not exactly the same. Operation and maintenance personnel need to deploy specific algorithms for each KPI time series one by one and set different warning thresholds. Therefore, the current anomaly detection solution usually has good detection effects only for specific types of KPI time series anomalies, but lacks good general generalization ability. Summary of the Invention

[0005] Embodiments of this application provide a KPI time series detection method and related devices, which are used to improve the general generalization ability of the KPI time series detection method.

[0006] In view of this, on the one hand, the present application provides a KPI time series detection method, including: obtaining the first key performance indicator (KPI) time series at the time point to be detected and the KPI time series set within a preset time period before the time point to be detected; extracting the data features of the first KPI time series to generate a test feature vector, and extracting the data features of each KPI time series in the KPI time series set to generate a first training set; obtaining the positive sample set and the negative sample set in the first training set, where the positive sample set includes KPI time series samples at normal times, and the negative sample set includes KPI time series samples at abnormal times; sampling the positive sample set and the negative sample set to obtain a second training set, where the number of positive samples in the second training set is the same as the number of negative samples; using the second training set to train a random forest algorithm to obtain an anomaly detection model; inputting the second training set and the test feature vector into the anomaly detection model to calculate the similarity matrix corresponding to the second training set and the test feature vector; and determining the first detection result of the first KPI time series according to the similarity matrix, the positive samples in the second training set, and the negative samples in the second training set.

[0007] On the other hand, the present application provides a detection device, including: a first acquisition module, configured to obtain the first key performance indicator (KPI) time series at the time point to be detected and the KPI time series set within a preset time period before the time point to be detected;

[0008] a feature extraction module, configured to extract the data features of the first KPI time series to generate a test feature vector, and extract the data features of each KPI time series in the KPI time series set to generate a first training set;

[0009] a second acquisition module, configured to obtain the positive sample set and the negative sample set in the first training set, where the positive sample set includes KPI time series samples at normal times, and the negative sample set includes KPI time series samples at abnormal times;

[0010] a sampling module, configured to sample the positive sample set and the negative sample set to obtain a second training set, where the number of positive samples in the second training set is the same as the number of negative samples;

[0011] a training module, configured to use the second training set to train a random forest algorithm to obtain an anomaly detection model;

[0012] A detection module, configured to input the second training set and the test feature vector into the anomaly detection model to calculate a similarity matrix corresponding to the second training set and the test feature vector; and determine a first detection result of the first KPI time series according to the similarity matrix, positive class samples in the second training set, and negative class samples in the second training set.

[0013] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the feature extraction module is specifically configured to extract statistical features, fitting features, and original features of the first KPI time series to generate the test feature vector;

[0014] The feature extraction module is specifically configured to extract statistical features, fitting features, and original features of each KPI time series in the KPI time series set to generate the first training set.

[0015] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the sampling module is further configured to sample the positive class sample set and the negative class sample set to obtain a third training set, where the number of positive class samples in the third training set is the same as the number of negative class samples;

[0016] The training module is further configured to use the third training set to train a random forest algorithm to update the anomaly detection model;

[0017] The detection module is further configured to input the third training set and the test feature vector into the updated anomaly detection model to output a second detection result of the first KPI time series, and use the first detection result and the second detection result as a detection result set; and so on, until the number of detection results in the detection result set reaches a preset number, and then determine a final detection result of the first KPI time series according to the detection results in the detection result set.

[0018] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the detection module is specifically configured to

[0019] Calculate a first similarity of positive class samples in the second training set and a second similarity of negative class samples in the second training set according to the similarity matrix;

[0020] Determine a first classification result of the test feature vector according to the first similarity and the second similarity;

[0021] Determine a first detection result of the first KPI time series according to the first classification result.

[0022] In a possible design, in another implementation of another aspect of the embodiments of the present application, the detection module is specifically configured to input the second training set and the test feature vector into the anomaly detection model to calculate a first proportional value that the first sample in the second training set and the first sample corresponding to the test feature vector fall into the same leaf node in the anomaly detection model, and use the first proportional value as the similarity between the first sample in the second training set and the first sample corresponding to the test feature vector;

[0023] Input the second training set and the test feature vector into the anomaly detection model to calculate a second proportional value that the second sample in the second training set and the second sample corresponding to the test feature vector fall into the same leaf node in the anomaly detection model, and use the second proportional value as the similarity between the second sample in the second training set and the second sample corresponding to the test feature vector;

[0024] And so on, until the similarities between each sample in the second training set and each sample corresponding to the test feature vector are calculated by traversal, and the similarities between each sample in the second training set and each sample corresponding to the test feature vector are summarized to obtain the similarity matrix.

[0025] In a possible design, in another implementation of another aspect of the embodiments of the present application, the detection module is specifically configured to determine a similarity set corresponding to the positive class samples in the second training set from the similarity matrix, and sum up each similarity in the similarity set corresponding to the positive class samples to obtain the first similarity;

[0026] Determine a similarity set corresponding to the negative class samples in the second training set from the similarity matrix, and sum up each similarity in the similarity set corresponding to the negative class samples to obtain the second similarity.

[0027] In a possible design, in another implementation of another aspect of the embodiments of the present application, the detection module is specifically configured to, when the first similarity is greater than the second similarity, determine that the first classification result of the test feature vector is normal;

[0028] When the first similarity is less than the second similarity, determine that the first classification result of the test feature vector is abnormal.

[0029] In a possible design, in another implementation of another aspect of the embodiments of the present application, the detection module is specifically configured to obtain a first value and a second value according to the detection results in the detection result set, where the first value is the number of detection results indicating that the first KPI time series is a normal KPI time series, and the second value is the number of detection results indicating that the first KPI time series is an abnormal KPI time series;

[0030] When the first value is greater than the second value, it is determined that the first KPI time series is a normal time series;

[0031] When the first value is less than the second value, it is determined that the first KPI time series is an abnormal time series.

[0032] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;

[0033] Wherein, the memory is used to store programs;

[0034] The processor is used to execute the programs in the memory, and the processor is used to execute the methods of the above aspects according to the instructions in the program code;

[0035] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate with each other.

[0036] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is enabled to execute the methods of the above aspects.

[0037] Another aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0038] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages: In the normal sequence samples and abnormal sequence samples, a balanced data set is generated by sampling, and then the sampled balanced data set is used as the training data of the random forest algorithm. It avoids directly using the original unbalanced data set as the training set, which may lead to an increase in the imbalance degree of the data set during the sampling process of the random forest algorithm. Therefore, it is also beneficial to improve the diversity of the data, enrich the differences between different models, and improve the generalization performance of the KPI time series anomaly detection algorithm. Description of the Drawings

[0039] Figure 1 It is a schematic diagram of the architecture of the solution implementation system in the embodiment of the present application;

[0040] Figure 2 It is a schematic diagram of an embodiment of the KPI time series detection method in the embodiment of the present application;

[0041] Figure 3 It is a schematic diagram of an embodiment of the detection device in the embodiment of the present application;

[0042] Figure 4 It is another schematic diagram of an embodiment of the detection device in the embodiment of the present application;

[0043] Figure 5 It is another schematic diagram of an embodiment of the detection device in the embodiment of the present application. Detailed implementation manners

[0044] The embodiment of the present application provides a KPI time series detection method and related devices, which are used to improve the general generalization ability of the KPI time series detection method.

[0045] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "correspond to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0046] For the convenience of understanding, some professional terms in the embodiment of the present application are described below:

[0047] KPI refers to a special time series with practical application significance obtained by regular sampling, such as the number of website visits per unit time, the number of daily active users (Daily Active User, DAU) of a product, the transaction volume per unit time, the transaction volume in a transaction service system, the page view volume of a website in a website service system, and so on.

[0048] For a long time, anomaly detection and alarm of key performance indicators (KPIs) have been a hot topic in academic and industrial research. Especially in the Internet field, it is necessary to monitor the fluctuations of data such as product behavior logs, application information, Polaris metrics, and function point usage in a timely manner every day. Once an anomaly occurs in the KPI time series, relevant personnel should be alerted immediately and adjustments should be made. The KPI time series is also an indicator that product, operation, and development personnel need to pay attention to at all times. Through the fluctuations of the KPI time series, it is possible to analyze in a timely manner whether there are anomalies in the product system and operation strategies, etc. The current process for anomaly detection of KPI time series is as follows: Combining the data characteristics of different KPI time series to determine whether to use other methods such as month-on-month or year-on-year. This solution requires determining the threshold for anomaly alarm in advance. Only when the fluctuation exceeds the threshold will it be recognized as an anomaly. The determination of the threshold requires the long-term experience accumulation of operators, and the thresholds for each KPI time series are not exactly the same. The operation and maintenance personnel need to deploy specific algorithms for each KPI time series one by one and set different alarm thresholds. Therefore, the current anomaly detection solution usually has good detection effects only for specific types of KPI time series anomalies, but lacks good general generalization ability.

[0049] To solve this technical problem, an embodiment of the present application provides a KPI time series detection method, including: obtaining the first key performance indicator KPI time series at the time point to be detected and the set of KPI time series within a preset time period before the time point to be detected; extracting the data characteristics of the first KPI time series to generate a test feature vector, and extracting the data characteristics of each KPI time series in the set of KPI time series to generate a first training set; obtaining the positive sample set and the negative sample set in the first training set, where the positive sample set includes KPI time series samples at normal times, and the negative sample set includes KPI time series samples at abnormal times; sampling the positive sample set and the negative sample set to obtain a second training set, where the number of positive samples in the second training set is the same as the number of negative samples; training a random forest algorithm using the second training set to obtain an anomaly detection model; inputting the second training set and the test feature vector into the anomaly detection model to calculate the similarity matrix corresponding to the second training set and the test feature vector; and determining the first detection result of the first KPI time series according to the similarity matrix, the positive samples in the second training set, and the negative samples in the second training set.

[0050] The method provided by the present application can be applied to, for example Figure 1The system architecture shown in the figure includes a terminal device and a detection device. Among them, the terminal device can be set as an enterprise device or a device running enterprise software. The enterprise software can run on the terminal device in the form of a browser or in the form of an independent application (APP), etc. The specific presentation form of the enterprise software is not limited here. The detection device can be independent of the terminal device and interact with the terminal device in the form of a server for data; or, the detection device is integrated inside the terminal device and interacts with the terminal device in the form of a client for data; or, the detection device can be independent of the terminal device and interact with the terminal device in the form of a terminal device for data. When the detection device and the terminal device implement the method provided in this application, the detection device obtains the KPI time series generated when the terminal device operates as an enterprise device or the detection device obtains the KPI time series generated when the terminal device runs enterprise software, and then the detection device detects whether the KPI time series is an abnormal time series. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a palm computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. The number of servers and terminal devices is also not limited. The solution provided in this application can be completed independently by the terminal device or completed in cooperation with the server by the terminal device. Regarding this, this application does not make specific limitations.

[0051] It can be understood that in the specific implementation of this application, related data such as KPI time series are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0052] Combined with the above introduction, the KPI time series detection method in this application will be introduced below. Please refer to Figure 2 , an embodiment of the KPI time series detection method in the embodiment of this application includes:

[0053] 201. Obtain the first key performance indicator (KPI) time series at the time point to be detected and the set of KPI time series within a preset time period before the time point to be detected.

[0054] During the operation of the device or application, the detection device can perform regular sampling on the KPI time series of the device or application, that is, the detection device obtains the first KPI time series at the time point to be detected and detects the regularly sampled first KPI time series. In this embodiment, the detection device also needs to obtain the set of KPI time series within a preset time period before the time point to be detected while obtaining the first KPI time series.

[0055] It can be understood that the set of KPI time series is used to update the anomaly detection model once, and then the first KPI time series is detected according to the updated anomaly detection model. Among them, the preset time period can be a time period close to the time point to be detected. For example, when the detection device needs to detect the KPI time series of application A at 10:30 am on July 5th, the preset time period can be from 9:00 am to 10:30 am on July 5th. And the set of KPI time series can be the set of KPI time series generated during the operation of the application within the above time period.

[0056] In an exemplary solution, the first KPI time series may include the operation log of application A at 10:30 am. For example, at this moment, the number of accesses to the application is 1000 times, the number of online accesses to the application is 10 million, the number of transactions generated in the application is 1000 times, and so on.

[0057] 202. Extract the data features of the first KPI time series to generate a test feature vector, and extract the data features of each KPI time series in the set of KPI time series to generate a first training set.

[0058] The detection device extracts data features from the first KPI time series and generates the test feature vector according to the data features. Similarly, the detection device uses the same feature extraction method to extract data features from each KPI time series in the set of KPI time series and generates the first training set according to the data features of the set of KPI time series.

[0059] Optionally, when the detection device extracts features from the first KPI time series and the KPI time series set, in order to enrich the feature quantity of the measurement feature vector and the samples in the first training set, the detection device can extract features of multiple attribute features such as statistical features, fitting features, and original features from the first KPI time series and the KPI time series, so as to construct a first training set with time features as independent variables and whether it is abnormal as dependent variables, and a test feature vector for which no detection result is obtained. In this embodiment, the statistical features can be used to indicate data features obtained through statistics such as the maximum value in each data, the minimum value in each data, and so on. The fitting features can be data features obtained through integrated statistics of the change trend of each data in the KPI time series and the seasonal change of each data. The original features can be used to indicate the values of each data in the KPI time series, or the features of the user corresponding to the application, such as the gender of the user, the location where the user is located, the version number of the application, and so on.

[0060] 203. Obtain the positive sample set and the negative sample set in the first training set. The positive sample set includes KPI time series samples at normal times, and the negative sample set includes KPI time series samples at abnormal times.

[0061] In this embodiment, the detection device can label each KPI time series sample in the first training set. At this time, the label is used to indicate whether the KPI time series is a normal time series or an abnormal time series. In this way, the KPI time series samples in the first training set labeled with the label indicating the KPI time series as a normal time series are used as positive samples, and the KPI time series samples labeled with the label indicating the KPI time series as an abnormal time series are used as negative samples. Then the detection device summarizes the positive samples into the positive sample set and the negative samples into the negative sample set.

[0062] In an exemplary solution, the first training set includes 100 samples, there are 90 positive samples, and there are 10 negative samples. Then the positive sample set includes 90 positive samples, and the negative sample set includes 10 negative samples.

[0063] 204. Sample the positive sample set and the negative sample set to obtain a second training set, in which the number of positive samples is the same as the number of negative samples.

[0064] To ensure the balance of the training samples, the detection device needs to keep the number of positive samples and negative samples in the training set for training consistent. Therefore, in this embodiment, it is proposed to sample the same number of times from the positive sample set and the negative sample set, and merge the sampled positive samples and negative samples to obtain the second training set.

[0065] In an exemplary solution, if the detection device determines that the number of samples in the second training set is 100, then 50 positive samples need to be sampled from the positive sample set and 50 negative samples need to be sampled from the negative sample set.

[0066] 205. Use the second training set to train a random forest algorithm to obtain an anomaly detection model.

[0067] The detection device passes the sample data in the second training set through the target model corresponding to the random forest algorithm, calculates the loss based on the predicted classification result obtained from the sample data and the true data corresponding to the sample data in the second training set, and then adjusts the parameters of the target model in the reverse direction according to the loss to obtain the anomaly detection model. It can be understood that the training process of the anomaly detection model in this embodiment is the same as that in the prior art, and will not be elaborated here specifically.

[0068] 206. Input the second training set and the test feature vector into the anomaly detection model to calculate the similarity matrix corresponding to the second training set and the test feature vector.

[0069] After the detection device trains the anomaly detection model according to the second training set, input the second training set and the test feature vector into the anomaly detection model simultaneously to obtain the first detection result of the first KPI time series.

[0070] In this embodiment, to ensure the accuracy of the detection result, the detection device uses a similarity judgment method. The specific method can be as follows: The detection device inputs the second training set and the test feature vector into the anomaly detection model to calculate the similarity matrix between each sample in the second training set and the test feature vector.

[0071] Optionally, the specific way for the detection device to calculate the similarity matrix between each sample in the second training set and the test feature vector can be as follows: The detection device inputs the second training set and the test feature vector into the anomaly detection model to calculate the first proportional value that the first sample in the second training set and the first sample corresponding to the test feature vector fall into the same leaf node in the anomaly detection model, and uses the first proportional value as the similarity between the first sample in the second training set and the first sample corresponding to the test feature vector; inputs the second training set and the test feature vector into the anomaly detection model to calculate the second proportional value that the second sample in the second training set and the second sample corresponding to the test feature vector fall into the same leaf node in the anomaly detection model, and uses the second proportional value as the similarity between the second sample in the second training set and the second sample corresponding to the test feature vector; and so on, until the similarities between each sample in the second training set and each sample corresponding to the test feature vector are calculated by traversal, and the similarities between each sample in the second training set and each sample corresponding to the test feature vector are summarized to obtain the similarity matrix. In an exemplary solution, assume that the second training set includes sample A and sample B. When the detection device inputs the second training set and the test feature vector into the anomaly detection model for the first similarity calculation, sample A and the test feature vector fall into the same leaf node in the anomaly detection model, and sample B and the test feature vector fall into different leaf nodes in the anomaly detection model; when the detection device inputs the second training set and the test feature vector into the anomaly detection model for the second similarity calculation, sample A and the test feature vector fall into the same leaf node in the anomaly detection model, and sample B and the test feature vector fall into different leaf nodes in the anomaly detection model; when the detection device inputs the second training set and the test feature vector into the anomaly detection model for the third similarity calculation, sample A and the test feature vector fall into different leaf nodes in the anomaly detection model, and sample B and the test feature vector fall into different leaf nodes in the anomaly detection model; when the detection device inputs the second training set and the test feature vector into the anomaly detection model for the fourth similarity calculation, sample A and the test feature vector fall into different leaf nodes in the anomaly detection model, and sample B and the test feature vector fall into the same leaf node in the anomaly detection model; and so on, repeating the calculation until the preset number of times (such as 100 times) is reached. At this time, count the number of times that sample A and the test feature vector fall into the same leaf node in the anomaly detection model (assume it is 45 times), and calculate the proportion of the number of times that sample A and the test feature vector fall into the same leaf node in the anomaly detection model to the number of calculation times (i.e., 45 / 100 = 0.45). At this time, this proportional value is used as the similarity between sample A and the test feature vector. Similarly, the similarity between sample B and the test feature vector can also be obtained.By analogy calculation, calculate the similarity between each sample in the second training set and the test feature vector, and summarize the similarities to obtain the similarity matrix.

[0072] 207. Determine the first detection result of the first KPI time series according to the similarity matrix, the positive class samples in the second training set, and the negative class samples in the second training set.

[0073] In this embodiment, after the detection device calculates the similarity matrix, the detection device can calculate the first similarity between the positive class samples in the second training set and the test feature vector and the second similarity between the negative class samples in the second training set and the test feature vector according to the similarity matrix respectively; determine the first classification result of the test feature vector according to the first similarity and the second similarity; determine the first detection result of the first KPI time series according to the first classification result.

[0074] That is, after obtaining the similarity matrix, the specific method for the detection device to calculate the first similarity between the positive class samples in the second training set and the test feature vector and the second similarity between the negative class samples in the second training set and the test feature vector according to the similarity matrix can be as follows: determine the similarity set corresponding to the positive class samples in the second training set from the similarity matrix, and sum up each similarity in the similarity set corresponding to the positive class samples to obtain the first similarity; determine the similarity set corresponding to the negative class samples in the second training set from the similarity matrix, and sum up each similarity in the similarity set corresponding to the negative class samples to obtain the second similarity. In an exemplary solution, assume that there are 50 positive class samples and 50 negative class samples in the second training set, and the test feature vector corresponds to a KPI time series, then the similarity matrix is a 100*1 matrix, that is, it includes 100 similarity values. Then the first similarity between the positive class samples and the test feature vector is the sum of 50 similarities of the positive class samples, and the second similarity between the negative class samples and the test feature vector is the sum of 50 similarities of the negative class samples.

[0075] It can be understood that when calculating the first similarity and the second similarity, the detection device can also use other methods for calculation, such as taking the average value or variance of the similarities between the positive class samples and the test feature vector, etc. The specific method is not limited here.

[0076] After the detection device obtains the first similarity and the second similarity, it determines the classification result of the test feature vector according to the first similarity and the second similarity. Specifically, it can be as follows: when the first similarity is greater than the second similarity, it is determined that the first classification result of the test feature vector is normal; when the first similarity is less than the second similarity, it is determined that the first classification result of the test feature vector is abnormal. In an exemplary solution, assume that the first similarity between the positive class samples in the second training set and the test feature vector is 20, and the second similarity between the negative class samples in the second training set and the test feature vector is 21. Then the first classification result of the test feature vector is abnormal; assume that the first similarity between the positive class samples in the second training set and the test feature vector is 20, and the second similarity between the negative class samples in the second training set and the test feature vector is 19. Then the first classification result of the test feature vector is normal.

[0077] In this embodiment, in order to improve the accuracy of the detection result of the first KPI time series, the detection device may perform repeated calculations on steps 203 to 206 to obtain multiple detection results, and finally determine the final detection result of the first KPI time series according to the detection results. The specific operation can be as follows: obtain a first value and a second value according to the detection results in the detection result set, where the first value is the number of detection results indicating that the first KPI time series is a normal KPI time series, and the second value is the number of detection results indicating that the first KPI time series is an abnormal KPI time series; when the first value is greater than the second value, it is determined that the first KPI time series is a normal time series; when the first value is less than the second value, it is determined that the first KPI time series is an abnormal time series. In an exemplary solution, assume that the detection device repeats the operations of steps 203 to 206 100 times. Then the detection device can obtain 100 detection results for the first KPI time series. The 100 detection results include detection results indicating that the first KPI time series is a normal time series and also include detection results indicating that the first KPI time series is an abnormal time series; assume that the number of detection results indicating that the first KPI time series is a normal time series is 65, and the number of detection results indicating that the first KPI time series is an abnormal time series is 35. Then it is determined that the first KPI time series is a normal time series. If the number of detection results indicating that the first KPI time series is a normal time series is 45, and the number of detection results indicating that the first KPI time series is an abnormal time series is 55. Then it is determined that the first KPI time series is an abnormal time series.

[0078] In this embodiment, in the normal sequence samples and abnormal sequence samples, sampling is performed to generate a balanced data set, and then the sampled balanced data set is used as the training data for the random forest algorithm. This avoids directly using the original unbalanced data set as the training set, which may exacerbate the imbalance of the data set during the sampling process of the random forest algorithm. Therefore, it is also beneficial to improve the diversity of the data, enrich the differences between different models, and improve the generalization performance of the KPI time series anomaly detection algorithm. At the same time, statistical features, fitting features, original features, etc. of the KPI time series are extracted. These features can well reflect the dispersion degree, change trend, forward and backward correlation, and implicit characteristics of the KPI time series, providing effective data features for the anomaly detection of the KPI time series, thereby increasing the accuracy of the KPI time series detection. Further, the random forest similarity matrix is used to measure the sample similarity as the initial model learner, and combined with the ensemble learning idea, the classification results of multiple similarity matrices are aggregated to output the final result, effectively improving the accuracy of the classification result.

[0079] The beneficial effects of the method provided by this application will be described below through a specific experimental process:

[0080] The five algorithms in the experiment are the month-on-month algorithm and year-on-year algorithm based on fixed configurations, the time series prediction algorithm Prophet developed by Facebook, the AnomalyDetection algorithm developed by Twitter, the isolation forest algorithm, and the Local Outlie Factor (LOF) algorithm. Among them, the month-on-month algorithm, year-on-year algorithm, and Prophet algorithm need to set threshold parameters in advance for KPI sequence anomaly detection. The threshold of the month-on-month algorithm is set to 10%, that is, if the difference ratio between the T moment and the T-1 moment exceeds 10%, it is regarded as an outlier; the threshold of the year-on-year algorithm is set to 10%, that is, if the difference ratio between the T moment and the T-t moment (t represents the time period of 1 day) exceeds 10%, it is regarded as an outlier; the threshold based on the Prophet algorithm is set to 10%, that is, when the difference ratio between the Prophet predicted value and the actual value of the KPI sequence exceeds 10%, it is regarded as an outlier.

[0081] Select a real Internet KPI dataset from the website. This competition dataset collects KPI data from real scenarios of many Internet companies, and professional personnel judge and label abnormal data points, which are provided after being desensitized. The interval between every two time points is 1 minute or 5 minutes, and three KPI sequences are selected as the evaluation dataset to evaluate the classification performance of the algorithm for identifying outliers. Referring to the literature on Microsoft and other KPI anomaly detection algorithms, three evaluation metrics, namely Accuracy, F1-measure, and F2-measure, are selected to evaluate the performance of the algorithm. The test results are shown in Table 1, Table 2, and Table 3 as follows:

[0082] Table 1

[0083] Method Accuracy <![CDATA[F1-measure]]> <![CDATA[F2-measure]]> MoM Algorithm 0.07 0.08 0.18 YoY Algorithm 0.05 0.08 0.18 Anomaly Detection Algorithm 0.98 0.67 0.58 LOF Algorithm 0.77 0.26 0.46 Isolation Forest Algorithm 0.97 0.72 0.79 Prophet Threshold Algorithm 0.88 0.4 0.61 Algorithm Provided by This Application 0.99 0.92 0.88

[0084] Table 2

[0085] Method Accuracy <![CDATA[F1-measure]]> <![CDATA[F2-measure]]> MoM Algorithm 0.24 0.39 0.61 YoY Algorithm 0.24 0.39 0.61 Anomaly Detection Algorithm 0.80 0.43 0.35 LOF Algorithm 0.73 0.43 0.43 Isolation Forest Algorithm 0.75 0.36 0.26 Prophet Threshold Algorithm 0.7 0.45 0.49 Algorithm Provided by This Application 0.73 0.58 0.68

[0086] Table 3

[0087] Method Accuracy <![CDATA[F1-measure]]> <![CDATA[F2-measure]]> MoM Algorithm 0.14 0.25 0.46 YoY Algorithm 0.14 0.25 0.46 Anomaly Detection Algorithm 0.86 0.06 0.04 LOF Algorithm 0.66 0.39 0.54 Isolation Forest Algorithm 0.76 0.38 0.44 Prophet Threshold Algorithm 0.56 0.29 0.42 Algorithm Provided by This Application 0.91 0.60 0.51

[0088] It can be understood that in the above Table 1, Table 2, and Table 3, the bold numbers are used to indicate that the algorithm performs best in this indicated comparison. Based on the results of the above Table 1, Table 2, and Table 3, it can be seen that the algorithm provided in this application has better detection and recognition effects compared to the other five commonly used algorithms.

[0089] The detection device in this application will be described in detail below. Please refer to Figure 3 , Figure 3 which is a schematic diagram of an embodiment of the detection device in an embodiment of this application. The detection device 20 includes:

[0090] A first acquisition module 201, configured to acquire the first key performance indicator (KPI) time series at the time point to be detected and the set of KPI time series within a preset time period before the time point to be detected;

[0091] A feature extraction module 202, configured to extract the data features of the first KPI time series to generate a test feature vector, and extract the data features of each KPI time series in the set of KPI time series to generate a first training set;

[0092] A second acquisition module 203, configured to acquire the positive class sample set and the negative class sample set in the first training set, where the positive class sample set includes KPI time series samples at normal times, and the negative class sample set includes KPI time series samples at abnormal times;

[0093] A sampling module 204, configured to sample the positive class sample set and the negative class sample set to obtain a second training set, where the number of positive class samples in the second training set is the same as the number of negative class samples;

[0094] A training module 205, configured to train a random forest algorithm using the second training set to obtain an anomaly detection model;

[0095] A detection module 206, configured to input the second training set and the test feature vector into the anomaly detection model to calculate a similarity matrix corresponding to the second training set and the test feature vector; and determine a first detection result of the first KPI time series according to the similarity matrix, the positive class samples in the second training set, and the negative class samples in the second training set.

[0096] In an embodiment of the present application, a detection device is provided. By using the above device, in normal sequence samples and abnormal sequence samples, sampling is performed to generate a balanced data set, and then the sampled balanced data set is used as training data for the random forest algorithm. This avoids directly using the original unbalanced data set as the training set, which may exacerbate the imbalance of the data set during the sampling process of the random forest algorithm. Therefore, it is also beneficial to improve the diversity of the data, enrich the differences between different models, and improve the generalization performance of the KPI time series anomaly detection algorithm.

[0097] Optionally, based on the corresponding embodiment above, in another embodiment of the detection device 20 provided in the embodiment of the present application, Figure 3 The feature extraction module 202 is specifically configured to extract statistical features, fitting features, and original features of the first KPI time series to generate the test feature vector;

[0098] The feature extraction module 202 is specifically configured to extract statistical features, fitting features, and original features of each KPI time series in the KPI time series set to generate the first training set.

[0099] In an embodiment of the present application, a detection device is provided. By using the above device, feature extraction such as statistical features, fitting features, and original features is performed on the KPI time series. These features can well reflect the dispersion degree, change trend, forward and backward correlation, and implicit characteristics of the KPI time series, etc., providing effective data features for the anomaly detection of the KPI time series, thereby increasing the accuracy of the KPI time series detection.

[0100] Optionally, based on the above,

[0101] Optionally, in the above Figure 3Based on the corresponding embodiment, in another embodiment of the detection device 20 provided in the embodiments of the present application, the sampling module 204 is further configured to sample the positive sample set and the negative sample set to obtain a third training set, where the number of positive samples in the third training set is the same as the number of negative samples;

[0102] The training module 205 is further configured to use the third training set to train a random forest algorithm to update the anomaly detection model;

[0103] The detection module 206 is further configured to input the third training set and the test feature vector into the updated anomaly detection model to output a second detection result of the first KPI time series, and the first detection result and the second detection result are used as a detection result set; and so on, until the number of detection results in the detection result set reaches a preset number, then determine the final detection result of the first KPI time series according to the detection results in the detection result set.

[0104] In the embodiments of the present application, a detection device is provided. By using the above device, independent repeated sampling is performed on normal sequence samples and abnormal sequence samples to generate a balanced data set, and then the sampled balanced data set is used as the training data of the random forest algorithm. This avoids directly using the original unbalanced data set as the training set, which may exacerbate the imbalance of the data set during the sampling process of the random forest algorithm. Therefore, it is also beneficial to improve the diversity of the data, enrich the differences between different models, and improve the generalization performance of the KPI time series anomaly detection algorithm.

[0105] Optionally, based on the corresponding embodiment above, in another embodiment of the detection device 20 provided in the embodiments of the present application, Figure 3 The detection module 206 is specifically configured to calculate a first similarity of positive samples in the second training set and a second similarity of negative samples in the second training set according to the similarity matrix respectively;

[0106] Determine a first classification result of the test feature vector according to the first similarity and the second similarity;

[0107] Determine a first detection result of the first KPI time series according to the first classification result.

[0108] In the embodiments of the present application, a detection device is provided. By using the above device, the random forest similarity matrix is used to measure the sample similarity as the initial model learner, and combined with the idea of ensemble learning, the classification results of multiple similarity matrices are summarized and output as the final result, effectively improving the accuracy of the classification result.

[0109]

[0110] ​Optionally, based on the above Figure 3 corresponding embodiment, in another embodiment of the detection device 20 provided by the embodiments of the present application,

[0111] The detection module 206 is specifically configured to input the second training set and the test feature vector into the anomaly detection model to calculate a first ratio value of the first sample in the second training set and the first sample corresponding to the test feature vector falling in the same leaf node in the anomaly detection model, and use the first ratio value as the similarity between the first sample in the second training set and the first sample corresponding to the test feature vector;

[0112] Input the second training set and the test feature vector into the anomaly detection model to calculate a second ratio value of the second sample in the second training set and the second sample corresponding to the test feature vector falling in the same leaf node in the anomaly detection model, and use the second ratio value as the similarity between the second sample in the second training set and the second sample corresponding to the test feature vector;

[0113] And so on, until the similarities between each sample in the second training set and each sample corresponding to the test feature vector are calculated by traversal, and the similarities between each sample in the second training set and each sample corresponding to the test feature vector are summarized to obtain the similarity matrix.

[0114] In the embodiments of the present application, a detection device is provided. By using the above device and combining the idea of ensemble learning, the similarities between multiple training samples and test samples are integrated to obtain the similarity matrix, thereby effectively improving the accuracy of the classification result.

[0115] Optionally, based on the above Figure 3 corresponding embodiment, in another embodiment of the detection device 20 provided by the embodiments of the present application, the detection module 206 is specifically configured to determine a similarity set corresponding to the positive class samples in the second training set from the similarity matrix, and sum the similarities in the similarity set corresponding to the positive class samples to obtain the first similarity;

[0116] Determine a similarity set corresponding to the negative class samples in the second training set from the similarity matrix, and sum the similarities in the similarity set corresponding to the negative class samples to obtain the second similarity.

[0117] In the embodiments of the present application, a detection device is provided. By using the above device, it is also possible to determine the type of each positioning point according to the positioning speed information of each positioning point. Therefore, for the case where the type of the positioning point cannot be directly collected, the type of each positioning point can still be determined through relevant calculations, thereby improving the feasibility and operability of the solution.

[0118] Optionally, based on the above Figure 3 corresponding embodiment, in another embodiment of the detection device 20 provided by the embodiments of the present application, the detection module 206 is specifically configured to determine that the first classification result of the test feature vector is normal when the first similarity is greater than the second similarity;

[0119] When the first similarity is less than the second similarity, it is determined that the first classification result of the test feature vector is abnormal.

[0120] In the embodiments of the present application, a detection device is provided. By using the above device, the classification result is judged according to the similarity, thereby effectively improving the accuracy of the classification result.

[0121] Optionally, based on the above Figure 3 corresponding embodiment, in another embodiment of the detection device 20 provided by the embodiments of the present application,

[0122] the detection module 206 is specifically configured to obtain a first value and a second value according to the detection results in the detection result set, where the first value is the number of detection results indicating that the first KPI time series is a normal KPI time series, and the second value is the number of detection results indicating that the first KPI time series is an abnormal KPI time series;

[0123] When the first value is greater than the second value, it is determined that the first KPI time series is a normal time series;

[0124] When the first value is less than the second value, it is determined that the first KPI time series is an abnormal time series.

[0125] In the embodiments of the present application, a detection device is provided. By using the above device and combining the idea of ensemble learning, the similarities between multiple training samples and test samples are integrated to obtain the similarity matrix, thereby effectively improving the accuracy of the classification result.

[0126] The detection device provided by the present application can be used in a server. Please refer to Figure 4 , Figure 4FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The programs stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0127] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0128] The steps performed by the detection device in the above embodiments may be based on the Figure 4 server structure shown.

[0129] The detection device provided by the present application can be used for terminal devices. Please refer to Figure 5 . For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. In the embodiments of the present application, a smart phone is taken as an example of the terminal device for illustration:

[0130] Figure 5 FIG. shows a block diagram of a part of the structure of a smart phone related to the terminal device provided by the embodiment of the present application. Referring to Figure 5 , the smart phone includes: a radio frequency (RF) circuit 410, a memory 420, an input unit 430, a display unit 440, a sensor 450, an audio circuit 460, a wireless fidelity (WiFi) module 470, a processor 480, and a power supply 490 and other components. Those skilled in the art can understand that Figure 5The smartphone structure shown does not limit the smartphone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0131] The following will specifically introduce each component of the smartphone in conjunction with Figure 5 :

[0132] The RF circuit 410 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is given to the processor 480 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 410 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 410 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0133] The memory 420 can be used to store software programs and modules. The processor 480 executes various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 420. The memory 420 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the smartphone (such as audio data, phone book, etc.). In addition, the memory 420 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0134] The input unit 430 can be used to receive input numeric or character information and generate key signal inputs related to the user settings and function controls of the smart phone. Specifically, the input unit 430 may include a touch panel 431 and other input devices 432. The touch panel 431, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 431), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 431 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, then sends it to the processor 480, and can receive and execute the commands sent by the processor 480. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 431. In addition to the touch panel 431, the input unit 430 may further include other input devices 432. Specifically, the other input devices 432 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0135] The display unit 440 can be used to display the information input by the user or the information provided to the user and various menus of the smart phone. The display unit 440 may include a display panel 441. Optionally, the display panel 441 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 431 can cover the display panel 441. When the touch panel 431 detects a touch operation thereon or nearby, it is transmitted to the processor 480 to determine the type of touch event. Subsequently, the processor 480 provides corresponding visual output on the display panel 441 according to the type of touch event. Although in Figure 5 the touch panel 431 and the display panel 441 are implemented as two independent components to realize the input and input functions of the smart phone, in some embodiments, the touch panel 431 and the display panel 441 can be integrated to realize the input and output functions of the smart phone.

[0136] The smart phone may further include at least one sensor 450, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 441 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 441 and / or the backlight when the smart phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the smart phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the smart phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0137] The audio circuit 460, the speaker 461, and the microphone 462 can provide an audio interface between the user and the smart phone. The audio circuit 460 can transmit the electrical signal converted from the received audio data to the speaker 461, and the speaker 461 converts it into a sound signal for output; on the other hand, the microphone 462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 460 and then converted into audio data. After the audio data is output to the processor 480 for processing, it is sent through the RF circuit 410 to, for example, another smart phone, or the audio data is output to the memory 420 for further processing.

[0138] WiFi belongs to short - range wireless transmission technology. The smart phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 470, which provides users with wireless broadband Internet access. Although Figure 5 the WiFi module 470 is shown, it can be understood that it does not belong to the essential components of the smart phone and can be completely omitted within the scope of not changing the essence of the invention according to needs.

[0139] The processor 480 is the control center of the smart phone, connecting various parts of the entire smart phone using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 420, and by calling data stored in the memory 420, it executes various functions of the smart phone and processes data, thereby monitoring the smart phone as a whole. Optionally, the processor 480 may include one or more processing units; optionally, the processor 480 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 480 either.

[0140] The smart phone further includes a power supply 490 (such as a battery) for powering each component. Optionally, the power supply can be logically connected to the processor 480 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system.

[0141] Although not shown, the smart phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.

[0142] In the above embodiments, the steps performed by the detection device may be based on the Figure 5 shown terminal device structure.

[0143] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0144] An embodiment of the present application further provides a computer program product including a program. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0145] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.

[0146] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in electrical, mechanical or other forms.

[0147] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0148] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0149] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0150] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A KPI time series detection method, characterized in that, Including: Obtain the first KPI time series at the time point to be detected and the set of KPI time series within a preset time period before the time point to be detected; Extract the data features of the first KPI time series to generate a test feature vector, and extract the data features of each KPI time series in the set of KPI time series to generate a first training set; Obtain the positive sample set and the negative sample set in the first training set, where the positive sample set includes KPI time series samples at normal times, and the negative sample set includes KPI time series samples at abnormal times; Sample the positive sample set and the negative sample set to obtain a second training set, where the number of positive samples in the second training set is the same as the number of negative samples; Use the second training set to train the random forest algorithm to obtain an anomaly detection model; Input the second training set and the test feature vector into the anomaly detection model to calculate the similarity matrix corresponding to the second training set and the test feature vector; Determine the first detection result of the first KPI time series according to the similarity matrix, the positive samples in the second training set, and the negative samples in the second training set.

2. The method according to claim 1, characterized in that, The extracting the data features of the first KPI time series to generate a test feature vector includes: Extract the statistical features, fitting features, and original features of the first KPI time series to generate the test feature vector; The extracting the data features of each KPI time series in the set of KPI time series to generate a first training set includes: Extract the statistical features, fitting features, and original features of each KPI time series in the set of KPI time series to generate the first training set.

3. The method according to claim 1, wherein After inputting the second training set and the test feature vector into the anomaly detection model to output the first detection result of the first KPI time series, the method further includes: Sample the positive sample set and the negative sample set to obtain a third training set, where the number of positive samples in the third training set is the same as the number of negative samples; Use the third training set to train the random forest algorithm to update the anomaly detection model; Input the third training set and the test feature vector into the updated anomaly detection model to output the second detection result of the first KPI time series, and use the first detection result and the second detection result as a detection result set; And so on, until the number of detection results in the detection result set reaches a preset number, then determine the final detection result of the first KPI time series according to the detection results in the detection result set.

4. The method according to any one of claims 1 to 3, characterized in that, The determining the first detection result of the first KPI time series according to the similarity matrix, the positive samples in the second training set, and the negative samples in the second training set includes: Calculate the first similarity between the positive samples in the second training set and the test feature vector and the second similarity between the negative samples in the second training set and the test feature vector according to the similarity matrix; Determine the first classification result of the test feature vector according to the first similarity and the second similarity; Determine the first detection result of the first KPI time series according to the first classification result.

5. The method according to any one of claims 1 to 3, characterized in that, The step of inputting the second training set and the test feature vector into the anomaly detection model to calculate the similarity matrix between the second training set and the test feature vector includes: Input the second training set and the test feature vector into the anomaly detection model to calculate the first proportion value that the first sample in the second training set and the corresponding first sample of the test feature vector fall into the same leaf node in the anomaly detection model, and use the first proportion value as the similarity between the first sample in the second training set and the corresponding first sample of the test feature vector; Input the second training set and the test feature vector into the anomaly detection model to calculate the second proportion value that the second sample in the second training set and the corresponding second sample of the test feature vector fall into the same leaf node in the anomaly detection model, and use the second proportion value as the similarity between the second sample in the second training set and the corresponding second sample of the test feature vector; And so on, until the similarities between each sample in the second training set and the corresponding samples of the test feature vector are calculated by traversal, and summarize the similarities between each sample in the second training set and the corresponding samples of the test feature vector to obtain the similarity matrix.

6. The method according to claim 4, wherein The step of calculating the first similarity between the positive class samples in the second training set and the test feature vector and the second similarity between the negative class samples in the second training set and the test feature vector according to the similarity matrix includes: Determine the set of similarities corresponding to the positive class samples in the second training set from the similarity matrix, and sum up each similarity in the set of similarities corresponding to the positive class samples to obtain the first similarity; Determine the set of similarities corresponding to the negative class samples in the second training set from the similarity matrix, and sum up each similarity in the set of similarities corresponding to the negative class samples to obtain the second similarity.

7. The method according to claim 4, characterized in that, The step of determining the first classification result of the test feature vector according to the first similarity and the second similarity includes: When the first similarity is greater than the second similarity, determine that the first classification result of the test feature vector is normal; When the first similarity is less than the second similarity, determine that the first classification result of the test feature vector is abnormal.

8. The method according to claim 3, wherein The step of determining the final detection result of the first KPI time series according to the detection results in the detection result set includes: Obtain a first value and a second value according to the detection results in the detection result set, where the first value is the number of detection results indicating that the first KPI time series is a normal KPI time series, and the second value is the number of detection results indicating that the first KPI time series is an abnormal KPI time series; When the first value is greater than the second value, determine that the first KPI time series is a normal time series; When the first value is less than the second value, determine that the first KPI time series is an abnormal time series.

9. A detection device, characterized in that, including: A first acquisition module, configured to acquire a first KPI time series at a time point to be detected and a set of KPI time series within a preset time period before the time point to be detected; A feature extraction module, configured to extract data features of the first KPI time series to generate a test feature vector, and extract data features of each KPI time series in the set of KPI time series to generate a first training set; A second acquisition module, configured to acquire a positive sample set and a negative sample set in the first training set, where the positive sample set includes KPI time series samples at normal times, and the negative sample set includes KPI time series samples at abnormal times; A sampling module, configured to sample the positive sample set and the negative sample set to obtain a second training set, where the number of positive samples in the second training set is the same as the number of negative samples; A training module, configured to train a random forest algorithm using the second training set to obtain an anomaly detection model; A detection module, configured to input the second training set and the test feature vector into the anomaly detection model to calculate a similarity matrix corresponding to the second training set and the test feature vector; determine a first detection result of the first KPI time series according to the similarity matrix, positive samples in the second training set, and negative samples in the second training set.

10. The device according to claim 9, characterized in that, The feature extraction module is specifically configured to extract statistical features, fitting features, and original features of the first KPI time series to generate the test feature vector; The feature extraction module is specifically configured to extract statistical features, fitting features, and original features of each KPI time series in the set of KPI time series to generate the first training set.

11. The device according to claim 9, characterized in that, The sampling module is further configured to sample the positive sample set and the negative sample set to obtain a third training set, where the number of positive samples in the third training set is the same as the number of negative samples; The training module is further configured to train a random forest algorithm using the third training set to update the anomaly detection model; The detection module is further configured to input the third training set and the test feature vector into the updated anomaly detection model to output a second detection result of the first KPI time series, and use the first detection result and the second detection result as a detection result set; And so on, until the number of detection results in the detection result set reaches a preset number, then determine a final detection result of the first KPI time series according to the detection results in the detection result set.

12. The device according to any one of claims 9 to 11, characterized in that The detection module is specifically configured to calculate a first similarity of positive samples in the second training set and a second similarity of negative samples in the second training set according to the similarity matrix; Determine a first classification result of the test feature vector according to the first similarity and the second similarity; Determine a first detection result of the first KPI time series according to the first classification result.

13. The device according to any one of claims 9 to 11, characterized in that The detection module is specifically configured to input the second training set and the test feature vector into the anomaly detection model to calculate a first proportion value that the first sample in the second training set and the first sample corresponding to the test feature vector fall into the same leaf node in the anomaly detection model, and use the first proportion value as the similarity between the first sample in the second training set and the first sample corresponding to the test feature vector; Input the second training set and the test feature vector into the anomaly detection model to calculate a second proportion value that the second sample in the second training set and the second sample corresponding to the test feature vector fall into the same leaf node in the anomaly detection model, and use the second proportion value as the similarity between the second sample in the second training set and the second sample corresponding to the test feature vector; And so on, until the similarities between each sample in the second training set and each sample corresponding to the test feature vector are calculated by traversal, and summarize the similarities between each sample in the second training set and each sample corresponding to the test feature vector to obtain the similarity matrix.

14. The device according to claim 12, wherein The detection module is specifically configured to determine a similarity set corresponding to the positive class samples in the second training set from the similarity matrix, and sum up each similarity in the similarity set corresponding to the positive class samples to obtain the first similarity; Determine a similarity set corresponding to the negative class samples in the second training set from the similarity matrix, and sum up each similarity in the similarity set corresponding to the negative class samples to obtain the second similarity.

15. The device according to claim 12, characterized in that The detection module is specifically configured to, when the first similarity is greater than the second similarity, determine that the first classification result of the test feature vector is normal; When the first similarity is less than the second similarity, determine that the first classification result of the test feature vector is abnormal.

16. The device according to claim 11, wherein The detection module is specifically configured to obtain a first value and a second value according to the detection results in the detection result set, where the first value is the number of detection results indicating that the first KPI time series is a normal KPI time series, and the second value is the number of detection results indicating that the first KPI time series is an abnormal KPI time series; When the first value is greater than the second value, determine that the first KPI time series is a normal time series; When the first value is less than the second value, determine that the first KPI time series is an abnormal time series.

17. A computer device, characterized in that, Comprising: A memory, a processor, and a bus system; Wherein, the memory is used for storing programs; The processor is used for executing the programs in the memory, and the processor is used for executing the method according to any one of claims 1 to 8 according to the instructions in the program code; The bus system is used for connecting the memory and the processor, so that the memory and the processor can communicate with each other.

18. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 8.

19. A computer program product, characterized in that, The computer program product includes instructions that, when executed on a computer device, cause the computer device to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Adaptive sampling for imbalance mitigation and dataset size reduction in machine learning

    US20200342265A1

  • Forcasting time series data

    US20210099894A1