Anomaly determination method, training method, device, electronic device, and storage medium

CN114328123BActive Publication Date: 2026-09-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111680366.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2026-09-11
Estimated Expiration
2041-12-30

Smart Images

  • Figure CN114328123B_ABST
    Figure CN114328123B_ABST
Patent Text Reader

Abstract

The present disclosure provides an anomaly determination method, a training method and device of an anomaly determination model, an electronic device, a storage medium and a program product, relates to the technical field of data processing, and in particular to the technical field of deep learning, big data and the like. The specific implementation scheme is as follows: features of target data to be processed are extracted to obtain a target feature vector, the target data to be processed including index data at a current moment and index data of a plurality of historical time series associated with the current moment; and based on the target feature vector, a label of the index data at the current moment is obtained, the label being used to represent whether the index data at the current moment is abnormal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of deep learning and big data. Specifically, it relates to anomaly detection methods, training methods for anomaly detection models, devices, electronic devices, storage media, and program products. Background Technology

[0002] With the rapid development of internet services, users are becoming increasingly reliant on the business systems that support these services, thus demanding improved system stability. Providing early warnings for abnormal situations in business systems is one way to enhance their stability. Summary of the Invention

[0003] This disclosure provides an anomaly determination method, an anomaly determination model training method, an apparatus, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of this disclosure, an anomaly determination method is provided, comprising: extracting features from target data to be processed to obtain a target feature vector, wherein the target data to be processed includes indicator data at the current moment and indicator data of multiple historical time series associated with the current moment; and obtaining a label for the indicator data at the current moment based on the target feature vector, wherein the label is used to characterize whether the indicator data at the current moment is abnormal.

[0005] According to another aspect of this disclosure, a training method for an anomaly determination model is provided, comprising: training the anomaly determination model using target training samples to obtain the trained anomaly determination model, wherein the target training samples include sample data and sample labels, wherein the sample data includes sample indicator data at the current time and sample indicator data of multiple historical time series associated with the current time, and the sample labels include the sample labels at the current time, the sample labels being used to characterize whether the sample indicator data at the current time is anomaly.

[0006] According to another aspect of this disclosure, an anomaly determination apparatus is provided, comprising: an extraction module, configured to extract features of target data to be processed to obtain a target feature vector, wherein the target data to be processed includes indicator data at the current moment and indicator data of multiple historical time series associated with the current moment; and a determination module, configured to obtain a label of the indicator data at the current moment based on the target feature vector, wherein the label is used to characterize whether the indicator data at the current moment is abnormal.

[0007] According to another aspect of this disclosure, a training apparatus for an anomaly determination model is provided, comprising: a training module for training the anomaly determination model using target training samples to obtain a trained anomaly determination model, wherein the target training samples include sample data and sample labels, wherein the sample data includes sample indicator data at the current time and sample indicator data of multiple historical time series associated with the current time, and the sample labels include the sample labels at the current time, the sample labels being used to characterize whether the sample indicator data at the current time is anomaly.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein the memory stores instructions executable by said at least one processor, said instructions being executed by said at least one processor to enable said at least one processor to perform a method as disclosed herein.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods as disclosed herein.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as disclosed herein.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This illustration schematically shows an exemplary system architecture to which anomaly determination methods and apparatus can be applied according to embodiments of the present disclosure;

[0014] Figure 2 A flowchart illustrating an anomaly determination method according to an embodiment of the present disclosure is shown schematically;

[0015] Figure 3 A flowchart illustrating an anomaly determination method according to another embodiment of this disclosure is shown schematically;

[0016] Figure 4 A flowchart illustrating the acquisition of reference data according to another embodiment of this disclosure is shown schematically;

[0017] Figure 5A flowchart illustrating a training method for an anomaly determination model according to another embodiment of the present disclosure is shown schematically;

[0018] Figure 6 A block diagram of an anomaly determination apparatus according to an embodiment of the present disclosure is shown schematically;

[0019] Figure 7 A block diagram schematically illustrates a training apparatus for an anomaly determination model according to an embodiment of the present disclosure; and

[0020] Figure 8 A block diagram of an electronic device suitable for implementing an anomaly determination method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0022] This disclosure provides an anomaly determination method, an anomaly determination model training method, an apparatus, an electronic device, a storage medium, and a program product.

[0023] According to embodiments of this disclosure, an anomaly determination method is provided, comprising: extracting features of target data to be processed to obtain a target feature vector, wherein the target data to be processed includes indicator data at the current moment and indicator data of multiple historical time series associated with the current moment; and obtaining a label of the indicator data at the current moment based on the target feature vector, wherein the label is used to characterize whether the indicator data at the current moment is abnormal.

[0024] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0025] Figure 1 The illustration schematically depicts an exemplary system architecture to which anomaly determination methods and apparatus can be applied according to embodiments of the present disclosure.

[0026] It is important to note that Figure 1The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the anomaly determination method and apparatus can be applied may include a terminal device, but the terminal device can implement the anomaly determination method and apparatus provided by the embodiments of this disclosure without interacting with the server.

[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0029] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0030] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0031] It should be noted that the anomaly determination method provided in this embodiment can generally be executed by terminal devices 101, 102, or 103. Accordingly, the anomaly determination device provided in this embodiment can also be disposed in terminal devices 101, 102, or 103.

[0032] Alternatively, the anomaly determination method provided in this embodiment can generally be executed by server 105. Correspondingly, the anomaly determination device provided in this embodiment can generally be located in server 105. The anomaly determination method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the anomaly determination device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0033] For example, when a user logs into an application on terminal devices 101, 102, or 103 and performs operations such as browsing or searching, server 105 monitors various metrics of the application and obtains historical time-series metric data. Based on this metric data, it issues warnings about whether the application might experience anomalies at the current moment. Alternatively, a server or server cluster capable of communicating with terminal devices 101, 102, 103, and / or server 105 can analyze the historical time-series metric data and ultimately issue warnings about whether the application might experience anomalies at the current moment.

[0034] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0035] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0036] Figure 2 A flowchart illustrating an anomaly determination method according to an embodiment of the present disclosure is shown schematically.

[0037] like Figure 2 As shown, the method includes operations S210 to S220.

[0038] In operation S210, the features of the target data to be processed are extracted to obtain the target feature vector. The target data to be processed includes the indicator data at the current moment and the indicator data of multiple historical time series associated with the current moment.

[0039] In operation S220, based on the target feature vector, the label of the indicator data at the current time is obtained, where the label is used to characterize whether the indicator data at the current time is abnormal.

[0040] According to embodiments of this disclosure, the metric data can refer to application metrics, such as the application's crash rate, but it is not limited to this; the number of times the application freezes can also be used as metric data. By predicting whether the metric data is abnormal, it is possible to predict whether the application's operation is abnormal, thereby achieving early warning.

[0041] According to embodiments of this disclosure, the indicator data at the current moment and the indicator data of multiple historical time series associated with the current moment can be used as target data to be processed, and the abnormality of the indicator data at the current moment can be determined based on the target data to be processed.

[0042] According to embodiments of this disclosure, the multiple historical time series associated with the current moment can be multiple historical time series divided at equal intervals, and each historical time series can have the same time length. However, it is not limited to this. The multiple historical time series associated with the current moment can also be multiple historical time series divided from historical time according to different time periods.

[0043] According to embodiments of this disclosure, multiple historical time series associated with the current moment may include historical time series within the same time period as the current moment, and at least one historical time series within a different time period. These multiple historical time series may have the same duration and all fall within the same historical time segment. For example, multiple historical time series may include a first historical time series of the first three hours before the current moment, a second historical time series of the first three hours of the same moment on the previous day, a third historical time series of the last three hours of the same moment on the previous day, a fourth historical time series of the first three hours of the same moment on the previous week, and a fifth historical time series of the last three hours of the same moment on the previous week. By utilizing the indicator data from historical time series within the same time period as the current moment, the characteristics of the indicator data most associated with the current moment can be learned. Combining the characteristics of the indicator data from at least one historical time series within a different time period as the current moment, more periodic indicator data can also be considered, such as indicators with large fluctuations and personalized fluctuation cycles, such as those related to holidays and release cycles. This allows for application in the field of anomaly identification of periodically continuous indicator data, improving the effectiveness of early warnings for abnormal moments.

[0044] According to embodiments of this disclosure, a time series feature extraction module, such as the Tsfresh module, can be used to extract features. This module can be designed to extract feature sets of various types, including statistical features, fitting features, and classification features. The operations provided in this disclosure for extracting features from the target data to be processed can be used to establish a feature set that adapts to abnormal scenario patterns, thereby increasing the diversity and accuracy of features. This reduces the complexity of determining whether the indicator data at the current moment is abnormal and increases the interpretability of the obtained labels. This allows for adaptation to various abnormal scenarios, such as mean shifts, upward trends, sharp drops, and changes in fluctuation frequency.

[0045] According to an embodiment of this disclosure, for operation S220, an anomaly determination model can be used to obtain the label of the indicator data at the current moment based on the target feature vector.

[0046] According to embodiments of this disclosure, the anomaly determination model may include at least one of the following: logistic regression model, decision tree model, etc.

[0047] According to embodiments of this disclosure, a decision tree model, such as XGBoost (eXtreme GradientBoosting), can be utilized.

[0048] According to embodiments of this disclosure, before performing operation S210 to extract features of the target data to be processed and obtain the target feature vector, a regression method can be used to filter the data to be processed to obtain the target data to be processed.

[0049] The anomaly determination method provided in this embodiment can integrate regression methods with anomaly determination models. In addition to using the anomaly determination model to determine whether the indicator data at the current moment is abnormal, a filtering operation is added to improve the accuracy of determining abnormal indicator data.

[0050] According to embodiments of this disclosure, multiple regression algorithms in a regression method can be used to filter the data to be processed to obtain the target data to be processed.

[0051] For example, data to be processed is extracted from the initial target data according to a predetermined time window length. The data to be processed is then input into multiple regression algorithms to obtain multiple regression results, where each regression result corresponds one-to-one with a different regression algorithm.

[0052] According to embodiments of this disclosure, the predetermined time window length can be a fixed time window length, such as 3 hours, but it is not limited to this; it can also be a time window length that varies periodically. Data to be processed can be extracted from the initial target data according to the predetermined time window length and the same time period within different time cycles.

[0053] For example, using a predetermined time window of 3 hours, the indicator data for the current moment is extracted from the initial target data to be processed. This includes indicator data from the first historical time series (3 hours before the current moment), the second historical time series (3 hours before the same moment on the previous day), the third historical time series (3 hours after the same moment on the previous day), the fourth historical time series (3 hours before the same moment on the previous week), and the fifth historical time series (3 hours after the same moment on the previous week). Each of the first to fifth historical time series has a 3-hour time frame, and indicator data for each of these 180 moments is determined at one-minute intervals. This also includes indicator data for the current moment, the moments on the previous day, and the moments on the previous week. The total number of indicator data moments determined is 903.

[0054] According to embodiments of this disclosure, in response to determining that all of the multiple regression results are normal regression results for the indicator data used to characterize the current moment, subsequent operations S210 and S220 can be stopped.

[0055] According to an embodiment of this disclosure, in response to determining that at least one of the multiple regression results is a target regression result used to characterize the indicator data at the current moment as abnormal, the data to be processed is taken as the target data to be processed, and subsequent operations S210 and S220 are performed.

[0056] According to embodiments of this disclosure, at least one of the plurality of regression algorithms may include at least one of the following: polynomial fitting algorithm, exponential moving average algorithm, isolated forest algorithm, and standard deviation method.

[0057] For example, the data to be processed is input into a multinomial fitting algorithm to obtain a first regression result. The data is then input into an exponential moving average algorithm to obtain a second regression result. The data is further input into an isolated forest algorithm to obtain a third regression result. Finally, the data is input into a 3Sigma algorithm using the standard deviation method to obtain a fourth regression result. If at least one of the first to fourth regression results is the target regression result, operations S210 and S220 are executed. If all four regression results represent normal regression results for the indicator data at the current moment, subsequent operations S210 and S220 are stopped.

[0058] The regression method provided in this disclosure can be used to filter the data to be processed. Multiple regression algorithms can be used to reduce the reliance on human experience. It can be applied to the analysis scenarios of various types of indicator data with different dimensions, personalized fluctuation cycles, and different fluctuation degrees, thereby improving the universality of determining whether the indicator data at the current moment is abnormal.

[0059] Figure 3 A flowchart illustrating an anomaly determination method according to another embodiment of this disclosure is shown schematically.

[0060] like Figure 3 As shown, initial data to be processed 310 can be obtained, which includes indicator data at the current moment and initial indicator data of multiple historical time series associated with the current moment. However, it is not limited to this. The initial data can also be indicator data at various moments of a historical continuous time series. Historical labels 320 for characterizing whether the initial indicator data is abnormal can also be obtained from the initial indicator data of multiple historical time series. The initial data to be processed 310 can be preprocessed based on the historical labels 320 to obtain preprocessed data 330. The target preprocessed data in the preprocessed data 330 is updated using the benchmark data to obtain target initial data to be processed 340. Data to be processed 350 is extracted from the target initial data to be processed 340 according to a predetermined time window length. The data to be processed 350 is filtered using a regression method to obtain target data to be processed 360. Features of the target data to be processed 360 are extracted to obtain target feature vector 370. Input the target feature vector 370 into the anomaly determination model to obtain the label of the indicator data at the current time, such as label 381 representing normality and label 382 representing anomaly.

[0061] According to embodiments of this disclosure, the target preprocessed data in the preprocessed data may refer to the initial index data of at least one historical time series that is in a different time period than the current time. For example, the initial index data of a second historical time series in the first three hours of the same time as the current time on the previous day, the initial index data of a third historical time series in the last three hours of the same time as the current time on the previous day, the initial index data of a fourth historical time series in the first three hours of the same time as the current time on the previous week, and the initial index data of a fifth historical time series in the last three hours of the same time as the current time on the previous week.

[0062] According to embodiments of this disclosure, the target preprocessed data in the preprocessed data can be updated using benchmark data. Specifically, abnormal initial indicator data in the target preprocessed data is replaced with normal indicator data from the benchmark data to obtain the target initial data to be processed. All indicator data in the target initial data to be processed that correspond to the target preprocessed data are normal indicator data.

[0063] According to embodiments of this disclosure, updating the target preprocessed data in the preprocessed data using benchmark data can ensure that the indicator data of historical time series within the same time period as the current time includes both normal and abnormal indicator data; while the indicator data of at least one historical time series within a different time period from the current time, such as the indicator data of the second historical time series, the indicator data of the third historical time series, the indicator data of the fourth historical time series, and the indicator data of the fifth historical time series, each includes only normal indicator data.

[0064] By utilizing the method for determining the data to be processed provided in this embodiment, the fluctuation pattern of normal indicator data can be learned by using at least one historical time series that is in a different time period from the current time, and the real-time fluctuation phenomenon can be learned by using a historical time series that is in the same time period as the current time. By comparing the difference between real-time fluctuation and normal fluctuation, the accuracy of identifying abnormal indicator data can be improved.

[0065] Figure 4 A flowchart illustrating the acquisition of reference data according to another embodiment of this disclosure is shown schematically.

[0066] like Figure 4 As shown, initial data to be processed 410 can be obtained, which includes indicator data at the current time and initial indicator data of multiple historical time series associated with the current time. However, it is not limited to this; the initial data can also be indicator data at various times of a continuous historical time series. Historical labels 420 for characterizing whether the initial indicator data is abnormal can also be obtained from the initial indicator data of multiple historical time series. The initial data to be processed 410 can be preprocessed based on the historical labels 420 to obtain preprocessed data 430. Abnormal data filtering is performed on the preprocessed data 410 to obtain filtered data 440. The filtered data 440 is aggregated to form aggregated data 450. Based on the aggregated data 450, multiple initial benchmark data are obtained. The multiple initial benchmark data may include first initial benchmark data 461, second initial benchmark data 462, and third initial benchmark data 463. The multiple initial benchmark data can be weighted and summed to obtain benchmark data 470.

[0067] According to embodiments of this disclosure, by performing anomaly filtering on preprocessed data, abnormal indicator data can be filtered out, while normal indicator data is retained.

[0068] According to embodiments of this disclosure, filtered data from the same moment across different time periods can be aggregated to form aggregated data. For example, aggregating indicator data at 3:00 AM can aggregate indicator data from the current day at 3:00 AM, indicator data from the previous day at 3:00 AM, ..., up to indicator data from the previous two months at 3:00 AM.

[0069] According to embodiments of this disclosure, indicator data can be extracted from aggregated data according to time periods and moments, and exponential smoothing can be performed to obtain first initial benchmark data, second initial benchmark data, and third initial benchmark data. The first initial benchmark data can be generated by extracting indicator data from the aggregated data for the previous N1 weeks and smoothing the indicator data for the previous N1 weeks to obtain the first initial benchmark data. The second initial benchmark data can be generated by extracting indicator data from the aggregated data for the same moment in the previous N2 days and smoothing it to obtain the second initial benchmark data. The third initial benchmark data can be generated by extracting indicator data from the aggregated data for the previous N3 moments and smoothing it to obtain the third initial benchmark data. The smoothing process can be exponential smoothing, but is not limited to this; it can also be a moving average. N1, N2, and N3 can be any integer.

[0070] According to embodiments of this disclosure, weights can be assigned to the first initial benchmark data, the second initial benchmark data, and the third initial benchmark data, and benchmark data can be obtained using a weighted summation method. The indicator data in the benchmark data are all normal indicator data. Abnormal indicator data in the target preprocessed data can be replaced using the normal indicator data from the benchmark data.

[0071] By using the benchmark data acquisition method provided in this embodiment, the indicator data of multiple time periods are aggregated, and the indicator data of different time periods are weighted and summed, which makes the benchmark data more reasonable and consistent with the rules.

[0072] According to embodiments of this disclosure, preprocessing the initial data to be processed to obtain preprocessed data may include the following operations.

[0073] For example, missing values ​​are imputed in the initial data to be processed, resulting in imputed data. The imputed data is then smoothed, resulting in smoothed data. Finally, the smoothed data is normalized, resulting in preprocessed data.

[0074] According to embodiments of this disclosure, complete data can be obtained by imputing missing values ​​in the initial data to be processed. This missing value imputation method can address the problem of decreased accuracy in identifying outlier data due to missing data.

[0075] According to embodiments of this disclosure, the imputed data can be smoothed to obtain smoothed data with glitch removal. Kalman filtering can be used for smoothing. Smoothing can highlight the fluctuation patterns of indicator data over time, reduce the impact of glitch on these fluctuation patterns, and thus improve the accuracy of subsequent identification of abnormal indicator data.

[0076] According to embodiments of this disclosure, smoothed data can be normalized so that the numerical range of indicator data in the preprocessed data is, for example, between 0 and 1, thereby improving the ability of subsequent filtering operations using regression methods to transfer between different types of indicator data.

[0077] Figure 5 A flowchart illustrating a training method for an anomaly determination model according to an embodiment of the present disclosure is shown.

[0078] like Figure 5 As shown, the method includes operations S510 to S520.

[0079] When operating S510, the regression method is used to filter the training samples to obtain the target training samples.

[0080] In operation S520, the anomaly determination model is trained using the target training samples to obtain the trained anomaly determination model. The target training samples include sample data and sample labels. The sample data includes the sample index data at the current time and the sample index data of multiple historical time series associated with the current time. The sample labels include the sample labels at the current time, which are used to characterize whether the sample index data at the current time is anomaly.

[0081] According to embodiments of this disclosure, the training method for the anomaly determination model may include only operation S520, but is not limited thereto, and may also include operations S510 and S520.

[0082] According to embodiments of this disclosure, the anomaly determination model may include at least one of the following: a logistic regression model and a decision tree model.

[0083] According to embodiments of this disclosure, the XGBoost model in the decision tree model can be employed.

[0084] According to embodiments of this disclosure, multiple historical time series associated with the current moment may include historical time series within the same time period as the current moment, and at least one historical time series within a different time period. These multiple historical time series may have the same duration and all fall within the same historical time period. For example, multiple historical time series may include a first historical time series of the first three hours before the current moment, a second historical time series of the first three hours of the same moment on the previous day, a third historical time series of the last three hours of the same moment on the previous day, a fourth historical time series of the first three hours of the same moment on the previous week, and a fifth historical time series of the last three hours of the same moment on the previous week. By utilizing sample indicator data from historical time series within the same time period as the current moment, the characteristics of the sample indicator data most associated with the current moment can be learned. Combining the sample indicator data from at least one historical time series within a different time period, the characteristics of more periodic indicator data can also be considered, such as sample indicator data with large fluctuations and personalized fluctuation cycles, such as holidays and release cycles. This allows the trained anomaly detection model to be applied to the field of anomaly identification of periodically continuous indicator data, improving identification accuracy.

[0085] According to an embodiment of this disclosure, for operation S520, the anomaly determination model is trained using training samples to obtain the trained anomaly determination model, which can be performed by the following operations.

[0086] For example, extracting features from sample data to obtain target sample feature vectors; and using the target sample feature vectors and sample labels to train an anomaly determination model to obtain the trained anomaly determination model.

[0087] According to embodiments of this disclosure, a time series feature extraction module, such as the Tsfresh module, can be used to perform the operation of extracting features from sample data. The time series feature extraction module provided in this disclosure can achieve automated feature extraction using methods such as variance filtering, and can establish an adaptive feature set of abnormal scene patterns, thereby reducing the risk of overfitting in the anomaly determination model during training.

[0088] According to an embodiment of this disclosure, for operation S510, the training samples are screened using a regression method to obtain the target training samples, which can be performed by the following operations.

[0089] For example, the sample data in the training samples are input into multiple regression algorithms to obtain multiple sample regression results, wherein the multiple sample regression results correspond one-to-one with the multiple regression algorithms, wherein the training samples are extracted from the target initial training samples according to a predetermined time window length; and in response to determining that at least one of the multiple sample regression results is the target sample regression result, the training samples are used as the target training samples, wherein the target sample regression result is used to characterize the sample indicator data at the current time as an anomaly.

[0090] According to embodiments of this disclosure, at least one of the plurality of regression algorithms may include at least one of the following: polynomial fitting algorithm, exponential moving average algorithm, isolated forest algorithm, and standard deviation method.

[0091] According to embodiments of this disclosure, in response to determining that multiple sample regression results are all normal sample regression results used to characterize the sample index data at the current moment, the training sample can be deleted without further processing.

[0092] The training method for the anomaly detection model provided in this disclosure utilizes regression methods to screen training samples, which increases the number of anomalous sample indicator data and reduces the imbalance of target training samples used to train the anomaly detection model. This avoids the problem of disrupting fluctuation patterns caused by oversampling or anomaly injection to increase the number of anomalous sample indicator data. Furthermore, using multiple regression algorithms for screening—that is, employing various screening methods—can identify target training samples with different fluctuation types, dimensions, or personalized fluctuation periods, thereby improving the training convergence speed of the anomaly detection model.

[0093] According to embodiments of this disclosure, each of the plurality of regression algorithms can be trained using the training samples provided in the embodiments of this disclosure to obtain a trained regression algorithm. In the operation of training anomaly determination models using target training samples, multiple trained regression algorithms can be used to filter target training samples from the training samples, thereby improving the convergence speed of training the anomaly determination model.

[0094] According to embodiments of this disclosure, training samples can be determined through the following operations.

[0095] For example, obtaining initial sample data; preprocessing the initial sample data to obtain preprocessed sample data; and using benchmark sample data to update the target preprocessed sample data in the preprocessed sample data to obtain the target initial sample data.

[0096] According to embodiments of this disclosure, target preprocessed sample data in preprocessed sample data can be updated using benchmark sample data. Specifically, abnormal initial sample indicator data in the target preprocessed sample data is replaced with normal sample indicator data from the benchmark sample data to obtain target initial sample data. The sample indicator data in the target initial sample data corresponding to the target preprocessed sample data are all normally fluctuating sample indicator data.

[0097] According to embodiments of this disclosure, the target preprocessed sample data in the preprocessed sample data may refer to the initial sample index data of at least one historical time series that is in a different time period from the current time.

[0098] According to embodiments of this disclosure, sample indicator data from at least one historical time series at a different time period than the current time can be used to learn the fluctuation patterns of normal sample indicator data, while sample indicator data from historical time series at the same time period as the current time can be used to learn the fluctuation phenomena of real-time sample indicator data. By comparing the differences between real-time fluctuations and normal fluctuations, the accuracy of identifying abnormal indicator data can be improved.

[0099] According to embodiments of this disclosure, benchmark sample data can be obtained in the following manner.

[0100] For example, the preprocessed sample data is filtered for outliers to obtain filtered sample data; the filtered sample data from multiple historical moments is aggregated to form aggregated sample data; based on the aggregated sample data, multiple initial benchmark sample data are obtained; and the multiple initial benchmark sample data are weighted and summed to obtain benchmark sample data.

[0101] According to embodiments of this disclosure, by performing anomaly filtering on preprocessed sample data, abnormal sample indicator data can be filtered out, while normal sample indicator data is retained.

[0102] According to embodiments of this disclosure, sample index data can be extracted from aggregated sample data at different time periods and times, and exponential smoothing can be performed to obtain first initial benchmark sample data, second initial benchmark sample data, and third initial benchmark sample data.

[0103] By using the benchmark sample data acquisition method provided in this embodiment, sample indicator data from multiple time periods can be aggregated, and the sample indicator data from different time periods can be weighted and summed, making the benchmark sample data more reasonable and consistent with the rules.

[0104] According to embodiments of this disclosure, preprocessing the initial sample data to obtain preprocessed sample data may include the following operations.

[0105] For example, missing values ​​are imputed in the initial sample data to obtain imputed sample data; the imputed sample data is smoothed to obtain smoothed sample data; and the smoothed sample data is normalized to obtain preprocessed sample data.

[0106] According to embodiments of this disclosure, before performing operation S510, the initial sample data can be preprocessed, and the smoothed sample data can be normalized to improve the ability to transfer different types of sample indicator data when performing operation S510.

[0107] Figure 6 A block diagram of an anomaly determination apparatus according to an embodiment of the present disclosure is shown schematically.

[0108] like Figure 6 As shown, the anomaly determination device 600 may include an extraction module 610 and a determination module 620.

[0109] The extraction module 610 is used to extract features from the target data to be processed to obtain a target feature vector. The target data to be processed includes the indicator data at the current time and the indicator data of multiple historical time series associated with the current time.

[0110] The determination module 620 is used to obtain the label of the indicator data at the current time based on the target feature vector, wherein the label is used to characterize whether the indicator data at the current time is abnormal.

[0111] According to embodiments of this disclosure, the anomaly determination device may further include a screening module.

[0112] The filtering module is used to filter the data to be processed using regression methods to obtain the target data to be processed.

[0113] According to embodiments of this disclosure, the filtering module may include an input unit and a filtering unit.

[0114] The input unit is used to input the data to be processed into multiple regression algorithms to obtain multiple regression results. Each regression result corresponds to a different regression algorithm. The data to be processed is extracted from the initial target data according to a predetermined time window.

[0115] A filtering unit is used to select the data to be processed as the target data to be processed in response to determining that at least one of the multiple regression results is the target regression result, wherein the target regression result is used to characterize the indicator data at the current moment as an anomaly.

[0116] According to embodiments of this disclosure, the anomaly determination device may further include an acquisition module, a preprocessing module, and an update module.

[0117] The acquisition module is used to acquire the initial data to be processed.

[0118] The preprocessing module is used to preprocess the initial data to be processed, and obtain the preprocessed data.

[0119] The update module is used to update the target preprocessed data in the preprocessed data using the baseline data, so as to obtain the target initial data to be processed.

[0120] According to embodiments of this disclosure, the preprocessing module may include a filling unit, a smoothing unit, and a normalization unit.

[0121] The imputation unit is used to impute missing values ​​in the initial data to be processed, resulting in imputed data.

[0122] The smoothing unit is used to smooth the filled data to obtain smoothed data.

[0123] The normalization unit is used to normalize the smoothed data to obtain preprocessed data.

[0124] According to embodiments of this disclosure, the anomaly determination device may further include a filtering module, an aggregation module, a pre-benchmark module, and a weighting module.

[0125] The filtering module is used to filter out abnormal data from the preprocessed data to obtain filtered data.

[0126] The aggregation module is used to aggregate filtered data from multiple historical moments to form aggregated data.

[0127] The pre-benchmark module is used to obtain multiple initial benchmark data based on the aggregated data.

[0128] The weighting module is used to perform weighted summation on multiple initial benchmark data to obtain benchmark data.

[0129] According to embodiments of this disclosure, at least one of the plurality of regression algorithms includes at least one of the following: polynomial fitting algorithm, exponential moving average algorithm, isolated forest algorithm, and standard deviation method.

[0130] Figure 7 A block diagram of a training apparatus for an anomaly determination model according to an embodiment of the present disclosure is shown schematically.

[0131] like Figure 7 As shown, the training device 700 for the anomaly determination model may include a sample screening module 710 and a training module 720.

[0132] The sample screening module 710 is used to screen training samples using regression methods to obtain target training samples.

[0133] The training module 720 is used to train an anomaly determination model using target training samples to obtain the trained anomaly determination model. The target training samples include sample data and sample labels. The sample data includes sample indicator data at the current time and sample indicator data of multiple historical time series associated with the current time. The sample labels include the sample labels at the current time, which are used to characterize whether the sample indicator data at the current time is anomaly.

[0134] According to embodiments of this disclosure, the training apparatus for the anomaly determination model may include only the training module 720, but is not limited thereto, and may also include a sample screening module 710 and a training module 720.

[0135] According to embodiments of this disclosure, the training module may include a sample extraction unit and a sample determination unit.

[0136] The sample extraction unit is used to extract features from the sample data to obtain the target sample feature vector.

[0137] The sample determination unit is used to train the anomaly determination model using the feature vector of the target sample and the sample label, and obtain the trained anomaly determination model.

[0138] According to embodiments of this disclosure, the sample screening module may include a sample input unit and a sample screening unit.

[0139] The sample input unit is used to input the sample data in the training samples into multiple regression algorithms to obtain multiple sample regression results. The multiple sample regression results correspond one-to-one with the multiple regression algorithms. The training samples are extracted from the target initial training samples according to a predetermined time window length.

[0140] The sample selection unit is used to select training samples as target training samples in response to determining at least one of the multiple sample regression results as the target sample regression result, wherein the target sample regression result is used to characterize the sample indicator data at the current time as an anomaly.

[0141] According to embodiments of this disclosure, the training apparatus for the anomaly determination model may further include a sample acquisition module, a sample preprocessing module, and a sample update module.

[0142] The sample acquisition module is used to acquire initial sample data.

[0143] The sample preprocessing module is used to preprocess the initial sample data to obtain preprocessed sample data.

[0144] The sample update module is used to update the target preprocessed sample data in the preprocessed sample data using the benchmark sample data, so as to obtain the target initial sample data.

[0145] According to embodiments of this disclosure, the sample preprocessing module may include a sample imputation unit, a sample smoothing unit, and a sample normalization unit.

[0146] The sample imputation unit is used to impute missing values ​​in the initial sample data to obtain imputed sample data.

[0147] The sample smoothing unit is used to smooth the imputed sample data to obtain smoothed sample data.

[0148] The sample normalization unit is used to normalize the smoothed sample data to obtain preprocessed sample data.

[0149] According to embodiments of this disclosure, the training apparatus for the anomaly determination model may further include a sample filtering module, a sample aggregation module, a sample pre-benchmark module, and a sample weighting module.

[0150] The sample filtering module is used to filter out abnormal data from the preprocessed sample data to obtain filtered sample data.

[0151] The sample aggregation module is used to aggregate filtered sample data from multiple historical moments to form aggregated sample data.

[0152] The sample pre-benchmark module is used to obtain multiple initial benchmark sample data based on the aggregated sample data.

[0153] The sample weighting module is used to perform weighted summation on multiple initial benchmark sample data to obtain benchmark sample data.

[0154] According to embodiments of this disclosure, at least one of the plurality of regression algorithms includes at least one of the following: polynomial fitting algorithm, exponential moving average algorithm, isolated forest algorithm, and standard deviation method.

[0155] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0156] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0157] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.

[0158] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0159] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0160] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0161] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0162] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as anomaly determination methods or anomaly determination model training methods. For example, in some embodiments, the anomaly determination method or anomaly determination model training method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the anomaly determination method or anomaly determination model training method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other suitable manner (e.g., by means of firmware) to perform an anomaly determination method or an anomaly determination model training method.

[0163] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0164] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0165] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0167] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0168] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0169] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0170] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An anomaly determination method, comprising: The data to be processed is input into multiple regression algorithms to obtain multiple regression results. The multiple regression results correspond one-to-one with the multiple regression algorithms. The data to be processed is extracted from the target initial data to be processed according to a predetermined time window length. The data to be processed includes the indicator data at the current moment and the indicator data of multiple historical time series associated with the current moment. The indicator data includes the application crash rate or the number of stutters. In response to determining that at least one of the plurality of regression results is the target regression result, the data to be processed is taken as the target data to be processed, wherein the target regression result is used to characterize the indicator data at the current moment as an anomaly; Extract the features of the target data to be processed to obtain the target feature vector; and Based on the target feature vector, a label is obtained for the indicator data at the current moment, wherein the label is used to characterize whether the indicator data at the current moment is abnormal.

2. The method according to claim 1, further comprising: Obtain the initial data to be processed; The initial data to be processed is preprocessed to obtain preprocessed data; as well as The target preprocessed data in the preprocessed data is updated using the benchmark data to obtain the target initial data to be processed.

3. The method according to claim 2, wherein, The preprocessing of the initial data to be processed to obtain preprocessed data includes: The initial data to be processed is imputed to obtain the imputed data; The imputed data is smoothed to obtain smoothed data; and The smoothed data is normalized to obtain the preprocessed data.

4. The method according to claim 2 or 3, further comprising: The preprocessed data is then filtered for abnormal data to obtain filtered data. The filtered data from multiple historical moments are aggregated to form aggregated data; Based on the aggregated data, multiple initial baseline data are obtained; as well as The multiple initial benchmark data are weighted and summed to obtain the benchmark data.

5. The method according to claim 1, wherein, At least one of the plurality of regression algorithms includes at least one of the following: Polynomial fitting algorithm, exponential moving average algorithm, isolated forest algorithm, standard deviation method.

6. A training method for an anomaly determination model, comprising: The sample data from the training samples are input into multiple regression algorithms to obtain multiple sample regression results. Each of these multiple sample regression results corresponds one-to-one with one of the multiple regression algorithms. The training samples are extracted from the initial target training samples according to a predetermined time window length. In response to determining that at least one of the plurality of sample regression results is the target sample regression result, the training sample is used as the target training sample, wherein the target sample regression result is used to characterize the sample indicator data at the current moment as an anomaly; The anomaly detection model is trained using the target training samples to obtain the trained anomaly detection model. The target training samples include sample data and sample labels. The sample data includes sample indicator data at the current moment and sample indicator data of multiple historical time series associated with the current moment. The sample labels include the sample labels at the current moment, which are used to characterize the sample indicator data at the current moment as abnormal. The sample indicator data includes the application's crash rate or number of stutters.

7. The method according to claim 6, wherein, The anomaly detection model is trained using training samples, resulting in the following trained anomaly detection model: Extract the features from the sample data to obtain the target sample feature vector; and The anomaly determination model is trained using the target sample feature vector and the sample label to obtain the trained anomaly determination model.

8. The method according to claim 6, further comprising: Obtain initial sample data; The initial sample data is preprocessed to obtain preprocessed sample data; as well as The target preprocessed sample data in the preprocessed sample data is updated using the benchmark sample data to obtain the target initial training sample.

9. The method according to claim 8, wherein, The preprocessing of the initial sample data to obtain preprocessed sample data includes: The initial sample data is imputed to obtain the imputed sample data; The imputed sample data is smoothed to obtain smoothed sample data; and The smoothed sample data is normalized to obtain the preprocessed sample data.

10. The method according to claim 8 or 9, further comprising: The preprocessed sample data is filtered for abnormal data to obtain filtered sample data. The filtered sample data from multiple historical moments in the filtered sample data are aggregated to form aggregated sample data; Based on the aggregated sample data, multiple initial benchmark sample data are obtained; and The multiple initial benchmark sample data are weighted and summed to obtain the benchmark sample data.

11. An anomaly determination device, comprising: The filtering module is used to filter the data to be processed using regression methods to obtain target data to be processed. The data to be processed includes indicator data at the current moment and indicator data of multiple historical time series associated with the current moment. The indicator data includes the application's crash rate or number of lags. An extraction module is used to extract features from the target data to be processed, thereby obtaining a target feature vector; and The determination module is used to obtain a label for the indicator data at the current time based on the target feature vector, wherein the label is used to characterize whether the indicator data at the current time is abnormal; The filtering module includes: An input unit is used to input the data to be processed into multiple regression algorithms to obtain multiple regression results, wherein each regression result corresponds one-to-one with one of the multiple regression algorithms, and the data to be processed is extracted from the initial target data to be processed according to a predetermined time window length; and A filtering unit is configured to, in response to determining that at least one of the plurality of regression results is a target regression result, use the data to be processed as the target data to be processed, wherein the target regression result is used to characterize the indicator data at the current moment as an anomaly.

12. The apparatus of claim 11, further comprising: The acquisition module is used to acquire the initial data to be processed. The preprocessing module is used to preprocess the initial data to be processed to obtain preprocessed data; as well as The update module is used to update the target preprocessed data in the preprocessed data using the reference data to obtain the target initial data to be processed.

13. The apparatus according to claim 12, wherein, The preprocessing module includes: The filling unit is used to fill in missing values ​​in the initial data to be processed to obtain the filled data. A smoothing unit is used to smooth the filled data to obtain smoothed data; and The normalization unit is used to normalize the smoothed data to obtain the preprocessed data.

14. The apparatus according to claim 12 or 13, further comprising: The filtering module is used to filter out abnormal data from the preprocessed data to obtain filtered data. The aggregation module is used to aggregate filtered data from multiple historical moments in the filtered data to form aggregated data. The pre-benchmark module is used to obtain multiple initial benchmark data based on the aggregated data; as well as The weighting module is used to perform weighted summation on the multiple initial benchmark data to obtain the benchmark data.

15. The apparatus according to claim 11, wherein, At least one of the plurality of regression algorithms includes at least one of the following: Polynomial fitting algorithm, exponential moving average algorithm, isolated forest algorithm, standard deviation method.

16. A training apparatus for an anomaly determination model, comprising: The sample selection module is used to select training samples using regression methods to obtain target training samples; The training module is used to train the anomaly detection model using the target training samples, thereby obtaining the trained anomaly detection model. The target training samples include sample data and sample labels. The sample data includes sample indicator data at the current moment and sample indicator data of multiple historical time series associated with the current moment. The sample labels include the sample labels at the current moment, which are used to characterize the sample indicator data at the current moment as abnormal. The sample screening module includes: A sample input unit is used to input sample data from the training samples into multiple regression algorithms to obtain multiple sample regression results, wherein each of the multiple sample regression results corresponds one-to-one with the multiple regression algorithms, and the training samples are extracted from the target initial training samples according to a predetermined time window length; and A sample screening unit is configured to, in response to determining that at least one of the plurality of sample regression results is a target sample regression result, use the training sample as the target training sample, wherein the target sample regression result is used to characterize the sample indicator data at the current moment as abnormal, and the sample indicator data includes the application's crash rate or number of stutters.

17. The apparatus according to claim 16, wherein, The training module includes: A sample extraction unit is used to extract features from the sample data to obtain a target sample feature vector; and The sample determination unit is used to train the anomaly determination model using the target sample feature vector and the sample label to obtain the trained anomaly determination model.

18. The apparatus of claim 16, further comprising: The sample acquisition module is used to acquire initial sample data; The sample preprocessing module is used to preprocess the initial sample data to obtain preprocessed sample data; as well as The sample update module is used to update the target preprocessed sample data in the preprocessed sample data using the benchmark sample data, so as to obtain the target initial training sample.

19. The apparatus according to claim 18, wherein, The sample preprocessing module includes: The sample imputation unit is used to impute missing values ​​in the initial sample data to obtain imputed sample data. A sample smoothing unit is used to smooth the imputed sample data to obtain smoothed sample data; and The sample normalization unit is used to normalize the smoothed sample data to obtain the preprocessed sample data.

20. The apparatus according to claim 18 or 19, further comprising: The sample filtering module is used to filter out abnormal data from the preprocessed sample data to obtain filtered sample data. The sample aggregation module is used to aggregate filtered sample data from multiple historical moments in the filtered sample data to form aggregated sample data. The sample pre-benchmark module is used to obtain multiple initial benchmark sample data based on the aggregated sample data; as well as The sample weighting module is used to perform weighted summation on the multiple initial benchmark sample data to obtain the benchmark sample data.

21. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the anomaly determination method of any one of claims 1 to 5 or the training method of the anomaly determination model of any one of claims 6 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the anomaly determination method according to any one of claims 1 to 5 or the anomaly determination model training method according to any one of claims 6 to 10.

23. A computer program product comprising a computer program that, when executed by a processor, implements the anomaly determination method according to any one of claims 1 to 5 or the training method of the anomaly determination model according to any one of claims 6 to 10.

Citation Information

Patent Citations

  • Abnormality detection model training method, anomaly detection method and device

    CN113743607A