Anomaly and Change Detection Using the Robustness of Sparse Decomposition
By decomposing the time series into potential components and identifying anomalies based on the significance threshold, the problems of inaccuracy and inefficiency in existing systems when identifying time series anomalies are solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202110358292.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-17
- Filing Date
- 2021-04-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-04-01
AI Technical Summary
Existing analytical computing systems have inaccuracy and inefficiency when identifying abnormalities in time series, especially when dealing with multiple periodic trends or data level changes, conventional systems are prone to misidentifying data spikes and levels of changes as abnormalities.
Anomaly data are identified by decomposing the time series into latent components and satisfying the significance threshold based on one or both of the spike/depression and horizontal changes of the latent components. The system intelligently makes real values obey potential component constraints, eliminates non-real values, and thus improves the accuracy and efficiency of abnormal detection.
It improves the accuracy and efficiency of abnormal detection, reduces the possibility of misidentification of abnormalities, and effectively utilizes computing resources, so that significant abnormalities in the time series can be more accurately identified.
Smart Images

Figure CN113806122B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to anomaly and change detection that utilizes the robustness of sparse decomposition. Background Art
[0002] In recent years, analytical computing systems have improved the accuracy of identifying trends and seasonal variations by configuring new algorithms for analyzing time series, which include datasets of metrics recorded over time. For example, conventional analytical computing systems can identify and present anomalies from time series representing user actions regarding website, network-accessible applications, or other network-based device operations. By way of illustration, some existing systems can decompose a time series of a large amount of network metric data into multiple components as a basis for identifying anomalous metrics in the time series, such as anomalous user actions outside of an expected trend.
[0003] Although conventional analytical computing systems can identify anomalies in a time series, such systems may inaccurately and ineffectively identify outliers within the time series by applying conventional anomaly detection algorithms. For example, when the time series represents multiple periodic trends or when the time series does not reflect changes in data level, conventional systems sometimes misidentify anomalies in the time series data. Specifically, although some existing systems analyze potential component sequences from a time series to identify anomalous data spikes and level changes, these existing systems typically require a large amount of user input. For example, some existing analytical computing systems require users to input a periodic frequency and a maximum number of anomalies as a basis for identifying anomalies within the time series. However, even with such user input, existing analytical computing systems still continue to inaccurately identify data spikes and level changes as false positives in terms of anomalies within the time series. Such systems may misidentify data spikes and level changes by uniformly applying an anomaly detection algorithm to all values within the time series, even if some of these values may be certain non-real values (e.g., missing values, unavailable values, non-numeric values, infinite values) or other values that distort the correct identification of the data trend.
[0004] By way of illustration, some analytical computing systems may misidentify anomalous metrics from a time series due to the inability to account for missing values or non-real values. For example, in some cases, a time series may include both real values and non-real values. By performing an algorithm that is uniformly applied to missing values, real values, non-real values, etc., conventional systems typically label spikes or dips and level changes as anomalies in the time series when these outliers reflect the idiosyncratic application of the algorithm to missing values or non-real values rather than outliers.
[0005] By applying anomaly detection techniques and protocols to values that may ensure incorrect results, conventional analytical computing systems cannot effectively utilize computing resources. For example, as just discussed, conventional systems typically analyze the entire sequence of potential components of an entire time series (usually including non-real and / or non-significant values) to identify various anomalies. However, by analyzing potentially unnecessary values in the entire sequence of potential components, conventional systems waste computing resources to identify potential errors and / or non-significant anomalies.
[0006] In addition to this inaccuracy and inefficiency, conventional analytical computing systems are also unable to effectively divide a dataset into a training sequence and a test sequence to train and test an anomaly detection algorithm on the dataset. For example, conventional systems often divide a time series into (i) data corresponding to a training period for tuning an anomaly detection algorithm and (ii) data corresponding to a test period for identifying anomalies in the test period. When the training period of a time series does not accurately represent the test period of the time series due to periodic data, special events, or other events, such conventional systems may not be able to identify anomalies by dividing the data into periods.
[0007] Independent of the inefficient division of the dataset, some conventional analytical computing systems apply anomaly detection algorithms rigidly. For example, as described above, some conventional systems apply an anomaly detection algorithm to a time series regardless of the type of underlying data or changes in values in the time series. By ignoring changes in data type or value, some conventional systems may misidentify different periodic variations, zero values, or non-real values as indicating anomaly values. In addition, as described above, some conventional systems can only identify anomalies in the test period of a time series dataset after training an anomaly detection algorithm only on the training period of the time series dataset. However, the strict reliance on the test and training periods of the data within a time series ignores important changes in the time series that may lead to critical analytical insights.
[0008] Conventional systems have these and other problems. Summary of the Invention
[0009] One or more embodiments of systems, non-transitory computer-readable media, and methods are described that solve the foregoing or other problems or provide other benefits. In particular, the disclosed systems determine latent components of a metric time series and identify anomalous data within the metric time series based on one or both of spikes / dips and level changes of the latent components meeting a significance threshold. To identify such latent components, in some cases, the disclosed systems account for the range of value types by intelligently subjecting real values to latent component constraints for time series decomposition and intelligently excluding non-real values from the latent component constraints. The disclosed systems may further identify significant anomalous data values from the latent components of the metric time series by jointly determining whether one or both of a subsequence of the spike component sequence and level changes of the level component sequence meet the significance threshold. In some cases, the disclosed systems may further modify the metric time series and its latent components by considering other time-based data fluctuations or data types, including data in a time series reflecting special day or time effects, leading or trailing zeros, and low event counts, to improve anomaly detection. In another embodiment, the disclosed systems may estimate missing data values from the metric time series by considering, for example, complex periodic patterns or level changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] As briefly described below, the detailed description provides additional specificity and detail for one or more embodiments by using the drawings.
[0011] Figure 1 FIG. 1 shows an example environment in which an anomaly detection system according to one or more embodiments may operate;
[0012] Figures 2A - 2F FIG. 2 shows an example metric time series and associated component sequences according to one or more embodiments;
[0013] Figure 3 FIG. 3 shows an overview of an anomaly detection system that determines latent components of a metric time series and identifies significant anomalies based on the latent components according to one or more embodiments;
[0014] Figure 4 FIG. 4 shows a process of identifying different data types, performing an optimization algorithm to determine latent components of a metric time series, and determining significant anomalies according to the latent components of the metric time series according to one or more embodiments;
[0015] Figures 5A - 5D FIG. 5 shows a graphical user interface for selecting a metric time series, determining latent components of the metric time series, and identifying anomalous data associated with the selected metric time series according to one or more embodiments;
[0016] Figure 6 A schematic diagram of an anomaly detection system according to one or more embodiments is shown;
[0017] Figure 7 A flowchart is shown for determining latent components of a metric time series according to one or more embodiments and identifying anomalous data within the metric time series based on one or both of a spike and a level change of the latent components meeting a significance threshold;
[0018] Figure 8 A flowchart is shown for simultaneously determining whether a subsequence of a spike component sequence and a level change corresponding to a level component sequence represent a significant anomaly according to one or more embodiments; and
[0019] Figure 9 A block diagram of an example computing device for implementing one or more embodiments of the present disclosure is shown. Detailed Description
[0020] The present disclosure describes one or more embodiments of an anomaly detection system that decomposes a metric time series into latent components and determines one or both of a spike and a level change indicating an anomalous data value from the latent components based on a significance threshold. Such latent components can include at least a spike component sequence and a level component sequence. As part of decomposing the metric time series, the anomaly detection system can consider the range of value types by: (i) intelligently subjecting real values in the metric time series to latent component constraints that define the relationship between the metric time series and the latent components, and (ii) intelligently excluding non-real values from the latent component constraints.
[0021] As a basis for identifying anomalous data values, the anomaly detection system can also jointly determine whether one or both of a subsequence of the spike component sequence and a level change of the level component sequence meet the significance threshold. In one or more embodiments, the anomaly detection system modifies the values of the metric time series and its latent components by considering special days or time effects, leading or trailing zeros, or low event counts to avoid common pitfalls in anomaly detection. By identifying significant anomalous values from the latent components of the metric time series and performing other operations described herein, the anomaly detection system improves the accuracy, efficiency, and flexibility of conventional anomaly detection. In another embodiment, the disclosed system can estimate missing data values from the metric time series by considering, for example, complex periodic patterns or level changes.
[0022] For further illustration, in some cases, an anomaly detection system can retrieve or access a metric time series that includes metric data values representing user actions within a digital network corresponding to a time period (e.g., data on user actions regarding a website or other network platform). The anomaly detection system also determines one or more potential components of the metric time series, such as a spike component sequence and a level component sequence. The anomaly detection system can also determine whether a subsequence of the spike component sequence meets a spike significance threshold and whether a level change corresponding to the level component sequence meets a level change significance threshold. Based on either or both of the subsequence meeting the spike significance threshold and the level change meeting the level change significance threshold, the anomaly detection system can also generate anomaly data values to be displayed on a client computing device.
[0023] In one or more embodiments, the anomaly detection system can determine at least one spike component sequence and one level component sequence as potential components of the retrieved metric time series. For example, the anomaly detection system can determine the potential components of the metric time series by applying an optimization algorithm to the metric time series. In one or more embodiments, the optimization algorithm is configured to determine potential components that, when combined, form the metric time series. For example, when executed, the optimization algorithm iteratively minimizes an objective function that decomposes the metric time series into different potential component sequences representing spikes, level changes, periodic trends, and error.
[0024] In some cases, the optimization algorithm is subject to potential component constraints that indicate the number, type, and additional qualities of the potential components associated with the metric time series, such as by constraining the metric time series to equal different potential component sequences. Based on the optimization algorithm and the potential component constraints, the anomaly detection system can identify at least one spike component sequence and one level component sequence. In another embodiment, the anomaly detection system can also apply the optimization algorithm based on the potential component constraints to identify a periodic component sequence and an error component sequence as potential components of the metric time series.
[0025] As described above, the anomaly detection system can also configure the potential component constraints of the metric time series to apply to real values and exclude non-real values. For illustration, the metric time series can include any number of real values. Additionally, the same metric time series can include any number of non-real values, such as unavailable values (e.g., "NA" or "unavailable"), non-numeric values (e.g., "NaN" or "not a number", such as zero divided by zero), and infinite values (e.g., "INF" or "infinity"). In at least one embodiment, the anomaly detection system can configure the potential component constraints of the optimization algorithm to intelligently exclude non-real values. In such cases, when decomposing the metric time series into one or more potential components, only real values need to satisfy the potential component constraints.
[0026] As further described above, an anomaly detection system can identify significant values of potential components of a metric time series. For example, the anomaly detection system can simultaneously identify significant spikes and dips from a spike component sequence and horizontal changes from a horizontal component sequence. In one or more embodiments, the anomaly detection system can identify significant outliers of the spike component sequence by determining whether a subsequence of the spike component sequence meets a spike significance threshold. Similarly, the anomaly detection system can identify significant outliers of the horizontal component sequence by determining whether a horizontal change corresponding to the horizontal component sequence meets a horizontal change significance threshold.
[0027] In one or more embodiments, the anomaly detection system can determine whether the value of a potential component meets a relative significance threshold. For example, the anomaly detection system can determine that a subsequence of the spike component sequence meets the spike significance threshold by: (i) generating a stationary time series equal to the combination of the spike component sequence and the residual error, and (ii) determining whether the data values of the stationary time series deviate from a dataset that conforms to a normalized distribution. As further explained below, in some embodiments, the stationary time series (y s ) is equal to the spike component sequence (d) plus the residual error value (μ). The anomaly detection system can also determine that a horizontal change corresponding to the horizontal component sequence meets the horizontal change significance threshold by: (i) generating a significant horizontal change value, and (ii) determining that the absolute value of the horizontal change corresponding to the horizontal component sequence exceeds or is equal to the significant horizontal change value.
[0028] In addition, the anomaly detection system can also preprocess or filter data from the metric time series or its potential components to improve the accuracy of anomaly detection. For example, the metric time series and associated potential components can include values that cause the system to misdetect (or fail to detect) anomalies. By way of illustration, the metric time series can include regularly repeating expected spike or dip values (e.g., like those associated with special days that repeat annually such as holidays) and / or zero groups (e.g., like leading or trailing zeros). In one or more embodiments, in combination with applying an optimization algorithm to increase the accuracy of anomaly detection relative to the metric time series, the anomaly detection system can identify and ignore regularly repeating spike or dip values or remove zero groups within the metric time series or its potential components.
[0029] In addition, even if the metric time series has a small number of values, the anomaly detection system can modify one or more constraints of the optimization algorithm to more accurately identify the potential components of the metric time series. For example, the metric time series can include multiple values corresponding to multiple events relative to a particular application or website. In some embodiments, the number of events can be less than a threshold number (e.g., 15 events). In response to determining that the number of values in the metric time series is less than the threshold number, the anomaly detection system can adjust or trim the confidence interval of the expected time series derived from the given metric time series. By adjusting or trimming the confidence interval, the disclosed anomaly detection system can effectively reduce the number of false anomalies detected within the metric time series.
[0030] In some cases, the anomaly detection system performs these and other operations without dividing the data of the metric time series into training and test periods. For example, conventional analytical computing systems typically use a first portion of the metric time series to train an anomaly detection algorithm and a second portion of the metric time series to test the same anomaly detection algorithm. This approach has problems because it assumes a correlation between the first and second portions of the metric time series. The present anomaly detection system avoids this approach by incorporating the entire metric time series into the identification of significant outliers of the potential components without dividing the metric time series into data values for the training period and data values for the test period, and thus, utilizes the entire data range available within the metric time series.
[0031] As described above, the anomaly detection system has many advantages and benefits compared to conventional systems and methods. For example, the anomaly detection system improves the accuracy of the analytical computing system in detecting significant anomalous data values based on the potential components of the metric time series. By decomposing the metric time series into a sequence of potential components, some existing systems lack sufficient algorithms or reference points to determine the sequence of potential components and identify statistically significant anomalous data values. Since the sequence of potential components typically lacks existing statistical significance thresholds, in the absence of such thresholds for the spike component sequence or the periodic component sequence, especially when processing simultaneously, some anomaly detection algorithms may misidentify anomalous data values.
[0032] In some cases, the anomaly detection system can use a novel application of the significance threshold to determine the statistically significant outliers of the sequence of potential components rather than misidentifying such anomalies. For example, in some cases, the anomaly detection system determines whether a stationary time series (equal to the combination of the spike component sequence and the residual error values) deviates from a dataset that conforms to a distribution (e.g., a dataset of similar time series that conforms to a standardized distribution). As further explained below, in some embodiments, the stationary time series (y s) is equal to the spike component sequence (d) plus the residual error value (μ). Simultaneously or independently, the anomaly detection system determines whether the level change corresponding to the level component sequence deviates from the significant level change value according to the significance threshold of the level change. As further explained below, in some embodiments, the significant level change value represents the critical value of the Gaussian distribution at a given significance level multiplied by the significant error calculation of the residual error value (SE(μ)). By using such a significance threshold, the anomaly detection system generates anomaly data that avoids the pitfalls associated with conventional systems, such as misidentifying anomalies or missing significant anomalies.
[0033] Independent of identifying statistically significant spikes, significant dips, or significant level changes, the anomaly detection system can intelligently apply optimization algorithms to metric time series that include both real values and certain non-real values. As described above, in some cases, the anomaly detection system subjects the optimization algorithm to latent component constraints to decompose the metric time series into latent components that sum to the time series. By intelligently subjecting real values to latent component constraints and intelligently excluding non-real values from latent component constraints, the anomaly detection system avoids misidentifying anomalies from non-real values. Compared to conventional systems that decompose metric time series that include non-real values into strict constituent latent components, the anomaly detection system selectively applies latent component constraints to the real values of the metric time series. By autonomously decomposing the metric time series into latent components and identifying significant anomalies from such latent components, the anomaly detection system avoids user input that can cause conventional systems to misidentify anomalies.
[0034] As described above, in some embodiments, the anomaly detection system improves the accuracy and efficiency of existing anomaly detection algorithms by avoiding the training and testing time periods utilized by conventional analysis computing systems. Instead of relying on the first part of the metric time series, which may inaccurately inform the analysis of the second part of the metric time series, the anomaly detection system analyzes the entire metric time series simultaneously without any specific training period. For example, as described below, the anomaly detection system performs an optimization algorithm that iteratively minimizes an objective function to identify the latent components of the time series. In some embodiments, the anomaly detection system obviates the need for training and testing the optimization algorithm because the optimization algorithm builds the residual error into the iterative minimization of the objective function. By avoiding any dependence on potentially irrelevant training data, the anomaly detection system also avoids wasting computing resources to train the system to generate anomaly data incorrectly.
[0035] By avoiding these testing and training periods, the analysis method of the anomaly detection system is more flexible. For example, as just mentioned, in some cases, the anomaly detection system does not divide the metric time series into data values for the training period and data values for the testing period to generate anomaly data. Thus, by being freed from potentially irrelevant training data, the anomaly detection system utilizes more powerful methods than conventional systems.
[0036] In addition to the improved efficiency and accuracy described above, in some embodiments, the disclosed anomaly detection system can infer missing values by imputing values based on past and future data, while taking into account complex periodic patterns and level changes. Conventional anomaly detection algorithms cannot interpolate or extrapolate such missing values as described herein and cannot correctly identify significant anomalies in the metric time series by imputing missing values.
[0037] For example, the metric time series may be missing values or other data as part of a weekly, monthly, or other periodic pattern (e.g., weekly spikes or dips). If the number of visitors (or users of an application) to a web page typically increases every other Monday (due to a special offer every two weeks), the metric time series for one such Monday may be missing the increase due to data loss or other interfering events. By using only the values before or after the missing value or by using the values from the previous week to impute the missing value, conventional systems or conventional anomaly detection algorithms incorrectly consider (or fail to consider) such missing values as part of a weekly or other periodic pattern. In contrast, the disclosed anomaly detection system can correctly infer the missing values from past or future data. For example, when the missing value is part of a bi-weekly pattern, the disclosed anomaly detection system can identify such a bi-weekly pattern and consider the values two weeks ago or two weeks in the future to impute the missing value.
[0038] As shown by the previous discussion, the present disclosure uses various terms to describe the features and advantages of the anomaly detection system. Additional details regarding the meaning of such terms are now provided. For example, as used herein, the term "metric time series" refers to a collection of data indexed over time. In particular, a metric time series can include a collection of data representing users and / or user actions on a particular application or website that occur at various times within a specific time period. By way of illustration, a metric time series can include the number of hyperlink clicks within a particular web page each day of the year. In one or more embodiments, the metric time series includes the value of each data point collected within the time period. Thus, a metric time series of daily hyperlink clicks over a year can have 365 values, where each value represents the number of hyperlink clicks for the associated day.
[0039] In one or more embodiments, an anomaly detection system may decompose a metric time series into one or more latent components or sequences of latent components. As used herein, the terms "latent component" and "sequence of latent components" refer to components of a time series or other data set that contribute to the observations in the metric time series. For example, an anomaly detection system may decompose a metric time series into one or more of a level component sequence, a spike component sequence, a periodic component sequence, and an error component sequence. Each sequence of latent components may represent a different contribution to the metric time series. Thus, the data in the latent components of a time series may not be directly observable.
[0040] As used herein, the term "level component sequence" refers to a latent component of a metric time series whose values exhibit an increase or decrease in the mean of the metric time series, such as a piecewise increase in the mean. For example, the values in a level component sequence may include or represent a level change. As used herein, a "level change" refers to an increase or decrease in the mean of a metric time series, such as a piecewise increase in the mean across a series of data points. By way of illustration, in a metric time series that includes a single level change (e.g., the mean of the values of the metric time series demonstrates a single piecewise increase), the corresponding level component sequence may include two values: the mean of the metric time series before the mean increase and the mean of the metric time series after the mean increase.
[0041] As used herein, the term "spike component sequence" refers to a latent component of a metric time series whose values exhibit spontaneous, anomalous, or other non-periodic increases or decreases. For example, each "spike" or anomalous increase in a spike component sequence may correspond to a value in the associated metric time series that is a spontaneous, anomalous, or other non-periodic increase compared to other values in the metric time series.
[0042] As used herein, the term "periodic component sequence" refers to a latent component of a metric time series whose values fluctuate in a manner related to a time period. For example, the values of a periodic component sequence may fluctuate by day, by week, by month, by year, by season (e.g., spring, summer, fall, winter), and / or by a reference date (e.g., six weeks before Christmas according to a recognized holiday).
[0043] As used herein, the term "error component sequence" refers to statistical noise, variance, or other residual latent components other than the interpretable latent components (e.g., level, periodic, periodic latent components). For example, an error component sequence may include values that, when combined with the values of other sequences of latent components, will result in the residual values of the original metric time series.
[0044] In one or more embodiments, an anomaly detection system can determine significant values of the latent components of a metric time series based on a significance threshold simultaneously. As used herein, the term "significance threshold" refers to a predetermined threshold relative to a latent component or a portion of a latent component, above or below which indicates a significant value. A portion of a latent component can be part of an indirect use of the significance threshold. For example, if a stationary time series corresponding to a spike component sequence deviates from a dataset conforming to a standardized distribution, the anomaly detection system can identify the values in the spike component sequence as significant.
[0045] As used herein, the term "anomaly data value" refers to an outlier or a set of outliers in a dataset. For example, an anomaly data value can be a data value that is anomalously different from an expected value within a given time. By way of illustration, an anomaly data value can represent an outlier data value in a metric time series that has a statistically significant difference from an expected value. Instead of identifying all potential anomalies relative to a metric time series, the anomaly detection system identifies significant anomalies based on one or more latent components of the metric time series.
[0046] As used herein, a "significant anomaly" refers to an identified anomaly that is statistically significant relative to other identified anomalies or data values. For example, a significant anomaly can include an identified anomaly from a metric time series having a small p-value or probability value, which indicates evidence against a null hypothesis (e.g., indicating a high likelihood that the identified anomaly is significant).
[0047] Additional details regarding the anomaly detection system will now be provided with reference to the accompanying drawings. For example, Figure 1 FIG. 100 shows a schematic diagram of an example system environment 100 (e.g., "the environment" 100) for implementing an anomaly detection system 102 according to one or more embodiments. Thereafter, a more detailed description of the components and processes of the anomaly detection system 102 is provided with respect to the subsequent drawings.
[0048] As Figure 1 shown, the environment 100 includes servers 106, an administrator computing device 108, a third-party web server 112, client computing devices 116a - 116n, and a network 114. Each component of the environment 100 can communicate via the network 114, and the network 114 can be any suitable network through which computing devices can communicate. Example networks are discussed in more detail below in connection with Figure 9 more detail.
[0049] As described above, environment 100 includes an administrator computing device 108 and client computing devices 116a - 116n. The administrator computing device 108 and the client computing devices 116a - 116n can be one of various computing devices, including smart phones, tablets, smart TVs, desktop computers, laptop computers, virtual reality devices, augmented reality devices, or other computing devices as described with respect to Figure 9 the other computing devices. In some embodiments, environment 100 can include any number or arrangement of client computing devices, each associated with a different user. The administrator computing device 108 and the client computing devices 116a - 116n can also communicate with the server(s) 106 and / or third - party web servers 112 via network 114. For example, the client computing devices 116a - 116n can communicate with the third - party web servers 112 via network 114 to view and interact with one or more websites hosted by the third - party web servers 112. The third - party web servers 112 can provide user interaction data associated with one or more websites to the server(s) 106 via network 114 for anomaly analysis. The administrator computing device 108 can receive anomaly data from the server(s) 106 via network 114 for display.
[0050] In one or more embodiments, the administrator computing device 108 includes a data analysis application 110. For example, a user of the administrator computing device 108 can query metric time - series data and analyze such data by interacting with the data analysis application 110. When the data analysis application 110 is executed, the administrator computing device 108 can communicate with the data analysis system 104 to receive and display metric time - series data, latent component data, and anomaly data. Additionally, the client computing devices 116a - 116n can each include a content application 118a - 118n. For example, the content application 118a can be an application for accessing and interacting with digital content, such as a web browser application, a social networking application, a file server application, etc.
[0051] As Figure 1 further shown, environment 100 includes the server(s) 106. The server(s) 106 can include one or more individual servers that can generate, store, receive, and transmit electronic data. For example, the server(s) 106 can receive data from the client computing device 116a in the form of user input such as a keystroke stream. Additionally, the server(s) 106 can transmit data to the administrator computing device 108.
[0052] As Figure 1Further shown, the server(s) 106 may also include an anomaly detection system 102 as part of the data analysis system 104. The data analysis system 104 may communicate with one or more of a third-party web server 112 and an administrator computing device 108 to receive, generate, modify, analyze, store, and transmit digital content. For example, the data analysis system 104 may communicate with the third-party web server 112 to assemble and analyze a metric time series associated with a user interaction with an application or website hosted by the third-party web server 112. Additionally, the anomaly detection system 102 and the data analysis system 104 may store and retrieve analysis information in an analysis database 120. For example, the analysis database 120 may store metric time series data, latent component data, and anomaly data.
[0053] Although Figure 1 the anomaly detection system 102 is depicted as being on the server(s) 106, in some embodiments, the anomaly detection system 102 may be implemented (e.g., located in whole or in part) on one or more other components of the environment 100. For example, the anomaly detection system 102 may be implemented, in whole or in part, by the administrator computing device 108.
[0054] In one or more embodiments, the third-party web server 112 is at least one of an application server, a communication server, a web hosting server, a social network server, or a digital content analysis server. For example, the third-party web server 112 may receive user interaction data associated with a user interaction with network content (e.g., hyperlink click, page landing, video completion) from one or more client computing devices 116a - 116n. The third-party web server 112 may also receive user information associated with the users of the client computing devices 116a - 116n. For example, the third-party web server 112 may receive user demographic information, user account information, and user profile information. In at least one embodiment, the third-party web server 112 may include multiple servers.
[0055] Figures 2A - 2E An example of a metric time series and associated latent components is shown. For example, Figure 2A a metric time series 202 is shown. In one or more embodiments, the metric time series 202 includes real values and non-real values as well as other parts with problematic indices. For illustration, the metric time series 202 includes real values in real value portions 206a, 206b, and 206c. In at least one embodiment, the real value portions 206a - 206c include values that may be integers or decimal numbers.
[0056] As Figure 2AFurther shown, the metric time series 202 also includes non-real values in other parts. For example, in the "unavailable" part 208, the metric time series 202 includes unavailable values. In one or more embodiments, the data analysis system 104 may receive an indicator of the unavailability of a numerical value associated with a specific index or position within the metric time series 202. This value may be unavailable for various reasons, such as data corruption or network connection problems during the associated time period.
[0057] In addition, as Figure 2A shown, the metric time series 202 includes a non-numeric part 210 and an infinity value part 212. For example, the non-numeric part 210 may include one or more non-numeric values (e.g., "not a number") associated with "0 / 0" or "NaN". Similarly, the infinity value part 212 may include one or more infinity values (e.g., "infinity") associated with "INF". In one or more embodiments, the data analysis system 104 may receive an indicator that a value associated with a specific index or position within the metric time series 202 is not a number or is an infinity value. Since the data analysis system 104 cannot accurately represent these types of inputs in the metric time series 202, the data analysis system 104 may fill these values with "NaN", "0 / 0", or "INF". The data analysis system 104 may use any suitable replacement value or character to represent non-real values.
[0058] In one or more embodiments, the data analysis system 104 may represent the metric time series as a trend over time. For example, as Figure 2B shown, the data analysis system 104 may generate a metric-time series 216 that represents the metrics of a specific metric time series. In one or more embodiments, the data analysis system 104 may generate the metric-time series 216 by plotting points in the metric-time series 216 for each value in the associated metric time series, where the x-value of the point is the date and / or time associated with the value, and the y-value of the point is the value.
[0059] As described above, and as described below, the anomaly detection system 102 may decompose the metric time series into one or more potential component sequences. In one or more embodiments, the anomaly detection system 102 may also generate a trend associated with each potential component sequence. For example, the anomaly detection system 102 may decompose the metric time series into a periodic component sequence, a level component sequence, a spike component sequence, and an error component sequence. Respectively as Figure 2C 、 2D, as shown in FIGS. 2E and 2F, the anomaly detection system 102 can generate a periodic component sequence 218, a horizontal component sequence 220, a spike component sequence 222, and an error component sequence 230.
[0060] For example, as Figure 2C shown, the periodic component sequence 218 can include periodically varying values that contribute to one or more index values in the corresponding metric time series. For example, the periodic component sequence 218 can represent a region 224 of the metric time series 216, and the shape of the waveform of this region 224 is similar to that of the periodic component sequence 218. As Figure 2D further shown, the horizontal component sequence 220 can include a piecewise increase corresponding to the average value of the metric time series. For example, the increase in the average value in the area 226 of the metric time series 216 can be represented by the increase shown in the horizontal component sequence 220. As Figure 2E further shown, the spike component sequence 222 can include non-repeating temporary changes in the metric values of the metric time series. For example, the spikes shown in the spike component sequence 222 can represent data points 228a, 228b, 228c, and 228d in the metric time series 216. As Figure 2F further shown, the error component sequence 230 can include statistical noise, variance, or other residual latent components other than the interpretable latent components, such as noise or variance not represented by the periodic, horizontal, and spike components shown in the above sequences 218 - 222.
[0061] As described above, in some embodiments, the anomaly detection system 102 improves the existing anomaly detection system by subjecting the optimization algorithm to intelligent latent component constraints for decomposing the metric time series into latent components that sum to the time series. For example, in some embodiments, by intelligently subjecting real values to latent component constraints and intelligently excluding non-real values from the latent component constraints, the anomaly detection system 102 avoids misidentifying anomalies from non-real values. Therefore, the resulting latent component sequence can provide higher accuracy in significant anomaly detection. According to one or more embodiments, Figure 3 shows an overview of the anomaly detection system 102 that determines the latent components of the metric time series and identifies significant anomalies based on the latent components.
[0062] Specifically, Figure 3 shows that the anomaly detection system 102 retrieves the metric time series 302. In one or more embodiments, the anomaly detection system 102 can retrieve the metric time series in response to receiving a query from the administrator computing device 108. For example, the anomaly detection system 102 can receive a query associated with the metric time series from the administrator computing device 108 and identify the corresponding metric time series from the analysis database 120 (e.g., as Figure 1as shown). Alternatively, in response to receiving a query, the anomaly detection system 102 may request a metric time series from a third-party web server 112.
[0063] In one or more embodiments, the anomaly detection system 102 may also determine potential components 304 of the metric time series. For example, the anomaly detection system 102 may determine at least one level component sequence and one spike component sequence as potential components of the metric time series by performing an optimization algorithm. In some cases, the anomaly detection system 102 performs an optimization algorithm that is subject to potential component constraints and excludes non-real values of the metric time series from the potential component constraints. By way of illustration, the anomaly detection system 102 configures the optimization algorithm to exclude any non-real values (e.g., NA, NaN) from the potential component constraints while decomposing the metric time series into one or more potential components. To further ensure that INF values (e.g., infinity values) in the metric time series are subsequently identified as anomalies, the anomaly detection system 102 may replace the spike component sequence with the metric time series in the objective function of the optimization algorithm. With these configurations of the optimization algorithm, the anomaly detection system 102 can accurately decompose the metric time series into at least one spike component sequence and one level component sequence even if the metric time series includes values representing non-real numbers.
[0064] As Figure 3 Further shown, the anomaly detection system 102 may identify statistically significant anomalies 306 from the spike component sequence and the level component sequence. In one or more embodiments, the anomaly detection system 102 may simultaneously apply different statistical analyses to one or more potential components of the metric time series to identify significant values of the potential components. The anomaly detection system 102 may then use anomaly detection to determine whether any significant values are still anomalous.
[0065] For example, the anomaly detection system 102 may identify a significant subsequence of the spike component sequence by determining whether one or more values in the spike component sequence meet a spike significance threshold. In one or more embodiments, for example, the anomaly detection system 102 generates a stationary time series equal to the combination of the spike component sequence and the residual error value. The anomaly detection system 102 also determines that the spike component sequence values meet the spike significance threshold by determining that the stationary time series deviates from a dataset that conforms to a distribution (e.g., a dataset of similar time series that conforms to a standardized distribution). In another embodiment, the anomaly detection system 102 may use other statistical methods to determine whether a subsequence of the spike component sequence meets the spike significance threshold such that the values of the subsequence are statistically significant.
[0066] The anomaly detection system 102 can also identify significant level changes corresponding to the horizontal component sequence. For example, the anomaly detection system 102 can determine that a horizontal change corresponding to the horizontal component sequence is significant by determining whether the horizontal change meets a horizontal change significance threshold. In one or more embodiments, the anomaly detection system 102 can determine that the horizontal change meets the horizontal change significance threshold by generating a significant level change value and determining that the absolute value of the horizontal change corresponding to the horizontal component sequence exceeds or is equal to the significant level change value. In such an example, the horizontal change significance threshold can be based on the critical value of a Gaussian distribution of a predetermined significance level.
[0067] As Figure 3 Further shown, the anomaly detection system 102 can generate anomaly detection algorithm results 308 indicating anomalies in the metric time series on a graphical user interface. For example, the anomaly detection system 102 can identify outlier values that represent a significant deviation from the expected trend associated with the metric time series. In one or more embodiments, the anomaly detection system 102 can also generate anomaly detection algorithm results including an indication of the outlier values to be displayed on a client computing device (e.g., such as Figure 1 the client computing device 108a shown) together with the metric time series 308 and its various potential components. For example, the anomaly detection system 102 can generate a trend of the metric time series with an indication of an anomaly spike in the metric time series. The anomaly detection system 102 can also generate anomaly detection algorithm results on the graphical user interface to include trends associated with the potential components.
[0068] Figure 4 Additional details of the process by which the anomaly detection system 102 determines significant anomalies based on the potential components of the metric time series are shown. For example, as Figure 4 shown, the anomaly detection system 102 can retrieve the metric time series 402. As described above, the anomaly detection system 102 can retrieve the metric time series in response to a query initiated by the administrator computing device 108. For example, the anomaly detection system 102 can receive, via the data analysis application 110, a query from the administrator computing device 108 for a metric time series associated with network user actions within a given time range. In response to receiving the query, the anomaly detection system 102 can retrieve metric time series information corresponding to the query from a third-party network server 112. In at least one embodiment, the metric time series information can be restricted to a specified time range. In other embodiments, in addition to other metadata (e.g., user information, geographic information), the metric time series information can also include other data outside the specified time range.
[0069] Before decomposing a metric time series into its latent components, in one or more embodiments, the anomaly detection system 102 may modify the metric time series to improve the accuracy of anomaly detection. For example, the anomaly detection system 102 may modify the metric time series 404 by identifying leading and / or trailing zeros in the metric time series. For example, in some embodiments, a monthly metric time series may be associated with an online marketing campaign related to a particular web page that begins in the middle of the month. The metric time series may only contain non-zero values for the last half of the month, and the initial indices of the metric time series may be filled with zeros because the online marketing campaign has not started on the days associated with these initial indices. If included in the anomaly detection, these leading zero values may result in inaccurate identification of anomalies associated with the metric time series. Thus, in response to identifying leading and / or trailing zeros in the metric time series (e.g., "yes" in 404), the anomaly detection system 102 may remove the identified leading and / or trailing zero values from the metric time series 406.
[0070] Additionally, the anomaly detection system 102 may improve the accuracy of the anomaly detection process by determining whether the number of values in the metric time series is less than a threshold number 408. For example, in one or more embodiments, when the number of values in the metric time series is below a threshold number (e.g., 15 events), an optimization algorithm that decomposes the metric time series into its latent components cannot accurately identify anomalies. In at least one embodiment, in response to determining that the number of values in the metric time series is equal to or less than the threshold number and that each value in the metric time series is non-negative (e.g., "yes" in 408), the anomaly detection system 102 may trim or reduce the confidence interval for the error constraint of the optimization algorithm 410. For example, as will be further discussed below, for the error constraint ||e||2 ≤ ρ, the anomaly detection system 102 may reduce the parameter ρ. In at least one embodiment, the practical effect of reducing the parameter ρ is to reduce the potential number of anomalies detected from the metric time series.
[0071] To decompose a metric time series into one or more latent components, the anomaly detection system 102 may perform an optimization algorithm that obeys latent component constraints. For example, as part of this performance, the anomaly detection system 102 may apply the latent component constraints to the values and may apply an objective function that is part of the optimization algorithm.
[0072] For example, the analysis system may configure the following optimization algorithm (1):
[0073]
[0074] Subject to y = s + t + d + e
[0075] ||e||2 ≤ ρ
[0076] As shown above, the objective function is: min {s,t,d,e} ||Fs||1 + w1||Δt||1 + w2||d||1. The strict latent component constraint is: subject to y = s + t + d + e. The error constraint is: ||e||2 ≤ ρ. Here, s is the periodic component sequence, t is the trend component sequence, d is the spike component sequence, e is the residual error component sequence, and the metric of the time series y and the corresponding latent component sequences are the metric values (e.g., observations) on a fixed window of finite size N (i.e., ). In the above objective function, the term F represents an N×N discrete Fourier transform matrix or other frequency transform matrix, where F multiplies the periodic component sequence s. ||Fs||1 is the periodic term calculated based on the periodic component sequence s. The term represents the first difference operator (i.e., the k th element of Δt is t(k + 1) = t(k)). The references to ||Δt||1 and ||d||1 represent the l1 norm of the trend change and the l1 norm of the spike component sequence, respectively. The parameter w1 represents the weight associated with the trend component sequence t, and the parameter w2 represents the weight associated with the spike component sequence d, and the parameter ρ represents the p-value or the statistical significance level.
[0077] The parameters w1, w2, and ρ are also parameters that can be adjusted to emphasize the various contributions of the latent components. In some embodiments, the weight w1 models or otherwise indicates the contribution of the trend component sequence to the metric time series. For example, reducing w1 indicates that the trend component sequence provides a greater contribution to the metric time series. The weight w2 models or otherwise indicates the contribution of the spike component sequence to the metric time series. For example, reducing w2 indicates that the spike component sequence provides a greater contribution to the metric time series.
[0078] As described above, the optimization algorithm (1) minimizes the objective function subject to the latent component constraint. Generally, the objective function is a convex function that includes the sum of various terms representing the latent components of the metric time series y. Thus, in multiple iterations, the optimization algorithm (1) minimizes the objective function such that all values of the resulting latent components satisfy the latent component constraint. If a value is missing in the metric time series y in a particular iteration, then according to the objective function, the optimization algorithm (1) selects an alternative value from Fs.
[0079] Therefore, the optimization algorithm (1) uses the l1 norm to enhance the sparsity of different potential components of the metric time series y, where the l1 norm is the absolute value of the relevant terms. By using the l1 norm of the vectors representing various potential components, the analysis system can identify the sparsely distributed values in the potential component sequence. For example, the periodic term Fs represents the frequency transformation of the periodic component sequence s. The frequency transformation transforms the periodic component sequence from the time domain to the frequency domain. The discrete Fourier transform ("DFT") is an example of Fs. In the optimization algorithm (1), by using the l1 norm of this frequency transformation (||Fs||1), a sparse representation of the periodic component sequence s in the discrete component Fourier domain or other frequency domains can be encouraged. By transforming the periodic component sequence from the time domain to the frequency domain, the optimization algorithm (1) is well-suited for representing periodic signals.
[0080] The above optimization algorithm (1) also uses the following assumption: the level component sequence t is piecewise constant. This assumption enables the shift in the average level of the metric values in the metric time series to be more accurately captured in the level component sequence. The piecewise constant assumption of the level component sequence t balances the simplicity of the model and overfitting to the data. From a theoretical perspective, the piecewise constant function considers all index values in the metric time series, so generality is not lost under this assumption. For practical considerations, slowly varying level values can be accurately represented by a series of infrequent level shifts or piecewise constant signals.
[0081] In the above optimization algorithm (1), the assumption that the level component sequence t is piecewise constant is achieved by using the level term ||Δt||1, where the level term ||Δt||1 is the l1 norm of the vector filled with the differences between adjacent level value pairs in the level component sequence t (i.e., Δt k = t(k + 1) - t(k)). In this example, for a slowly varying piecewise constant level component sequence t, (t(k + 1) - t(k)) is expected to be non-zero for very few values of k ∈ {0,..., N - 1}. Including the level term ||Δt||1 in the optimization algorithm encourages sparsity in the level component sequence t.
[0082] In this example, it is assumed that the spikes in the spike component sequence d occur infrequently, resulting in the spike component sequence d being sparse in the time domain. This assumption can be achieved by calculating the spike term as the l1 norm of the spike component sequence d.
[0083] The anomaly detection system 102 can calculate the error component sequence e according to the objective function by subtracting other potential component sequences s, t, and d from the metric time series y. The error component sequence e captures the noise in y and models the fitting error. In the optimization algorithm (1), by using the error constraint ||e||2 ≤ ρ as the upper limit of the energy of the error component sequence e, the influence of the error component sequence e can be controlled. To account for negative values, in some embodiments, the anomaly detection system 102 squares the error constraint as follows: Although the optimization algorithm (1) discussed above is a convex optimization problem, the anomaly detection system 102 can decompose the metric time series y into its potential components in any suitable manner.
[0084] Although the above optimization algorithm (1) decomposes the metric time series into its potential components, the optimization algorithm (1) fails to consider the non-real values of the metric time series y. Thus, as Figure 4 shown, instead of using the optimization algorithm (1), the anomaly detection system 102 executes the optimization algorithm (2) to decompose the metric time series into potential components 412. As shown, the optimization algorithm (2) is a modification of the optimization algorithm (1) that overcomes various drawbacks. For example, during execution, the anomaly detection system 102 can apply the potential component constraint to the real values 414, as Figure 4 and shown in the following optimization algorithm (2). As Figure 4 further shown and further described below, the anomaly detection system 102 can apply the objective function of the optimization algorithm (2) to handle the infinite values 416 in the metric time series.
[0085] As just noted, in one or more embodiments, the anomaly detection system 102 executes the optimization algorithm (2) to decompose the metric time series into potential components as follows:
[0086]
[0087] subject to y[ind] = s[ind] + t[ind] + d[ind] + e[ind]
[0088] ||e||2 ≤ ρ
[0089] As shown above, the optimization algorithm (2) uses the same objective function as the objective function described above for the optimization algorithm (1). Contrary to the optimization algorithm (1), the anomaly detection system 102 is configured for the latent component constraints of the optimization algorithm (2) (e.g., subject to y[ind] = s[ind] + t[ind] + d[ind] + e[ind]), such that the constraints are applied only to the real-valued indices of y. In this way, the optimization algorithm (2) will select Δt = d[i] = e[i] = 0 corresponding to the non-real values of y. Then, the anomaly detection system 102 calculates the periodic component sequence for the non-real-valued indices of y from the frequency domain. By configuring the optimization algorithm (2) in this way, after execution, the optimization algorithm (2) will replace the non-real values with the expected values taking into account the periodicity.
[0090] By executing the optimization algorithm (2), the anomaly detection system 102 appropriately takes into account the metric time series values including unavailable values and non-numeric values. However, the optimization algorithm (2) cannot correctly solve the problem of infinite values. Thus, the anomaly detection system 102 can also modify and apply the objective function of the optimization algorithm (2) to handle the infinite values 416 in the metric time series. For example, the anomaly detection system 102 can allow the infinite values in the metric time series y by replacing the spike component sequence d with the metric time series y in the objective function of the above optimization algorithm (2) (e.g., min {s,t,d,e} ||Fs||1 + w1||Δt||1 + w2||d||1). Thus, the objective function becomes: min {s,t,d,e} ||Fs||1 + w1||Δt||1 + w2||y||1. In at least one embodiment, this reconfiguration ensures that the infinite values in the metric time series are replaced with real numbers large enough to ensure that they are identified as significant anomaly spikes.
[0091] In one or more embodiments, the anomaly detection system 102 can execute the optimization algorithm (2) by constraining the metric time series to include the sum of the latent components. As described above, the latent component constraints of the optimization algorithm (e.g., subject to y[ind] = s[ind] + t[ind] + d[ind] + e[ind]) specify that the metric time series is the sum of various latent components. Here, the metric time series is the sum of the periodic component sequence, the level component sequence, the spike component sequence, and the error component sequence. In at least one embodiment, the anomaly detection system 102 can utilize the resulting error component sequence e to generate upper and lower bounds for the corresponding indices in the other latent component sequences for identifying anomalous data (e.g., index values outside the generated range).
[0092] By performing an optimization algorithm (2) that obeys latent component constraints, the anomaly detection system 102 can decompose the metric time series y into latent component sequences s, t, d, and e. In additional or alternative embodiments, the anomaly detection system 102 can perform an optimization algorithm to identify fewer or more latent component sequences of the metric time series y. Additional information regarding how the anomaly detection system 102 detects anomalies in the latent components of a metric time series is described by Shiv Kumar Saini, Sunav Choudhary, and Gaurush Hiranandani in U.S. Patent Application No. 15 / 804,012, filed November 6, 2017, "Extracting Seasonal, Level, and Spike Components from a Time Series of Metrics Data," the entire content of which is incorporated herein by reference.
[0093] After the latent component sequences of the metric time series have been identified, the anomaly detection system 102 can identify significant anomaly data values based on the latent component sequences. As Figure 4 shown, for example, the anomaly detection system 102 can identify significant spikes / dips and level changes 418 in the latent components. As described above, in some embodiments, the anomaly detection system 102 improves upon conventional systems by detecting anomaly data values in one or more latent components of the metric time series and determining whether the identified anomalies are statistically significantly different from a threshold or expected value.
[0094] In one or more embodiments, the anomaly detection system 102 identifies significant anomaly data values of the latent component sequences based on a significance threshold (e.g., as shown in action 418). For example, the anomaly detection system 102 can identify significant anomaly data values of the spike component sequence by determining whether a subsequence of the spike component sequence meets a spike significance threshold. In at least one embodiment, the anomaly detection system 102 generates a stationary time series equal to the combination of the spike component sequence and the residual error value. Then, the anomaly detection system 102 determines that the spike component sequence values meet the spike significance threshold by determining that the spike component sequence values deviate from a dataset that conforms to a standardized distribution.
[0095] For illustration, in one or more embodiments and to identify statistically significant spikes in the spike component sequence d, the anomaly detection system 102 first removes any serial correlation from the errors within the metric time series. For example, let e t = α0 + α1*e t-1 + μ t , where e tRepresents the error within the metric time series. The AR1 regression of the estimated error eliminates the serial correlation from the error. The anomaly detection system 102 can add the residual error value μ to the spike component sequence d to obtain a stationary time series: y s = d + μ. This stationary time series y s has no periodicity and level changes, and the d and μ sequences are uncorrelated. Therefore, the stationary time series y s is suitable for further statistical analysis because any small spikes introduced by the error e t are removed or otherwise resolved.
[0096] In at least one embodiment, the anomaly detection system 102 utilizes a Generalized Extreme Studentized Distribution (GESD) test to identify significant spikes and dips in the stationary time series y s . For example, the GESD test detects one or more outliers in a univariate dataset that conforms to an approximate standardized distribution. The anomaly detection system 102 can improve the accuracy of the GESD test by using the number of values in the spike component sequence as the maximum number of possible outliers required for the GESD test. Therefore, the anomaly detection system 102 can identify the outliers detected by the GESD test as significant spikes (e.g., significant values) in the stationary time series y s . In one or more embodiments, the anomaly detection system 102 can also associate the identified significant spikes with the corresponding indices of the spike component sequence to determine that a subsequence of the spike component sequence represents a statistically significant anomaly.
[0097] As described above, the anomaly detection system 102 can identify significant anomaly data values in the level component sequence by determining whether the level change corresponding to the level component sequence satisfies a level change significance threshold (e.g., as shown in action 418). For example, in at least one embodiment, when the level change represents a deviation from a significant level change value, the anomaly detection system 102 can determine that the level change satisfies the level change significance threshold. To illustrate, if then the anomaly detection system 102 can determine that the represented level change between two consecutive indices of the level component sequence is significant, where is the critical value of the Gaussian distribution at a significance level of α. If the represented level change between two consecutive indices of the level component sequence satisfies this threshold, the anomaly detection system 102 identifies the level change as significant.
[0098] As Figure 4As further shown, the anomaly detection system 102 can alternatively modify one or more of the metric time series and / or the latent component series to further improve the accuracy of anomaly detection. In particular, the anomaly detection system 102 determines whether there are matching previous spikes or dips 420 in a similar time period. For example, the anomaly detection system 102 can identify data values of the spike component series corresponding to the time period, and identify matching previous data values of the previous spike component series corresponding to a similar time period. Such previous data values of the previous spike component series can correspond or match the data values of the spike component series in terms of index, time, or position within a time window. In response to identifying these matching data values from a similar time period, in some embodiments, the anomaly detection system 102 does not identify the current spike / dip or the data values from the current spike component series as an anomaly 422 to be adjusted due to a periodic effect. In one or more embodiments, the anomaly detection system 102 can identify and address similar periodic effects in any latent component series such as the level component series. In at least one embodiment, the anomaly detection system 102 can identify and address the periodic effects in one or more latent component series jointly or simultaneously.
[0099] More specifically, the anomaly detection system 102 can determine that one or more data values from the spike component series d corresponding to a specific time effect or special day effect do not represent a significant anomaly (e.g., as in action 420). In one or more embodiments, the metric time series can include significant spikes and / or dips that repeat at regular intervals due to periodic events such as holidays, weekends, regular promotions, etc. In at least one embodiment, the anomaly detection system 102 can avoid identifying these regular spikes and / or dips as anomalies by determining whether a significant spike or dip in the current time period within the spike component series d corresponds to a significant spike or dip in a past similar time period.
[0100] For further illustration, for each significantly valued index identified in the current spike component sequence, the anomaly detection system 102 can identify the value represented by the same index in the past spike component sequences. If the value at the same index in the past spike component sequences is the same as that in the current spike component sequence, the anomaly detection system 102 can determine that there is a special day or time effect at the date and / or time represented by the index. In response to determining that there is a special day or time effect at that index (e.g., "yes" at 420), the anomaly detection system 102 can determine not to identify the index and its value in the current spike component sequence 422 as an anomaly. In one or more embodiments, the anomaly detection system 102 can identify special day or time effects on a year-by-year basis, month-by-month basis, week-by-week basis, etc.
[0101] After identifying significantly abnormal data values based on the latent components of the metric time series, the anomaly detection system 102 can generate the abnormal data values for display 424. For example, the anomaly detection system 102 can generate one or more interactive graphical user interfaces that include a visual representation of the abnormal data values related to the metric time series and its latent component sequences. In one or more embodiments, the anomaly detection system 102 can provide the generated abnormal data values for display on a client computing device (e.g., an administrator computing device 108 such as Figure 1 shown).
[0102] Figures 5A - 5D shows a graphical user interface that intuitively indicates the abnormal data associated with the selected metric time series. For example, as Figure 5A shown, the anomaly detection system 102 can generate a graphical user interface 502 for display on the administrator computing device 108. In one or more embodiments, the anomaly detection system 102 can generate a graphical user interface 502 having metric time series elements 504a, 504b, and 504c. For example, the anomaly detection system 102 can generate the metric time series elements 504a, 504b, and 504c in response to receiving a query from the administrator computing device 108 specifying an application or website, one or more types of user interactions associated with the application or website, target demographic information, and / or a time period. Thus, in response to detecting a selection of the metric time series element 504a, the anomaly detection system 102 can generate a metric time series 506 corresponding to the metric time series associated with the selected element.
[0103] As Figure 5AAs shown, the graphical user interface 502 may also include an analysis button 508. In one or more embodiments, in response to a detected selection of the analysis button 508, the anomaly detection system 102 may determine the potential components of the selected metric time series and identify significant anomalies based on the potential components. The data analysis system 104 may also generate significant anomalies for display on the administrator computing device 108 within the graphical user interface 502.
[0104] To further illustrate some advantages of the anomaly detection system 102, Figure 5B a display output of an existing analysis computing system is shown. For example, when performing an analysis of a metric time series associated with the Figure 5B metric time series 506 shown, the existing analysis computing system will identify the data as a training period at an hourly granularity, with a confidence level of 99, and approximately 336 points (e.g., half of the metric time series). The existing system analyzes the metric time series associated with the metric time series 506 to identify a fitted data sequence 507a, an upper bound 507b, and a lower bound 507c, as Figure 5B shown. Additionally, the existing system may identify multiple anomalies associated with the metric time series, which may or may not include non-real values. As Figure 5B shown, the existing system identifies a large number of anomalies that may not be significant. Moreover, the existing system may also have identified a large number of anomalies due to using the first half of the data values in the time series as the training period, even though these data values are not relevant to the second half of the data values in the time series, as indicated by an undetected level change (e.g., near the middle of the metric time series 506). In some cases, the existing system identifies the Figure 5B anomalies shown by performing the optimization algorithm (1) described above.
[0105] In contrast to the existing analysis computing system, the anomaly detection system 102 identifies significant anomalies from the metric time series. For example, as Figure 5C shown, the anomaly detection system 102 may identify the metric time series 506, the fitted data sequence 507a, the upper bound 507b, and the lower bound 507c. In some cases, the anomaly detection system 102 identifies the Figure 5C anomalies shown by performing the optimization algorithm (2) described above. Additionally, the anomaly detection system 102 may determine significant anomalies 510a, 510b, 510c, 510d, and 510e and a significant level change 512. As shown, the significant anomalies 510a, 510b, 510c, 510d, and 510e and the significant level change 512 represent Figure 5B fewer anomaly data values compared to those identified by the existing system shown in Figure 5C thereby providing more meaningful insights to the analyst.
[0106] In one or more embodiments, as Figure 5D shown, the anomaly detection system 102 can also generate trends associated with the latent components of the selected metric time series. For example, the anomaly detection system 102 can use the optimization algorithm (2) described above to generate trends for the level component sequence, spike component sequence, periodic component sequence, and error component sequence. The anomaly detection system 102 can display each of these trends overlaid on the metric time series 506, or can display each trend individually. By displaying the latent component trends associated with the metric time series 506, the anomaly detection system 102 provides an effective analysis tool that enables an administrator to quickly view significant spikes and level changes relative to the selected metric time series.
[0107] As regarding Figures 1 - 5D described, the anomaly detection system 102 performs operations for identifying significant anomalies in the latent components of the metric time series. Figure 6 A detailed schematic diagram of an embodiment of the anomaly detection system 102 described above is shown. Although shown above as being on the server(s) 106, the anomaly detection system 102 can be implemented by one or more different or additional computing devices (e.g., the administrator computing device 108). In one or more embodiments, the anomaly detection system 102 includes a decomposer 602, a significant value analyzer 612, a display manager 618, a sequence modifier 622, and an anomaly detector 630.
[0108] As described above, and as Figure 6 shown, the anomaly detection system 102 includes a decomposer 602. In one or more embodiments, the decomposer 602 includes an optimization algorithm 604 (e.g., optimization algorithm (2), which is a modification of optimization algorithm (1)), the optimization algorithm 604 in turn includes an objective function 606 (e.g., min {s,t,d,e} ||Fs||1 + w1||Δt||1 + w2||d||1), latent component constraints 608 (e.g., y[ind] = s[ind] + t[ind] + d[ind] + e[ind]), and error constraints 610 (e.g., ||e||2 ≤ ρ). As described above, the anomaly detection system 102 can modify, reconfigure, or execute one or more components of the decomposer 602 such that the optimization algorithm 604 can account for non-real values in the metric time series, as well as low event counts associated with the metric time series (e.g., the metric time series retrieved from the analysis database 120).
[0109] As described above, and as Figure 6As shown, the anomaly detection system 102 includes a significant value analyzer 612. In one or more embodiments, the significant value analyzer 612 can include a spike significance manager 614 and a level significance manager 616. For example, the spike significance manager 614 can identify significant values in a spike component sequence by determining that a value meets a spike significance threshold. Similarly, the level significance manager 616 can identify significant level changes corresponding to a level component sequence by determining that a level change meets a level change significance threshold.
[0110] As described above, and as Figure 6 shown, the anomaly detection system 102 includes a sequence modifier 622. In one or more embodiments, the sequence modifier 622 includes a special day manager 624, a leading / trailing zero manager 626, and a count data manager 628. The special day manager 624 can identify special days or time effects associated with a metric time series and determine not to label such values that reflect the special days or time effects as anomalies. The leading / trailing zero manager 626 can identify and remove leading and / or trailing zeros in the metric time series. The count data manager 628 can reduce the confidence interval of the error constraint 610 based on determining that the number of values in the metric time series is less than or equal to a threshold number.
[0111] As described above, and as Figure 6 shown, the anomaly detection system 102 includes a display manager 618, and the display manager 618 includes a display generator 620. In one or more embodiments, the display generator 620 configures a graphical user interface using metric time series data, latent component data, and anomaly data. The display generator 620 can also provide the graphical user interface to an administrator computing device for display and further interaction. In response to detected interactions with the graphical user interface, the display generator 620 can update the graphical user interface with additional or different data.
[0112] As described above, and as Figure 6 shown, the anomaly detection system 102 includes an anomaly detector 630. In one or more embodiments, the anomaly detector 630 includes functionality for simultaneously determining whether the values of one or more latent components of a metric time series represent anomalies. For example, the anomaly detector 630 can analyze a spike component sequence to identify one or more anomalous spikes and provide these identified anomalous spikes to the significant value analyzer 612 to further determine whether any of the anomalous spikes are significant in the metric time series.
[0113] Each component 602 - 630 of the anomaly detection system 102 can include software, hardware, or both. For example, components 602 - 630 can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices such as a client device or a server device. When executed by one or more processors, the computer-executable instructions of the anomaly detection system 102 can cause the computing device(s) to perform the methods described herein. Alternatively, components 602 - 630 can include hardware, such as a dedicated processing device for performing a particular function or set of functions. Alternatively, the components 602 - 630 of the anomaly detection system 102 can include a combination of computer-executable instructions and hardware.
[0114] In addition, the components 602 - 630 of the anomaly detection system 102 can be implemented, for example, as one or more operating systems, one or more standalone applications, one or more modules of an application, one or more plug-ins, one or more library functions or functions that can be called by other applications, and / or a cloud computing model. Thus, the components 602 - 630 can be implemented as standalone applications, such as desktop or mobile applications. In addition, the components 602 - 630 can be implemented as one or more web-based applications hosted on a remote server. The components 602 - 630 can also be implemented in a set of mobile device applications or “apps”. By way of illustration, the components 602 - 630 can be implemented in applications including, but not limited to, ADOBE ANALYTICS CLOUD, such as ADOBE ANALYTICS, ADOBE AUDIENCE MANAGER, ADOBE CAMPAIGN, ADOBE EXPERIENCE MANAGER, and ADOBE TARGET. “ADOBE”, “ANALYTICS CLOUD”, “ANALYTICS”, “AUDIENCE MANAGER”, “CAMPAIGN”, “EXPERIENCE MANAGER”, “TARGET”, and “CREATIVE CLOUD” are registered trademarks or trademarks of Adobe Systems Incorporated in the United States and / or other countries / regions.
[0115] Reference Figures 1 - 6 and the corresponding text and examples provide various different methods, systems, devices, and non-transitory computer-readable media of the anomaly detection system 102. In addition to the foregoing, one or more embodiments can also be described according to flowcharts including actions for achieving a particular result, as Figure 7 and 8 shown. Figure 7 and 8It can be performed with more or fewer actions. Additionally, the actions can be performed in a different order. Further, the actions described herein can be repeated, or performed in parallel with each other or with different instances of the same or similar actions.
[0116] As previously described, Figure 7 FIG. 700 is a flow diagram of an action sequence for identifying significant anomalies of one or more potential components of a metric time series according to one or more embodiments. Although Figure 7 FIG. 7 shows actions according to one embodiment, alternative embodiments can omit, add, reorder, and / or modify Figure 7 any of the actions shown. Figure 7 The actions of can be performed as part of a method. Alternatively, a non-transitory computer-readable medium can include instructions that, when executed by one or more processors, cause a computing device to perform Figure 7 the actions of. In some embodiments, a system can perform Figure 7 the actions of.
[0117] As Figure 7 shown, action sequence 700 includes an action 710 of retrieving a metric time series. For example, action 710 can involve retrieving a metric time series corresponding to a time period, the metric time series including metric data values representing user actions within a digital network.
[0118] As Figure 7 further shown, action sequence 700 includes an action 720 of determining a spike component sequence and a level component sequence associated with the metric time series. For example, action 720 can involve determining at least one spike component sequence and one level component sequence as potential components of the metric time series. In one or more embodiments, determining at least one spike component sequence and one level component sequence as potential components of the metric time series can include: applying an optimization algorithm that obeys potential component constraints for real values to the metric time series, the potential component constraints indicating that the metric time series includes at least a spike component sequence and a level component sequence; and based on the optimization algorithm, identifying the spike component sequence and the level component sequence as potential components of the metric time series. In at least one embodiment, applying the optimization algorithm to non-real values can include: excluding non-real values from the potential component constraints; and determining a periodic component sequence for non-real values based on the frequency domain.
[0119] As Figure 7Further shown, the action sequence 700 includes an action 730 of determining whether the spike component sequence and the horizontal component sequence include significantly abnormal data values. For example, the action 730 may involve determining whether a subsequence of the spike component sequence meets a spike significance threshold and whether a horizontal change corresponding to the horizontal component sequence meets a horizontal change significance threshold. In one or more embodiments, determining whether a subsequence of the spike component sequence meets the spike significance threshold may include: generating a stationary time series equal to the combination of the spike component sequence and the residual error value; making the stationary time series follow a dataset conforming to a standardized distribution; and determining data values of the stationary time series that deviate from the dataset conforming to the standardized distribution. Additionally, in one or more embodiments, determining that the horizontal change corresponding to the horizontal component sequence meets the horizontal change significance threshold may include: generating a significant horizontal change value; and determining that the absolute value of the horizontal change corresponding to the horizontal component sequence exceeds or is equal to the significant horizontal change value.
[0120] As Figure 7 Further shown, the action sequence 700 includes an action 740 of generating abnormal data values to be displayed on a client computing device based on one or more of the spike component sequence or the horizontal component sequence including multiple significantly abnormal data values. For example, the action 740 may involve generating abnormal data values from the metric time series to be displayed on a client computing device based on one or more of a subsequence of the spike component sequence meeting the spike significance threshold or the horizontal change corresponding to the horizontal component sequence meeting the horizontal change significance threshold.
[0121] Furthermore, in one or more embodiments, the action sequence 700 includes an action of applying the objective function of an optimization algorithm to infinite values in the metric time series by replacing the spike component sequence with the metric time series. Additionally, in at least one embodiment, the action sequence 700 includes an action of constraining the metric time series for real values to include the sum of the spike component sequence, the horizontal component sequence, the periodic component sequence, and the error component sequence.
[0122] In one or more embodiments, the action sequence 700 includes the following actions: identifying data values of the spike component sequence corresponding to a time period and previous data values of a previous spike component sequence corresponding to a similar time period; and determining that the data values of the spike component sequence do not represent an anomaly requiring adjustment due to a periodic effect. In at least one embodiment, the action sequence 700 further includes the following actions: at least determining the spike component sequence and the horizontal component sequence as potential components of the metric time series by applying an optimization algorithm to the metric time series without dividing the metric time series into data values for a training period and data values for a testing period.
[0123] Figure 8A flowchart of an action sequence 800 for simultaneously determining a subsequence of a spike component sequence and whether a horizontal change corresponding to a horizontal component sequence represents a significant anomaly according to one or more embodiments is shown. Although Figure 8 shows an action according to one embodiment, alternative embodiments may omit, add, reorder, and / or modify Figure 8 any of the actions shown. Figure 8 The actions of can be performed as part of a method. Alternatively, a non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause a computing device to perform Figure 8 the actions of. In some embodiments, the system may perform Figure 8 the actions of.
[0124] As Figure 8 shown, the action sequence 800 includes an action 810 of determining a spike component sequence and a horizontal component sequence as potential components of a metric time series. For example, the action 810 may involve: determining at least one horizontal component sequence and one spike component sequence as potential components of a metric time series subject to potential component constraints, and excluding non-real values of the metric time series from the potential component constraints. In one or more embodiments, determining at least one horizontal component sequence and one spike component sequence as potential components of a metric time series may include: applying an optimization algorithm subject to potential component constraints for real values to the metric time series, the potential component constraints indicating that the metric time series includes a spike component sequence, a horizontal component sequence, a periodic component sequence, and an error component sequence; and then, based on the optimization algorithm, identifying the spike component sequence, the horizontal component sequence, the periodic component sequence, and the error component sequence as potential components of the metric time series.
[0125] In at least one embodiment, the action sequence 800 further includes the following actions: applying an optimization algorithm subject to error constraints to the metric time series, the error constraints restricting the error component sequence to a confidence interval; determining that the number of metric data values within the metric time series is equal to or less than a threshold number of data values, and each metric data value is non-negative; and based on determining that the number of metric data values is equal to or less than the threshold number of data values, reducing the confidence interval of the error constraints.
[0126] Additionally, in at least one embodiment, the action sequence 800 includes the following actions: excluding unavailable values or non-numeric values from potential component constraints; and determining a periodic component sequence for non-real values based on the frequency domain. The action sequence 800 may also include an action of removing at least one of leading zero values or trailing zero values from the metric time series. Additionally, the action sequence 800 may further include the following actions: determining at least a horizontal component sequence and a spike component sequence as potential components of the metric time series by applying an optimization algorithm to the metric time series without dividing the metric time series into data values for a training period and data values for a testing period.
[0127] As Figure 8 shown, the action sequence 800 includes an action 820 of simultaneously determining whether a subsequence of the spike component sequence and a horizontal change corresponding to the horizontal component sequence represent a significant anomaly. For example, the action 820 may include simultaneously determining whether a subsequence of the spike component sequence and a horizontal change corresponding to the horizontal component sequence represent a significant anomaly by: determining whether a stationary time series equal to the combination of the spike component sequence and the residual error value deviates from a dataset that conforms to a distribution; and determining whether the horizontal change corresponding to the horizontal component sequence deviates from a significant horizontal change value according to a horizontal change significance threshold.
[0128] In at least one embodiment, determining whether a stationary time series equal to the combination of the spike component sequence and the residual error value deviates from a dataset that conforms to a distribution may include: determining the residual error value by applying an autoregressive model to an error component sequence within the metric time series; and applying a Generalized Extreme Studentized Deviate (“GESD”) test to the stationary time series equal to the combination of the spike component sequence and the residual error value.
[0129] Additionally, in at least one embodiment, determining whether the horizontal change corresponding to the horizontal component sequence deviates from a significant horizontal change value according to the horizontal change significance threshold may include: generating a standardized distribution according to a significance level; determining a significant horizontal change value from the standardized distribution; and determining that the absolute value of the horizontal change corresponding to the horizontal component sequence exceeds or is equal to the product of the significant horizontal change value and the residual error value.
[0130] As Figure 8 shown, the action sequence 800 includes an action 830 of generating anomaly data values based on the significant anomaly for display on a client computing device. For example, the action 830 may involve generating anomaly data values from the metric time series for display on a client computing device based on one or more of a subsequence of the spike component sequence or a horizontal change indicating a significant anomaly. In at least one embodiment, the action sequence 800 may include an action of generating a graphical user interface that visually indicates anomaly data values within the metric time series.
[0131] As an alternative to the above actions, in some embodiments, the anomaly detection system 102 performs steps for identifying significant anomaly data values from a horizontal component sequence or a spike component sequence within a metric time series. In particular, the algorithms and actions described above with respect to Figure 3 or Figure 4 may include actions (or structures) corresponding to steps for identifying significant anomaly data values from a horizontal component sequence or a spike component sequence within a metric time series.
[0132] Embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer including computer hardware (e.g., one or more processors and system memory), as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein can be at least partially implemented as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory) and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0133] A computer-readable medium can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present disclosure can include at least two distinctly different types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0134] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSDs”) (e.g., based on RAM), flash memory, phase change memory (“PCM”), other types of memory, other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.
[0135] "Network" is defined as one or more data links that enable the transfer of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (wired, wireless, or a combination of wired or wireless), the computer properly views that connection as a transmission medium. The transmission medium can include networks and / or data links that can be used to carry the desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general or special purpose computer. Combinations of the foregoing should also be included within the scope of computer-readable media.
[0136] In addition, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., "NIC") and then ultimately transferred to the computer system RAM and / or to a less volatile computer storage medium (device) at the computer system. Accordingly, it should be appreciated that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize the transmission medium.
[0137] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a particular function or group of functions. In some embodiments, computer-executable instructions are executed by a general purpose computer to transform the general purpose computer into a special purpose computer implementing elements of the present disclosure. Computer-executable instructions can be, for example, binary, intermediate format instructions (such as assembly language) or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0138] Those skilled in the art should understand that the present disclosure can be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present disclosure can also be practiced in a distributed system environment, in which local and remote computer systems linked by a network (through hardwired data links, wireless data links, or a combination of hardwired and wireless data links) each perform tasks. In a distributed system environment, program modules can be located in local and remote storage devices.
[0139] Embodiments of the present disclosure can also be implemented in a cloud computing environment. As used herein, the term "cloud computing" refers to a model that enables on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be used in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be quickly configured through virtualization and released with less management effort or service provider interaction, and then scaled accordingly.
[0140] The cloud computing model can consist of various features, such as on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model can also expose various service models, such as software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). Different deployment models (such as private cloud, community cloud, public cloud, hybrid cloud, etc.) can also be used to deploy the cloud computing model. Additionally, as used herein, the term "cloud computing environment" refers to an environment in which cloud computing is adopted.
[0141] Figure 9 A block diagram of an example computing device 900 that can be configured to perform one or more of the above processes is shown. It should be understood that one or more computing devices such as computing device 900 can represent the above computing devices (e.g., (multiple) servers 106, administrator computing device 108). In one or more embodiments, computing device 900 can be a mobile device (e.g., mobile phone, smartphone, PDA, tablet computer, laptop computer, camera, tracker, watch, wearable device, etc.). In some embodiments, computing device 900 can be a non-mobile device (e.g., desktop computer or another type of client device). Additionally, computing device 900 can be a server device that includes cloud-based processing and storage capabilities.
[0142] As Figure 9As shown, computing device 900 may include one or more processors 902, a memory 904, a storage device 906, an input / output interface 908 (or "I / O interface 908"), and a communication interface 910 that may be communicatively coupled via a communication infrastructure (e.g., bus 912). Although computing device 900 is shown in Figure 9 , the components shown are not intended to be limiting. Additional or alternative components may be used in other embodiments. Further, in some embodiments, computing device 900 includes fewer components than Figure 9 shown. The components of computing device 900 shown in Figure 9 will now be described in more detail. Figure 9 The components of computing device 900 shown in
[0143] In a particular embodiment, the processor(s) 902 includes hardware for executing instructions such as those that make up a computer program. By way of example and not limitation, to execute instructions, the processor(s) 902 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 904, or the storage device 906, and decode and execute the instructions.
[0144] Computing device 900 includes a memory 904 coupled to the processor(s) 902. The memory 904 may be used to store data, metadata, and programs executed by the processor(s). The memory 904 may include one or more of volatile and non-volatile memory, such as random access memory ("RAM"), read-only memory ("ROM"), solid state disk ("SSD"), flash memory, phase change memory ("PCM"), or other types of data storage. The memory 904 may be internal or distributed memory.
[0145] Computing device 900 includes a storage device 906 that includes storage for storing data or instructions. By way of example and not limitation, the storage device 906 may include the non-transitory storage media described above. The storage device 906 may include a hard disk drive (HDD), flash memory, a universal serial bus (USB) drive, or a combination of these or other storage devices.
[0146] As shown, computing device 900 includes one or more I / O interfaces 908 that are provided to allow a user to provide input to computing device 900 (e.g., a user stroke), receive output from computing device 900, and otherwise transfer data to and from computing device 900. These I / O interfaces 908 may include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of such I / O interfaces 908. The touch screen may be activated with a stylus or a finger.
[0147] The I / O interface 908 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 908 is configured to provide graphic data to the display for presentation to the user. The graphic data may represent one or more graphical user interfaces and / or any other graphic content that may serve a particular implementation.
[0148] The computing device 900 may also include a communication interface 910. The communication interface 910 may include hardware, software, or both. The communication interface 910 provides one or more interfaces for communication (e.g., packet-based communication) between the computing device and one or more other computing devices or one or more networks. By way of example and not limitation, the communication interface 910 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as WI-FI. The computing device 900 may also include a bus 912. The bus 912 may include hardware, software, or both that connect the components of the computing device 900 to each other.
[0149] In the foregoing specification, the present invention has been described with reference to specific example embodiments of the present invention. The various embodiments and aspects of the present invention have been described with reference to the details discussed herein, and the drawings illustrate the various embodiments. The above description and drawings are illustrative of the present invention and should not be construed as limiting the present invention. Many specific details have been described to provide a thorough understanding of the various embodiments of the present invention.
[0150] Without departing from the spirit or essential characteristics of the present invention, the present invention may be embodied in other specific forms. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Additionally, the steps / actions described herein may be repeated, or performed in parallel with each other or with different instances of the same or similar steps / actions. Accordingly, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes that fall within the equivalent meaning and scope of the claims should be included within their scope.
Claims
1. A non - transitory computer - readable medium storing instructions that, when executed by at least one processor, cause a computing device to: Retrieve a metric time series corresponding to a time period, the metric time series including metric data values that represent user actions within a digital network; Determine at least one spike component sequence and a level component sequence as potential components of the metric time series by: Applying an optimization algorithm subject to potential component constraints for real values, the potential component constraints indicating that the metric time series includes at least the spike component sequence and the level component sequence; and Based on the optimization algorithm, identifying the spike component sequence and the level component sequence as potential components of the metric time series; Determine whether a subsequence of the spike component sequence satisfies a spike significance threshold and whether a level change corresponding to the level component sequence satisfies a level change significance threshold; and Generate anomaly data values from the metric time series for display on a client computing device based on one or more of: the subsequence of the spike component sequence satisfying the spike significance threshold, and the level change corresponding to the level component sequence satisfying the level change significance threshold.
2. The non - transitory computer - readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to determine that a subsequence of the spike component sequence satisfies the spike significance threshold by: Generating a stationary time series that is equal to a combination of the spike component sequence and a residual error value; and Determining that the stationary time series deviates from a dataset conforming to a standardized distribution.
3. The non - transitory computer - readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to determine that a level change corresponding to the level component sequence satisfies the level change significance threshold by: Generating a significant level change value; and Determining that an absolute value of the level change corresponding to the level component sequence exceeds or is equal to the significant level change value.
4. The non - transitory computer - readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to apply the optimization algorithm to non - real values by: Excluding non - real values from the potential component constraints; and Based on the frequency domain, determining a periodic component sequence for the non - real values.
5. The non - transitory computer - readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to apply an objective function of the optimization algorithm to infinite values from the metric time series: Replace the spike component sequence with the metric time series.
6. The non-transitory computer-readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to apply the potential component constraint for real values to the optimization algorithm by: Constraining the metric time series to include the sum of the spike component sequence, the level component sequence, the periodic component sequence, and the error component sequence.
7. The non-transitory computer-readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to: Identify data values of the spike component sequence corresponding to the time period, and previous data values of a previous spike component sequence corresponding to a similar time period; and Determine that the data values of the spike component sequence do not represent an anomaly requiring adjustment due to periodic effects.
8. The non-transitory computer-readable medium according to claim 1, further storing instructions that, when executed by the at least one processor, cause the computing device to determine at least the spike component sequence and the level component sequence as potential components of the metric time series by: Applying the optimization algorithm to the metric time series without dividing the metric time series into data values for a training period and data values for a testing period.
9. A system for generating anomaly data values, comprising: At least one memory device storing a metric time series corresponding to a time period, the metric time series including metric data values representing user actions within a digital network; and At least one computing device configured to cause the system to: Determine at least one level component sequence and one spike component sequence as latent components of the metric time series subject to latent component constraints and exclude non-real values of the metric time series from the latent component constraints; Simultaneously determine whether a subsequence of the spike component sequence and a level change corresponding to the level component sequence represent significant anomalies by: Determining whether a stationary time series equal to a combination of the spike component sequence and a residual error value deviates from a dataset conforming to a certain distribution; and Determining whether the level change corresponding to the level component sequence deviates from a significant level change value according to a level change significance threshold; and Generate anomaly data values from the metric time series for display on a client computing device based on one or more of the subsequence of the spike component sequence and the level change representing significant anomalies.
10. The system according to claim 9, wherein the at least one computing device is further configured to cause the system to generate a graphical user interface that visually indicates the anomaly data values within the metric time series.
11. The system according to claim 9, wherein the at least one computing device is further configured to cause the system to determine whether the level change corresponding to the level component sequence deviates from the significant level change value according to the level change significance threshold by: Generate a standardized distribution according to the significance level; Determine the significant level change value from the standardized distribution; and Determine that the absolute value of the horizontal change corresponding to the horizontal component sequence exceeds or is equal to the product of the significant level change value and the residual error value.
12. The system according to claim 9, wherein the at least one computing device is further configured to cause the system to determine at least the horizontal component sequence and the spike component sequence as potential components of the metric time series by: Apply an optimization algorithm subject to potential component constraints for real values, the potential component constraints indicating that the metric time series includes the spike component sequence, the horizontal component sequence, a periodic component sequence, and an error component sequence; and Based on the optimization algorithm, identify the spike component sequence, the horizontal component sequence, the periodic component sequence, and the error component sequence as potential components of the metric time series.
13. The system according to claim 12, wherein the at least one computing device is further configured to: Apply the optimization algorithm subject to error constraints to the metric time series, the error constraints restricting the error component sequence to a confidence interval; Determine that the number of metric data values within the metric time series is equal to or less than a threshold number of data values and that each metric data value is non-negative; and Based on determining that the number of metric data values is equal to or less than the threshold number of data values, reduce the confidence interval for the error constraint.
14. The system according to claim 12, wherein the at least one computing device is further configured to: Exclude unavailable values or non-numeric values from the potential component constraints; and Determine the periodic component sequence for non-real values based on the frequency domain.
15. The system according to claim 9, wherein the at least one computing device is further configured to cause the system to remove at least one of leading zero values and trailing zero values from the metric time series.
16. The system according to claim 9, wherein the at least one computing device is further configured to cause the system to determine at least the horizontal component sequence and the spike component sequence as potential components of the metric time series by: Apply an optimization algorithm to the metric time series without dividing the metric time series into data values for a training period and data values for a testing period.
17. A method for determining anomalies from a metric time series in a digital media environment for analyzing business data, comprising: Access a metric time series corresponding to a time period, the metric time series including metric data points representing user actions within a digital network; Determine at least one spike component sequence and one level component sequence as latent components of the metric time series by: Applying an optimization algorithm subject to latent component constraints for real values to the metric time series, the latent component constraints indicating that the metric time series includes at least the spike component sequence and the level component sequence; and Based on the optimization algorithm, identify the spike component sequence and the level component sequence as latent components of the metric time series; Identify a plurality of significant anomaly data values from the level component sequence or the spike component sequence within the metric time series; and Based on identifying anomaly data values from the level component sequence or the spike component sequence, generate a visual representation of the anomaly data values from the metric time series for display on a client computing device.
18. The method according to claim 17, wherein generating the visual representation of the anomalous data value for display on the client computing device comprises: Generate a graphical user interface that visually indicates the abnormal data values within one or more of the following: the metric time series, the spike component series, and the horizontal component series.
19. The method according to claim 17, further comprising applying the optimization algorithm to non-real values by: excluding non-real values from the latent component constraints; and determining a sequence of periodic components for the non-real values based on the frequency domain.
Citation Information
Patent Citations
Extracting seasonal, level, and spike components from a time series of metrics data
US10628435B2
Framework for the automated determination of classes and anomaly detection methods for time series
EP3623964A1
Anomaly Detection in Big Data Time Series Analysis
US20200183946A1