Multivariate Anomaly Detection Based on Application Telemetry
By using electronic processors to classify and predict scoring models in telemetry data for software applications, the problem of high computing resource consumption in telemetry data processing is solved, and efficient and accurate abnormality detection is achieved.
Patent Information
- Application Number
- CN201980021127.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-22
- Filing Date
- 2019-03-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-03-05
AI Technical Summary
When processing large amounts of software application telemetry data, the abnormal detection efficiency is low and the computing resource consumption is high, making it difficult to efficiently identify key abnormalities.
An electronic processor is used to receive telemetry data, classify and convert it into multi-dimensional metrics, detect abnormalities using the prediction scoring model, and generate warning messages to reduce computing resource consumption.
It realizes efficient detection of abnormalities in software applications under high volume telemetry data, saves power, memory, communication bandwidth and processing resources, and improves the accuracy and efficiency of abnormal identification.
Smart Images

Figure CN111902805B_ABST
Abstract
Description
Technical Field
[0001] The embodiments described in this application relate to anomaly detection based on application telemetry. Background Art
[0002] Generally speaking, an anomaly (also known as: outlier, noise, deviation, or exception) is an item or event that is different from what is expected. In computer science, anomaly detection refers to identifying patterns that do not conform to expectations or data, events, or conditions that do not conform to other items in a group. In some cases, encountering an anomaly may indicate handling an exception, and thus can present a starting point for investigation. Usually, anomalies are detected by humans or computational systems that learn trajectories. Trajectories include logs of information that can come from applications, processes, operating systems, hardware components, and / or networks. Summary of the Invention
[0003] A simplified overview of one or more embodiments of the present disclosure is presented below in order to provide a basic understanding of these embodiments. This overview is not an exhaustive review of all contemplated embodiments, and is not intended to identify key or important elements of all embodiments, nor to delineate the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description presented later.
[0004] With the emergence of new technologies for data collection and the adoption of agile methods, servers generate large amounts of data quickly that can be used to evaluate code quality, products, and usage. The size of the data collected for anomaly detection is typically in the petabyte (pB) range. By observing billions of data points, millions of anomalies are generated, which is typically unfeasible due to the large number of anomalies. Therefore, new anomaly detection methods and systems are needed to process data received from software application telemetry at high volume and high speed and to reduce the computational overhead when processing these anomalies.
[0005] Among other things, the embodiments described in this application provide systems and methods for collecting telemetry data associated with several software applications at relatively high frequency intervals with a relatively high level of detail. The embodiments also provide methods and systems for simultaneously detecting anomalies associated with the software applications. One example includes monitoring crashes on various different software products, for example, monitoring crashes per minute across different platforms (e.g., Windows, Mac, Linux, etc.), software applications (e.g., Microsoft Office, Word, PowerPoint, Excel, etc.), countries, languages, processors, compilation versions, audiences, screen sizes, etc., and providing anomaly detection to remedy detected anomalies.
[0006] Some embodiments detect errors at a relatively deep level of a hierarchy of errors associated with a monitored software application running on a client device. In some cases, the error hierarchy has an inverted tree-branch structure, where the path from the top to the end of any branch is referred to as a "pivot" (also referred to in this application as a "dimension"). At the deepest level of each pivot (e.g., a "leaf node"), when an anomaly or error exists, the corresponding node is marked as "1". When there is no error at a node, the node is marked as "0". Once the nodes are marked, each pivot is investigated and the errors are rolled up towards the root node while monitoring and determining the frequency of anomalies, the severity of anomalies, and the usage score of operations associated with the monitored software application. Using this process, pivots that are more anomalous than others can be identified. Identifying these pivots provides actionable insights as to whether the errors are "non-critical" or "critical". Thus, only a smaller number of anomalies need to be addressed, reducing computational resources compared to traditional anomaly detection techniques.
[0007] One example embodiment includes a computer system for anomaly detection using application telemetry. The computer system includes an electronic processor configured to receive telemetry data originating from a plurality of client applications. The telemetry data includes data points representing errors associated with one or more operations on the plurality of client applications. The electronic processor is further configured to classify the telemetry data based on multiple categories of data, convert the multiple categories of data into one or more metrics spanning multiple dimensions, aggregate the metrics for the data of the categories across the multiple dimensions, access a predictive scoring model to generate or determine a predicted error associated with a dimension of interest, detect anomalies based on an item selected from the group consisting of the predicted error and a static threshold, and output a warning message associated with the anomaly. The computer system further includes a display device for displaying the warning message.
[0008] Another example embodiment includes a method for anomaly detection using application telemetry. The method includes: receiving telemetry data originating from a plurality of client applications, the telemetry data including data points representing errors associated with one or more operations on the plurality of client applications. The method further includes classifying the telemetry data based on multiple categories of data; converting the multiple categories of data into one or more metrics based on multiple dimensions; aggregating the metrics for the data of the categories across the multiple dimensions; accessing a predictive scoring model for a predicted error associated with a dimension of interest; detecting anomalies based on an item selected from the group consisting of the predicted error and a static threshold; and outputting a warning message associated with the anomaly.
[0009] Another example embodiment includes a non-transitory computer-readable medium comprising instructions that, when executed by one or more electronic processors, cause the one or more electronic processors to perform the following actions: receive telemetry data from a plurality of client applications, the telemetry data including data points representing errors associated with one or more operations of the plurality of client applications; classify the telemetry data based on multiple categories of data; convert the multiple categories of data into one or more metrics based on multiple dimensions; aggregate the metrics for the data of the category across all the dimensions; access a predictive scoring model for stored metrics associated with a dimension of interest; determine a prediction error associated with the dimension of interest; detect an anomaly based on an item selected from the group consisting of the prediction error and a static threshold; and send a warning message, generate a problem report, and store the problem report in a database.
[0010] By using the techniques disclosed in the present application, one or more devices can be configured to conserve resources, among other things, with respect to power resources, memory resources, communication bandwidth resources, processing resources, and / or other resources, while providing automatic detection of anomalies related to online software applications that use telemetry data. Technical effects other than those mentioned in the present application can also be learned from implementations of the techniques disclosed in the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The present disclosure will be better understood from the following detailed description taken in conjunction with the accompanying drawings, in which like reference numerals are used to designate like parts in the drawings.
[0012] Figure 1 FIG. shows an example environment for implementing an anomaly detection system for a software application that utilizes telemetry data.
[0013] Figure 2 is Figure 1 a schematic diagram of an example computer architecture of the anomaly detection system.
[0014] Figure 3A is an exemplary list of anomalies at various levels and times associated with various software applications.
[0015] Figure 3B FIG. shows an example of a pivot for rolling up anomalies to perform statistical inference.
[0016] Figure 3C FIG. shows another example of a pivot for rolling up anomalies to perform statistical inference.
[0017] Figure 4 is a flowchart of a method for creating a predictive scoring model according to some embodiments.
[0018] Figure 5 is a flowchart of a method for detecting anomalies using the prediction scoring model shown in Figure 4 . DETAILED DESCRIPTION
[0019] One or more embodiments are described and illustrated in the following description and drawings. These embodiments are not limited to the specific details provided in this application and can be modified in various ways. Additionally, there may be other embodiments not described in this application. Also, the functions performed by one component described in this application can be performed by multiple components in a distributed manner. Similarly, the functions performed by multiple components can be combined and performed by a single component. Likewise, a component described as performing a specific function can also perform additional functions not described in this application. For example, a device or structure that is "configured" in a certain way is configured at least in that way, but can also be configured in ways not listed. Further, some embodiments described in this application may include one or more electronic processors configured to perform the described functions by executing instructions stored in a non-transitory computer-readable medium. Similarly, the embodiments described in this application can be implemented as a non-transitory computer-readable medium that stores instructions executable by one or more electronic processors for performing the described functions. As used in this application, "non-transitory computer-readable medium" includes all computer-readable media but does not include transitory propagated signals. Thus, non-transitory computer-readable media can include, for example, hard disks, CD-ROMs, optical storage devices, magnetic storage devices, ROM (read-only memory), RAM (random access memory), register memory, processor caches, or any combination thereof.
[0020] In addition, the words and terms used in this application are for descriptive purposes and should not be considered restrictive. For example, the use of "including", "containing", "comprising", "having" and their variations in this application is intended to cover the items listed thereafter and their equivalents as well as additional items. The terms "connected" and "coupled" are used broadly and cover direct and indirect connections and couplings. Further, "connected" and "coupled" are not limited to physical or mechanical connections or couplings and can include electrical connections or couplings, whether direct or indirect. Additionally, electronic communication and notification can be performed using wired connections, wireless connections, or a combination thereof, and can be sent directly or through one or more intermediate devices via various types of networks, communication channels, and connections. Moreover, relational terms such as first and second, top and bottom, etc. may be used in this application only to distinguish one entity or action from another entity or action, and do not necessarily require or imply any actual such relationship or order between the entities or actions.
[0021] Among other things, embodiments of the present disclosure are directed to methods and systems for automatically detecting anomalies in software applications that use telemetry data. While many of the examples shown in this application are described in terms of cloud-based applications, the configurations disclosed in this application can be implemented in a variety of ways and in different applications. More specifically, the techniques and systems described in this application can be applied to a variety of software applications that experience anomalies and can be monitored, such as those running on a platform, device, or network that can be accessed and monitored or that can provide telemetry data that can be accessed and analyzed.
[0022] Figure 1 An example environment 100 for implementing an anomaly detection system for software applications that utilize telemetry data is shown. In environment 100, a first user 102(1) and a second user 102(2) (collectively referred to as "users 102") represent multiple users 102 who are able to access and use client computing devices 104(1) and 104(2) (collectively referred to as "client devices 104") to access one or more application servers 106(1), 106(2),... 106(N) (collectively referred to as "servers 106") in a data center 108. The servers 106 store and run one or more software applications 110. The software applications 110 can include, but are not limited to: personal information management (PIM) applications such as email applications, website hosting applications, storage applications, storage services, virtual machine applications, commercial product applications, entertainment services (e.g., music services, video services, game services, etc.), personal productivity services (e.g., travel services), social networking applications, or cloud-based applications. The users 102 access the software applications 110 via a network 112.
[0023] The terms "user", "consumer", "customer", or "subscriber" may be used interchangeably in this application to refer to the users 102, and one or more users 102 may subscribe to or otherwise register to access the one or more software applications 110 as "users" of the software applications 110. In this regard, a user may include a single user 102 or a group of multiple users 102, such as when an enterprise with hundreds of employees registers as a single user of a software application 110. Thus, the data center 108 can utilize a database or similar data structure to manage the registered users of the software applications 110, including the management of access credentials for individual users 102.
[0024] The client computing device 104 (sometimes referred to in this application as "client device 104") can be implemented as any number of computing devices, including but not limited to: personal computers, laptop computers, desktop computers, portable communication devices (e.g., mobile phones or tablets), set-top boxes, game consoles, smart TVs, wearable devices (e.g., smart watches, electronic smart glasses, fitness trackers, etc.), or other electronic devices capable of sending / receiving data over the network 112. The network 112 represents many different types of networks and includes wired and / or wireless networks that enable communication between various entities in the environment 100. In some configurations, the network 112 can include the Internet, local area networks (LANs), wide area networks (WANs), mobile telephone networks (MTNs), and other types of networks that may be used in combination with each other to facilitate communication between the server 106 and the client device 104. Although some configurations are described in the context of network-based systems, other types of client / server-based communication and associated application logic may also be used.
[0025] The data center 108 can include multiple geographically distributed server clusters, where one server cluster can include a sub-grouping of the servers 106. In this way, a large number of users 102 can evaluate the software application 110 from different geographical locations. The various resources of the data center 108 can be structured in any suitable organizational framework such that they can be tracked and managed. For example, the servers 106 of the data center 108 can be organized into multiple forests 114(1), 114(2),... 114(M) (collectively referred to hereinafter as "forests 114"), where one forest 114 represents an Active Directory group of a set of users 102 that employs a subset of the servers 106. The users 102 can be widely distributed geographically. As an illustrative example, a first set of forests 114 can represent users and servers in North America (e.g., a region), while a second set of forests 114 can represent other users and servers in South America (e.g., another region), and so on. Regions can be defined at different levels of granularity, such as continents, countries, states, cities, counties, streets, etc. Within each forest 114 is a set of sites 116, which represent a lower level of user and server grouping, and where each site 116 is a set of database availability groups (DAGs). Within each DAG 118 is a set of the servers 106. By managing the data center 108 in a hierarchical architecture, it is possible to more easily identify the location of problems or anomalies that occur in the software application 110.
[0026] Figure 1It is shown that the client device 104 is configured to execute a client application 122, which is configured to access the software application 110 via a network 112. For example, the client application 122 may include an email client application 122 that is built into or downloaded to the client device 104 after manufacture and is configured to access an email-based software application 110 to allow the user 102(2) to send and receive emails. Alternatively, the client application 122 may be a web browser that allows the client device 104 to access the software application 110 in response to the user 102(2) entering a Uniform Resource Locator (URL) in the address bar of the web browser.
[0027] In addition to connecting the client device 104 to the software application 110, the client application 122 may include a telemetry data module 124, which is configured to send telemetry data 126 to one or more servers 128(1), 128(2), … 128(P) of the anomaly detection system 130. The anomaly detection system 130 may be owned and operated by the service provider of the application 110, or by a third party entity with which the service provider has contracted to analyze the telemetry data 126 and detect anomalies from the telemetry data 126 on behalf of the service provider of the software application 110.
[0028] Generally speaking, the telemetry data 126 includes data generated as a result of the client application 122 accessing (connecting to, or disconnecting from) the software application 10, and as a result of the user 102 using the software application 110 via the client application 122. The telemetry data module 124 may cause the telemetry data 126 to be locally stored in the local memory of the client device 103 and / or sent to the server 128 periodically and / or in response to an event or rule. For example, the telemetry data module 124 may store the telemetry data 126 in local storage and / or send it every few minutes (e.g., every 5, 10, or 15 minutes) or at another suitable time interval as the software application 110 is being accessed by the client application 122 and used by the user 102. As another example, rules maintained by the client application 122 may specify that the telemetry data 126 will be stored locally and / or sent in response to an event (e.g., an event including a successful connection to the software application 110, or an event including the generation of a specific error code indicating a connection failure, etc.). Thus, as the client device 104 is used to access the software application 110 from various geographical locations, the anomaly detection system 130 may receive telemetry data 126 originating from multiple client devices 104.
[0029] Telemetry data 126 that is sent periodically and / or in response to an event or rule can include various types and amounts of data, depending on the implementation. For example, telemetry data 126 sent from an individual client device 104 can include, but is not limited to, a user identifier (e.g., the user's globally unique identifier (GUID)), a machine identifier identifying the client device 104 used to connect to the software application 110, the machine type (e.g., mobile phone, laptop, etc.) along with information about the build, manufacturer, model, etc., logs of successful connections, executions, failed requests, logs of errors and error codes, a server identifier, logs of user input commands received via the user interface of the client application 122, service connection data (e.g., login events, auto-discovery events, etc.), user feedback data (e.g., feedback about features of the software application 110), client configuration, logs of the time periods taken by the client application 122 to respond to user request events (longer time periods indicate that the client application 122 is hung or crashed), logs of the time periods taken by the server to respond to client requests, logs of the time periods for the following states: connected, disconnected, transient failure, attempt to connect, fault lockout, or waiting, etc. It should be noted that the telemetry data 126 does not include personal or private information other than the user identifier, and the collection of any data considered to be of a personal or private nature will not be collected without first obtaining the explicit consent of the user 102.
[0030] The servers 128 of the anomaly detection system 130 can be arranged in a cluster or as a server farm and span servers 128 across multiple clusters, and are shown as being equipped with one or more electronic processors 132 and one or more forms of computer-readable memory 134. The electronic processor 132 can be configured to execute instructions, applications, or programs stored in the memory 134. In some embodiments, the electronic processor 132 can include a hardware processor, which includes but is not limited to: an electronic central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), or a combination thereof.
[0031] The computer-readable memory 134 is an example of a computer storage medium. Generally, the computer-readable memory 134 can include computer-executable instructions that, when executed by the processor 132, perform the various functions and / or operations described in this application.
[0032] Components included in the computer-readable memory 134 can include a telemetry data collector 136 configured to collect or otherwise receive telemetry data 126 from which anomalies related to the software application 110 can be detected. The telemetry data collector 136 can be configured to receive telemetry data 126 from multiple client devices 104 as the client devices 104 are used to access the software application 110. The telemetry data 126 received by the telemetry data collector 136 can be maintained in one or more data stores of the anomaly detection system 130. Over time, a history of the telemetry data 126 is obtained, the telemetry data having timestamps corresponding to the times at which the telemetry data 126 was collected or sent from the telemetry data modules 124 of the client devices 104.
[0033] In some embodiments, the telemetry data 126 is divided into multiple categories of data such that a particular category can be specifically used to detect anomalies regarding the data of that category. Additionally, the categorization of the telemetry data 126 can be hierarchically organized, where different category levels can be defined. For example, a high-level category of data can include "errors", and the data for the category of "errors" can include multiple lower-level subcategories for each unique error or error code. Another high-level category can be defined for "users", and the data for the category of "users" can include multiple lower-level subcategories for each unique user ID. Similar category hierarchies can be defined, and any example of the telemetry data 126 described in this application can be associated with a single category, included in a higher-level category, and / or include lower-level subcategories in its own category.
[0034] Additionally, the raw telemetry data 126 can be transformed (or translated) into a set of metrics (e.g., counts or rates), which will be described in more detail below. For example, instances of a particular error code in the telemetry data 126 can be counted to produce a count of the particular error code. Data for a category for a particular error code can additionally or alternatively be analyzed over time to generate a rate for the particular error code as another type of metric. Any similar metrics for a given category can be generated from the raw telemetry data 126 using the methods described in this application.
[0035] In some embodiments, the memory 134 further includes an aggregation module 138 configured to aggregate certain categories of data according to components of interest in the system in which the software application 110 is implemented. Many components of interest can be defined for the system in which the software application 110 is implemented. For example, components of interest can include application crashes or error messages generated during the operation of the application program.
[0036] As an illustrative example, the aggregation module 138 can be configured to aggregate a count (e.g., a metric) of a particular error code (e.g., a category of data) based on the applications 110 accessed by a particular user 102(1). This allows a service provider of the application to monitor the operational status of the application 110 by analyzing the count of the particular error code with respect to the user 102(1) to see if there are anomalies in the data. An anomaly can be detected in an example scenario where multiple client devices 104 report an unusually high number of instances of a particular error code and / or an unusually long time period during which the particular error code is displayed (compared to the predicted count of the particular error code and / or the predicted time period during which the particular error code is to be displayed). This detected anomaly can be used to infer that a problem has occurred.
[0037] Figure 2Schematic diagram of an example computer architecture of the anomaly detection system 130. In some embodiments, the anomaly detection system 130 includes telemetry data 126, blob storage 131, machine learning training module 133 - anomaly detector 140, Structured Query Language (SQL) data warehouse 135, and application programming interface (API) 137 associated with the set of metrics associated with the telemetry data 126. In some embodiments, the blob storage 131 provides a storage mechanism (where data can be stored in, for example, one-hour chucks or one-day chucks) for storing unstructured data as objects (or blobs) in the cloud. The machine learning training module 133 reads a set of training data, which typically consists of data such as metrics for the previous 14 days (but it can be any arbitrary amount of data or even a random sampling of data). Metrics other than the dimensions of interest (e.g., build type or quantity) are aggregated, and additional metrics are calculated based on the distribution of values. The result is saved as a model in the blob storage 131 and used for the next slice / window of data (e.g., for the next day). The anomaly detector 140 is configured to determine a prediction error by comparing the prediction with the value of the aggregated metric received from the aggregation module 138. In some configurations, the prediction error includes the difference between the actual value of the metric obtained from the real-time telemetry data 126 and the predicted value of the metric from the prediction. If this difference (e.g., the prediction error) is greater than a threshold difference at a particular time, this condition can be considered an anomaly. In some configurations, an anomaly is detected when the prediction error exceeds the threshold difference over a predetermined time period (e.g., 30 minutes). In this way, an anomaly is detected if the anomaly state persists for a predetermined time period. It should be understood that since the generated prediction is not an absolute threshold (e.g., a single numerical value), normal fluctuations in the metrics are not misinterpreted as anomalies. The use of an absolute threshold for the prediction is insufficient in this regard because it either detects an anomaly when there is no problem with the monitored client application or fails to detect an anomaly that should have been detected. Therefore, the prediction generated by the anomaly detector 140 is more accurate than an absolute threshold prediction.
[0038] Figure 3A is an exemplary hierarchical diagram 300 showing various pivots for anomaly monitoring. As Figure 3A shown, the hierarchical diagram 300 has an inverted tree-like branching structure and a root node 302. In the example provided, the root node 302 is associated with an operating system (e.g., Windows or Win 32). The root node 302 branches into several branches (each branch having a pivot point) until each branch reaches a leaf node that can no longer branch further. In Figure 3AIn the example shown, the root node 302 branches into two branches to represent application programs 310 and 320. In this example, application program 310 is a spreadsheet application (e.g., Excel software application), and application program 320 is a word processing application (e.g., Word software application). Each software application can have several branches, e.g., production, testing, demonstration, etc. In Figure 3A the example provided, two branches are provided. The production version 312 of the software branches out from application program 310, and the test version 322 of the software branches out from application program 320.
[0039] The software builds 314 and 324 associated with the specific application programs branch out from the production version 312 of the software and the test version 322 of the software, respectively. An operation 330 (in this case, "open file") branches out from software build 314, and an operation 340 (in this case, "save file") associated with application program 320 branches out from software build 324. Additionally, "error type - 1" (represented by 336 and 350) associated with the scenario when various operations are still functional but not yet rejected is provided. For example, as Figure 3A shown, there is a single spike in the error on September 1 and no spikes on September 2 and 3. The application program 320 is generally available, except when the abnormally increased "file save" operations (at pivot point 340) on September 2 and 3 result in access being rejected due to "error type - 2" (as shown in block 342). Similarly, when the "open file" 330 operation is abnormal on all days (i.e., September 1, 2, and 3), access is rejected due to "error type - 2" (as shown in block 332).
[0040] Figure 3B An example of a pivot 380 for rolling up exceptions to perform statistical inference is shown. In Figure 3BIn the example shown, pivot 380 includes pivot points 381 - 387, which represent various types of parameters for which anomalies associated with software application 110 are monitored. Block 382 represents the software application 110 being monitored (e.g., Word, Excel, etc.). Block 383 represents the target object group (e.g., a list of users, customers, etc.). Block 384 represents the channel associated with the client that received a particular build of the software application 110. Block 385 represents the platform (or operating software) associated with the software application 110 (e.g., Win32, Android, etc.). Block 386 represents the team being affected (e.g., an internal or external team). Block 388 represents the application version (e.g., a particular build of the application). The systems and methods provided in this application allow for rolling up and across pivot 380 and the statistical inferences made at block 388 to analyze the various anomalies associated with each pivot point 381 - 387. In some embodiments, the roll-up is performed based on at least one of the severity of the anomaly, the frequency of the anomaly, and the user impact associated with the anomaly. In some embodiments, the frequency of the anomaly is determined by calculating the ratio of the total number of data points associated with the anomaly identified per dimension to the corresponding pivot.
[0041] Figure 3C Another example of a pivot 390 for rolling up anomalies to perform statistical inferences is shown. In Figure 3C the example shown, pivot 390 includes pivot points 391 - 398, which represent various types of parameters for which anomalies associated with software application 110 are monitored. Pivot point 391 represents the language version associated with the software application 110 (e.g., English, French, Japanese, etc.). Block 392 represents the country version associated with the software application 110 (e.g., United States, France, Japan, etc.). Block 393 represents the device category associated with the software application 110 (e.g., mobile or desktop). Block 394 represents the architecture associated with the software application 110 (e.g., the type of processor, such as ARM64, x86, etc.). Block 395 represents a particular installation type (e.g., Windows Installer or Click-to-Run version of Office). Block 396 represents the various versions of the operating system. Block 397 represents the rollout of various features (e.g., "control group" vs. "general"). Block 398 represents various file location information (e.g., internal / external cloud or a path with a Universal Naming Convention (UNC)). The systems and methods provided in this application allow for rolling up and across pivot 390 and the statistical inferences made at block 388 to analyze the various anomalies associated with each pivot point 391 - 398.
[0042] Figure 4It is a flowchart of an example method 400 for generating a predictive scoring model for certain categories of data aggregated for components of interest in a system implementing a software application 110. Method 400 is described with reference to the preceding figures. In one embodiment, method 400 is implemented in software executed by, for example, an electronic processor 132.
[0043] At block 402, a telemetry data collector 136 receives data from a plurality of client devices 104 over a period of time. For example, as a result of the client devices 104 executing corresponding client applications 122 to access the software application 110, telemetry data 126 is sent over the network 112 and received by the telemetry data collector 136. The telemetry data 126 includes data points representing error codes and error - free codes in each dimension.
[0044] At block 404, the received telemetry data is classified based on the category of the data. The category of the data provides information to the anomaly detection system 130 about what the problem is (e.g., a problem related to a specific error code). At block 406, the electronic processor 132 is configured to convert the data of each category into a set of metrics across all pivots.
[0045] At block 408, method 400 includes storing the set of metrics as historical telemetry data 126. Over time, samples of the telemetry data 126 are continuously received, processed, and stored, and the historical telemetry data 126 includes a large repository of processed telemetry data 126 that can be used for anomaly detection processing.
[0046] The electronic processor 132 is configured to access a portion of the historical telemetry data 126 of a category of data over a predetermined time window (as shown at block 410). For example, the historical telemetry data 126 of the count of a specific error code in the past month can be accessed at 410.
[0047] At block 412, the electronic processor 132 is configured to aggregate the metrics of the category data across all dimensions except the dimension of interest. For example, the count of a specific error code in the past month can be aggregated by a specific operation associated with the software application 110 across multiple users running that specific operation on their client devices 104.
[0048] At block 414, the electronic processor 132 is configured to calculate a tolerance level for the metrics aggregated for each dimension and data category to produce a prediction of what values would be expected under normal conditions. Continuing with the above example, the prediction generated at block 414 can include a time series of the expected count of a specific error code over time for a specific operation on the client device 104.
[0049] At block 416, method 400 includes storing a tolerance level for new data corresponding to an anomaly. Method 400 iterates over steps 410 - 414 to produce a prediction of the count of the same error code for different users.
[0050] Figure 5 is an example method 500 flowchart for automatically detecting anomalies using Figure 4 the prediction scoring model shown in. Method 500 is described with reference to the previous figures. At block 502, telemetry data 126 originating from multiple client devices 104 is received. In some embodiments, method 500 includes receiving telemetry data from multiple client applications. For example, as a result of client device 104 executing a corresponding client application 122 to access software application 110, telemetry data 126 is sent over network 112 and received by telemetry data collector 136. In some embodiments, telemetry data 126 is received at block 502 for a specific sampling interval. Thus, telemetry data 126 can represent real - time telemetry data originating at client device 104 and can be used for runtime anomaly detection. Telemetry data 126 includes data points representing errors, error codes, and error - free codes in each dimension.
[0051] At block 504, electronic processor 132 is configured to classify telemetry data 126 into various categories at anomaly detector 140 - 1. For example, at block 504, a category of data related to an error code associated with a "crash" in the software application can be selected. As another example, a category of data related to an operation performed (e.g., "open file", "save file", etc.) associated with software application 110 can be selected at block 504.
[0052] At block 506, the data of the category selected at block 504 is converted into a set of metrics across each dimension. Electronic processor 132 is configured to aggregate the metrics of the data of the category across all dimensions (as shown at block 508).
[0053] As shown at block 510, electronic processor 132 is configured to access a prediction scoring model for stored metrics associated with a dimension of interest (described by the Figure 4 flowchart shown in). Electronic processor 132 is configured to determine a prediction error (block 512) based on accessing the prediction scoring model provided in block 510.
[0054] At block 514, electronic processor 132 detects an anomaly based on the prediction error determined at block 512. In some embodiments, the anomaly is detected based on or using a static threshold in the absence of available historical data or when only a trivial amount of data is available. Method 500 iterates over steps 504 - 514 to detect anomalies associated with software application 110, and generates an indication of the anomaly and stores the indication in data warehouse 135. In some embodiments, electronic processor 132 is configured to detect data points having an error at the deepest level associated with one or more dimensions (or pivots), and roll up data points having an error from that deepest level associated with one or more dimensions. In some embodiments, the method includes, with electronic processor 132, determining the frequency of an anomaly by calculating a ratio of the anomalies identified per dimension to the total number of data points associated with the corresponding dimension.
[0055] At block 516, electronic processor 132 outputs a warning message. The warning message is sent to display device 139. In some embodiments, electronic processor 132 is configured to generate a problem report and store the error included in a database.
[0056] The various features and advantages of some embodiments are set forth in the claims below.
Claims
1. A computer system for anomaly detection using application telemetry data, the computer system comprising: An electronic processor configured to: Receive telemetry data from a plurality of client applications, the telemetry data including data points representing errors associated with one or more operations on the plurality of client applications, Associate each data point with a corresponding branch among a plurality of branches included in a hierarchical graph of data categories, For each branch among the plurality of branches included in the hierarchical graph, generate one or more metrics based on the number of data points associated with the branch, Access a prediction scoring model to determine a predicted error associated with a branch of interest, wherein the predicted error is the difference between a predicted metric associated with the branch of interest and the one or more metrics generated for the branch of interest, When the predicted error is greater than a threshold, detect an anomaly, Output a warning message associated with the anomaly; and A display device for displaying the warning message.
2. The computer system according to claim 1, wherein, The electronic processor is further configured to: Generate the prediction scoring model using historical telemetry data.
3. The computer system according to claim 1, wherein, The electronic processor is further configured to: Detect data points having errors at the deepest level associated with the plurality of branches.
4. The computer system according to claim 1, wherein, The anomaly is determined by rolling up data points having errors from the deepest level associated with the plurality of branches.
5. The computer system according to claim 4, wherein, The anomaly is rolled up based on items selected from the group consisting of severity of the anomaly, frequency of the anomaly, and user impact associated with the anomaly.
6. The computer system according to claim 5, wherein, The electronic processor is further configured to: Determine the frequency of the anomaly by calculating the ratio of the anomaly identified for each branch to the total number of data points associated with the corresponding branch.
7. The computer system according to claim 1, wherein The electronic processor is further configured to: Generate an indication of the anomaly and store the indication in a database.
8. A method for anomaly detection using application telemetry data, the method comprising: Receiving telemetry data from a plurality of client applications, the telemetry data including data points representing errors associated with one or more operations on the plurality of client applications, Associating each data point with a corresponding branch among a plurality of branches included in a hierarchical graph of data categories, For each branch among the plurality of branches included in the hierarchical graph, generating one or more metrics based on the number of data points associated with the branch, Accessing a prediction scoring model to determine a predicted error associated with a branch of interest, wherein the predicted error is the difference between a predicted metric associated with the branch of interest and the one or more metrics generated for the branch of interest, When the predicted error is greater than a threshold, detecting an anomaly, and Outputting a warning message associated with the anomaly.
9. The method according to claim 8, further comprising: Generating the prediction scoring model using historical telemetry data.
10. The method according to claim 8, further comprising: Detecting pivot points having errors at the deepest level associated with the plurality of branches.
11. The method according to claim 10, further comprising: Rolling up the data points with errors from the deepest level associated with the plurality of branches.
12. The method according to claim 8, further comprising: Rolling up the branches based on items selected from the group consisting of the severity of the anomaly, the frequency of the anomaly, and the user impact associated with the anomaly.
13. The method according to claim 8, further comprising: Generating an indication of the anomaly and storing the indication of the anomaly in a database.
14. The method according to claim 12, further comprising: Determining the frequency of the anomaly by calculating the ratio of the anomaly identified for each branch to the total number of data points associated with the corresponding branch.
15. A non-transitory computer-readable medium containing instructions that, when executed by one or more electronic processors, cause the one or more electronic processors to perform the following operations: Receiving telemetry data from a plurality of client applications, the telemetry data including data points representing errors associated with one or more operations on the plurality of client applications, Associating each data point with a corresponding branch among a plurality of branches included in a hierarchical diagram of data categories, For each branch among the plurality of branches included in the hierarchical diagram, generating one or more metrics based on the number of data points associated with the branch, Access a prediction scoring model to determine a prediction error associated with a branch of interest, where The predicted error is the difference between the predicted metric associated with the branch of interest and the one or more metrics generated for the branch of interest; When the predicted error is greater than a threshold, detecting an anomaly; And Sending a warning message associated with the anomaly.
Citation Information
Patent Citations
Anomaly Detection and Classification Using Telemetry Data
US20170250855A1