Computer application error root cause diagnostic tool
By generating and utilizing the set of fault profiles, the complexity and time-consuming error diagnosis problems in computer applications are solved, and efficient determination of root cause of errors and elasticity improvement of computer applications are achieved.
Patent Information
- Application Number
- CN202380079974.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-28
- Filing Date
- 2023-11-02
- Publication Date
- 2025-06-27
AI Technical Summary
Modern computer applications may cause errors or failures in operating environments due to computer resource problems, and the root causes of diagnosing these errors are very complex and time-consuming.
By generating and utilizing a set of failure profiles associated with a computer application, telemetry data is collected and compared to determine the root cause of the error. The fault profile includes information and telemetry data about a specific fault, used to diagnose and mitigate errors in computer applications.
This technology can efficiently determine the root cause of errors in computer applications, reduce the time and resource requirements for diagnosis and debugging, and make computer applications more flexible.
Smart Images

Figure CN120225994A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Since the beginning of computer programming practice, computer software applications have been affected by problems caused by unexpected situations or scenarios that occur during deployment to production. That is, after a computer application is developed and deployed in an operating environment (e.g., production), when it is executed by one or more computing devices (e.g., physical computer machines or virtual computer machines), the computer application may exhibit errors or experience failures. For example, a particular computer resource (such as a network component) used by a computer application may gradually or suddenly experience an increase in traffic, such as due to an increase in its use by different computer applications, rendering the computer resource unavailable to the computer application or responding slowly to requests from the computer application. Due to insufficient capacity to meet the needs of the computer application, the computer application may not be able to perform within the range expected by the user.
[0002] The operation of modern computing applications may require the utilization of thousands of computer resources. For example, due to increased connectivity of physical hardware devices, various layers and components of many modern applications have been distributed and / or virtualized across multiple distributed devices. An initial problem with any of these computer resources can affect the operation of the computer application, thereby causing other errors or failures. On the other hand, modern computing applications are typically designed to be resilient to problems with computer resources. The term "resilience" is generally used to describe the ability of a computer application to respond to problems in one or more of its computer resources and still provide the best possible service to its users. Thus, problems with computer resources may not affect the operational performance of the computer application. Therefore, when a computer application fails or exhibits errors, determining the root cause (such as identifying which specific computer resource is responsible for the error or failure) can be an extremely difficult task.
[0003] The diagnostic process can be very complex, especially for large distributed computer applications. In particular, the analysis and troubleshooting required to attribute the cause of a computer application error are typically performed manually and require a significant amount of time, computing resources, and a broad understanding of the operational state of the computer system on which the computer application operates, as well as an understanding of the computer application and its design. Moreover, many computer resources are not independent, so a problem with one computer resource can cause problems with other computer resources, further exacerbating the challenge of determining the cause of an error in a computer application. As a result, many computer application errors may not be diagnosed effectively or in a timely manner, making it impossible to identify and resolve the root cause of the error. Instead, the resilience of the computer application may not be as strong as expected or required by the user. SUMMARY OF THE INVENTION
[0004] The present invention content is provided to introduce, in a simplified form, a selection of concepts further described below in the detailed description. The invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used in isolation to assist in determining the scope of the claimed subject matter.
[0005] Embodiments described in this disclosure generally relate to improved techniques for diagnosing and mitigating errors in computer applications. In particular, the techniques disclosed herein help determine the root cause(s) of an error affecting a computer application and can do so more efficiently without the need for additional computer resources or the cumbersome processes of conventional techniques for error detection and attribution. In some embodiments, the root cause(s) of a computer application error are diagnosed or attributed by providing a set of fault profiles associated with the computer application. A fault profile includes information about a particular fault, such as the computer resource or other failure that may be the root cause of another error or condition detected during the operation of the computer application. A fault profile may also include telemetry data, an indication of telemetry data, or information derivable from telemetry data, in combination with a particular root cause fault that typically occurs on a computer system. In some embodiments, the fault profiles are generated during the pre-production development and testing of the computer application and are subsequently available during normal operation (such as when operating in a production environment) for diagnosing the root cause of an error.
[0006] When a computer application is operating, upon detecting an error or fault in the operation of the computer application, a set of telemetry data about the operation of the computer application in the operating state is collected. In some embodiments, the operating state telemetry data is compared with the telemetry data or information derivable from telemetry data in a set of fault profiles associated with the computer application. A relevant fault profile, such as the closest matching profile, can be determined based on the comparison. Through the relevant fault profile, information about a particular fault or error, such as a particular faulty computer resource, is then used to determine the root cause of the error or fault detected in the operation of the computer application. In some instances, the information in the relevant fault profile specifies logic or computer instructions for performing diagnostic tests, such as through a computer diagnostic service, to further determine the root cause of the error and / or mitigate the error.
[0007] According to an embodiment, a computer application operates on a computer system in a test environment. The test environment computer system includes various computer resources used by the computer application or the computer system. Various aspects of the test environment are controlled according to various test conditions, such as the state of particular computer resources and / or other computer applications operating on the computer system in the test environment.
[0008] The computer application operates in a test environment under typical or ideal conditions, which can be specified according to the test conditions. When the computer application operates on a computer system in the test environment under these conditions, a telemetry data set of the computer system is recorded within a time window that includes the operation of the computer application. This telemetry data reflects the state of the computer system during the normal or ideal operation of the computer application.
[0009] In addition to operating under typical or ideal conditions, the computer application also operates in a test environment where it encounters failure incidents, such as one or more faults injected into the computer system or the test environment. When the computer application is operating and encounters a failure incident, another telemetry data set regarding the computer system is recorded within a time window that includes the operation of the computer application. This telemetry data set represents the operation of the computer system during the failure state. For different types of failure incidents, this process can be repeated to record the failure state telemetry data set for each failure incident.
[0010] Then, the failure state telemetry data set for each session involving a specific failure incident is compared with the telemetry data set obtained during normal or ideal operation. Based on the comparison, data characteristics of the telemetry data or combinations of data characteristics of the telemetry data are identified that exist in the failure state telemetry data but do not exist in the telemetry data obtained during normal or ideal operation. These telemetry data characteristics that only exist in the failure state telemetry data include telemetry data characteristics whose presence is related to the occurrence of the failure incident and are referred to herein as "failure characteristics". Thus, specific failure characteristics include various aspects of the telemetry data indicating a specific failure incident and include individual characteristics or combinations of characteristics, which can include sequences or patterns of characteristics.
[0011] Based on the corresponding failure characteristics and information regarding a specific failure incident, such as information indicating a specific computer resource that suffered a failure via fault injection, a failure profile corresponding to the specific failure incident is generated. Failure profiles can be generated for different types of failure incidents such that a set of failure profiles is generated and associated with a specific computer application operating in the test environment. The set of failure profiles corresponding to a specific computer application can be stored in a computer data repository for diagnostic computer services to access.
[0012] The computer application is then placed into an operating environment, such as being deployed into production. When in production, the computer application is operated by end users for its intended use. When operating in the operating environment, once a failure or error is detected, an operational status telemetry data set is accessed. The operational status telemetry data is compared with a failure profile corresponding to the computer application. In particular, the operational status telemetry data characteristics are compared with the failure characteristics of the failure profile such that a related or closest matching failure profile is identified. Based on the related failure profile, information about a particular failure incident indicated in the failure profile is used to determine the root cause of the failure or error detected during the operation of the computer application.
[0013] In this way, the root cause of errors occurring in even large, complex, or distributed computer applications can be efficiently determined and mitigated, making the computer application more resilient. Further and advantageously, the root cause of the error can be more accurately determined and efficiently mitigated without the need for the substantial computing resources required to support conventional computer diagnostic services. Still further, embodiments of these techniques reduce the need for technicians with a broad understanding of the operating state of the computer system on which the computer application operates and an understanding of the computer application and its design for manual analysis and troubleshooting. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Aspects of the present disclosure are described in detail below with reference to the accompanying drawings, in which:
[0015] Figure 1 is a block diagram of an example computing environment suitable for implementing embodiments of the present disclosure;
[0016] Figure 2 is a diagram depicting an example computing architecture suitable for implementing embodiments of the present disclosure;
[0017] Figure 3 and Figure 4 is a flowchart depicting an example method for determining the root cause of a computer application error in accordance with embodiments of the present disclosure;
[0018] Figure 5 is a block diagram of an example computing environment suitable for implementing embodiments of the present disclosure; and
[0019] Figure 6 is a block diagram of an example computing environment suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] Aspects of the present disclosure relate to techniques for improving the operation of computer applications, including improved and programmed error detection, testing, debugging, and increasing the resiliency of computer applications. In particular, the embodiments described herein generally relate to improved techniques for diagnosing and mitigating errors in computer applications. These techniques may help determine the root cause(s) of errors affecting a computer application and can do so more efficiently without the need for additional computer resources or the cumbersome processes of conventional techniques for error detection and attribution.
[0021] As further explained herein, a computer application generally includes a computer software package that performs a specific function (e.g., for an end user or for another computer application). Although the term software is used herein, a computer application may be implemented using software, hardware, or a combination thereof. A computer application includes one or more computer programs, services, or routines that are executed by a computer processor to perform the computational operations of running the computer application. Some computer applications operate on a computer system that acts as a client computing device, a server, or a distributed computer system. In particular, a computer application may be layered, distributed, and / or virtualized. That is, the various components of a computer application may include virtualized components and / or be distributed across multiple remote hardware devices.
[0022] As used herein, terms such as "virtual" and "virtualized" in virtualized components, virtual machines, virtual networks, virtual processors, virtual storage systems, etc. are terms in the art, for example, and refer to computer components created using software on a physical computer (or physical distributed computing system), such as machines, networks, storage systems, computers, processors, etc., to emulate the functionality of another separate physical computer component, such as machines, networks, storage systems, computers, processors, etc. Thus, the emulated physical component is referred to as a virtual component.
[0023] According to various embodiments of the present disclosure, the root cause(s) of a computer application error can be diagnosed or attributed by providing a set of fault profiles associated with a computer application. A fault profile includes information about a particular fault, such as a computer resource or other failure incident that may be the root cause of another error or condition detected during the operation of the computer application. A fault profile may also include telemetry data, an indication of the telemetry data (such as a reference or pointer to the telemetry data), or information that may be derived from the telemetry data, in combination with a particular root cause fault that typically occurs on a computer system. As used herein, telemetry data includes one or more metrics, events, activity logs, traces, or other similar data about the state of a computer system or network, such as the usage or performance of an application or application component, and may include information about the computer application operating thereon. In some instances, the telemetry data includes timestamped or time-correlated data about particular events, metrics, activity logs, and other telemetry data characteristics, and may also include sequences, bursts, trajectories, or patterns of events, metrics, or activity data relative to time.
[0024] Fault profiles can be generated during pre-production or out-of-production development and testing of a computer application and subsequently used during production to diagnose the root cause of an error. For example, in one embodiment, a fault profile includes an indication of particular telemetry data captured during the operation of a computer application at the time a particular fault occurs (such as the occurrence of a particular failure incident) and may exclude telemetry data that is also present during the operation of the computer application where the particular fault is not present. In this way, the particular telemetry data of the fault profile is associated with the occurrence of the fault and can thus be considered to indicate the fault.
[0025] Once a computer application is released (or re - released) and operating in an operating environment, an operational state telemetry data set is accessed upon detecting an error in the operation of the computer application. The operational state telemetry data includes telemetry data collected during the operation of the computer application in an operating environment such as a production environment. In some embodiments, the operational state telemetry data is recorded continuously, periodically, or as needed while the computer application is operating in the operating environment. For example, in one embodiment, telemetry data is collected as needed, such as upon detecting an error, or when operating the first instance or the first few instances of the computer application in the production environment. In some embodiments, telemetry data is always recorded while the computer application is operating. Various aspects of the obtained operational state telemetry data can be compared with the telemetry data in a failure profile corresponding to the computer application or information that can be derived from the telemetry data. At least partially based on this comparison, a relevant failure profile can be determined. For example, in one embodiment, the comparison includes a similarity comparison such that the relevant failure profile is determined as the closest - matching failure profile. In some embodiments, logic including rules or conditions is used to determine the relevant failure profile. Through the relevant failure profile, information in the profile about a particular failure, such as a particular failed computer resource or other failure incident, is used to determine the root cause of the error detected in the operation of the computer application. In some instances, the information from the relevant failure profile specifies the logic or computer instructions for performing diagnostic tests by a computer diagnostic service to determine the root cause of the error and / or mitigate the error.
[0026] At a high level and according to an embodiment, a computer application that may be in pre - production can be operated on a computer system in a test environment. The test environment or more specifically the computer system operating in the test environment can include various computer resources used by the computer application or by the computer system. The term computer resource is used herein broadly to refer to computer system resources, which can be any physical or virtual resource with limited availability within a computer system. By way of example, computer resources can include connected devices, internal system components, files, network connections, memory regions, endpoints, computing services, and dependencies, any of which can be physical or virtual. Computer resources can also include components of a computer application, such as called functions, linked objects or programs, or computer libraries used by the computer application. Computer resources can also include a combination of virtual and physical computer resources.
[0027] A test environment computer system can be physical, virtual, or a combination thereof. For example, in one embodiment, the test environment computer system includes one or more virtual machines operating within the test environment. A computer application operating in the test environment can be in pre-production or can be out of production (or at least a particular instance of the computer application operating in the test environment is not operating in production). For example, in some instances, a released computer application is evaluated in the test environment to determine possible updates to the computer application or to facilitate an error diagnosis tool, such as generating a fault profile. In some instances, a computer application that already has a fault profile determined in a pre-production test environment may be brought back into the test environment to generate additional fault profiles. The new fault profiles can correspond to different errors or can contain telemetry-related data that is different from the previous fault profiles.
[0028] Various aspects of the test environment can be controlled based on various test conditions that specify characteristics of the test environment computer system on which the computer application operates, such as the status of computer resources and / or other computer applications operating on the computer system of the test environment. For example, a particular test condition can specify the configuration of the environment, which can include the (multiple) configuration of the various computer resources of the environment, such as whether a particular computer resource is available or unavailable or has reduced availability at different times during the operation of the computer application. The computer application operates in the test environment under typical conditions and / or ideal conditions, which can be specified based on the test conditions. Here, the term ideal is used to indicate that there are no known errors or faults in the test environment. For example, under ideal conditions, the computer resources required for the computer application are always provided, and / or the components of the computer system of the test environment can operate without errors or unexpected delays. In contrast, under typical conditions, the computer application encounters normal availability of computer resources. Thus, computer resources may sometimes be unavailable, as occurs during typical operation of the computer system.
[0029] When a computer application operates on a computer system in a test environment under these typical or ideal conditions, a telemetry data set regarding the computer system is recorded within a time window that includes the operation of the computer application. This telemetry data reflects the state of the computer system during the normal (or ideal) operation of the computer application. In some instances, multiple test sessions are performed in which the computer application is operated and telemetry data is recorded. With each test session, the computer application is operated on a computer system in a test environment under typical or ideal conditions, and a telemetry data set regarding the computer system is recorded within a time window that includes the operation of the computer application during the session. In this way, each session can produce a normal operation telemetry data set, resulting in multiple normal operation telemetry data sets. It is contemplated that in some instances, there may be various differences in the telemetry data between these normal operation telemetry data sets. These differences may be due to various changes, differences, or contexts present in the state of the computer system or the test environment from one session to another. But generally, these normal operation telemetry data sets represent aspects of the state of the computer system or the test environment during the expected normal operation of the computer application. In some embodiments, at least a portion of the multiple normal operation telemetry data sets are compared to one another such that outliers, such as specific telemetry data features that only exist in some of these sets, can be removed. In some embodiments, telemetry data features that are observed in each test session or nearly each test session are retained in the telemetry data set. In this way, a normal operation telemetry data set can be generated that includes telemetry data features that may be present in each instance or session of normal (or ideal) operation.
[0030] In addition to operating under typical or ideal conditions, the computer application also operates in a test environment where it is subject to failure incidents, such as introducing one or more faults into the computer system or the test environment. For example, the test conditions can specify that an initial fault (referred to herein as the fault source) is to occur, such as a particular computer resource failing or becoming unavailable for at least a portion of the test window. In some instances, the failure incident may cause additional errors in the operation of the computer application, which can include a performance degradation. In these instances, the failure incident is considered the root cause of the additional errors caused by the failure incident.
[0031] When a computer application is operating and suffers a failure incident, another telemetry data set regarding the computer system is recorded within a time window that includes the operation of the computer application. This telemetry data set represents the operation of the computer system during a failure or fault state and is referred to herein as fault state telemetry data. In some embodiments, the telemetry data is recorded before, during, and / or after the failure incident. For different failure incidents, this process can be repeated to record a fault state telemetry data set for each failure incident. In some instances, multiple test sessions of operating the computer application and recording the telemetry data can be performed for the same type of failure incident such that multiple telemetry data sets corresponding to a particular failure incident are captured. Multiple telemetry data sets for the same fault source can be compared such that outliers are filtered out, such as particular telemetry data characteristics that exist only in some of the sets. In some embodiments, telemetry data characteristics that are observed in each test session or nearly each test session are retained in the telemetry data set. In this way, a fault state telemetry data set can be generated that includes telemetry data characteristics that are more likely to be present in each instance of a failure incident.
[0032] Further, in some embodiments, the fault state telemetry data also includes information regarding the failure incident or fault source. In particular, it is contemplated that in some instances, the test conditions can specify a particular failure incident, such as injecting a particular fault, but the actual result of the fault injection operation may be different from the specified result. For example, the test conditions can specify that a particular component is made to experience a 50% increase in latency. However, due to other circumstances in the test environment, such as the operation of other applications, the actual latency of the component can be greater than or less than a 50% increase. Thus, the telemetry data can be recorded and used to determine the actual test conditions applied to the test environment based on the specified test conditions. Additionally or alternatively, in some embodiments, the recorded fault state telemetry data is used to confirm that a particular failure incident occurred according to the specified test conditions. In some embodiments, in cases where the actual test conditions are different from the specified test conditions, the actual test conditions are considered as the fault source and / or associated with the fault state telemetry data for generating a fault profile as described herein.
[0033] In some embodiments, at least some test conditions are determined based on chaos testing and / or fault testing. Chaos testing (sometimes referred to as chaos engineering) includes the process of testing the resilience of computer software by subjecting it to various conditions that cause failures or affect the operation of a computer application. For example, in various embodiments, this includes generating targeted failures, random interruptions, bottlenecks, or limitations that affect computer resources or a computer system. Chaos testing can include deliberately simulating and / or presenting potential adverse operating scenarios (or states) to components of a computer application or a computer system. This can include applying a perturbation model that deliberately "injects" faults into components of a computer application or a test environment. Such fault injection can include component outages and / or failures, injecting various time delays into components, increases (or decreases) in component usage (e.g., component exhaustion), limitations on component bandwidth, and the like. In some instances, failures are introduced to attempt to disrupt the operation of a computer application, such as causing it to crash. In some instances, failures are introduced to attempt to cause a degradation in the performance of a computer application.
[0034] Then, the fault state telemetry data set for each test session involving a particular failure incident is compared with a telemetry data set obtained during normal or ideal operation (e.g., normal operation telemetry data). Based on the comparison, data characteristics and / or data characteristic values of the telemetry data, or combinations of data characteristics and / or data characteristic values of the telemetry data, are identified that are present in the fault state telemetry data but not in the normal operation telemetry data. These telemetry data characteristics that are only present in the fault state telemetry data include the "fault characteristics" described herein. Thus, specific fault characteristics include various aspects of the telemetry data that indicate a particular failure incident. In some instances, a fault characteristic includes a single data characteristic or a combination of data characteristics, which can include a sequence or pattern of data characteristics. In some embodiments, a one-class support vector machine (SVM) is employed to determine fault characteristics or represent fault characteristics in a vectorized form.
[0035] Generate a fault profile corresponding to a particular failure incident based on the fault characteristics of the particular failure incident and based on information about the particular failure incident (such as information indicating a particular computer resource that has suffered the failure or other information about a particular root cause failure). Fault profiles can be generated for different types of failure incidents such that a set of fault profiles is generated and associated with a particular computer application operating in a test environment. In some embodiments, a fault profile includes a data structure (or a portion thereof) that includes information about the particular computer application, such as an application ID or version information, information about the fault characteristics, and information about a particular fault, which particular fault can be a root cause failure, such as an indication of the failure incident. In some embodiments, the data of the fault profile includes one or more vectors, such as an n-dimensional vector, a mapping, or a data constellation. In some embodiments, the fault profile is indexed based on the fault characteristics or the corresponding failure incident. For example, in an embodiment, the index is used to facilitate comparison of the fault profile with operational state telemetry data. As further described herein, some embodiments of the fault profile also include or have associated logic or computer instructions to facilitate detection of a particular fault characteristic, determination of a root cause failure, and / or mitigation of the root cause failure after the root cause failure has been determined. The fault profile corresponding to a particular computer application can be stored in a computer data repository for computer service diagnostic access.
[0036] Then, operate the computer application in an operating environment; for example, the computer application can be released (or re-released) into production. During operation, once an error is detected, a set of operational state telemetry data is recorded. In some embodiments, as the computer application operates, the telemetry data is recorded continuously or periodically such that the operational state telemetry data includes telemetry data before, during, and after the error incident. Aspects of the operational state telemetry data are compared to the fault profiles corresponding to the computer application. In particular, the operational state telemetry data characteristics are compared to the fault characteristics of the fault profiles such that a relevant fault profile is determined. In some embodiments, the relevant fault profile includes the closest matching fault profile, which is the fault profile having the most fault characteristics that also exist in the operational state telemetry data. In some embodiments, a diagnostic computer service is used to perform the comparison. In some embodiments, the diagnostic computer service performs additional analysis on the operational state telemetry data or performs tests on the operating environment (or the computer system operating within the operating environment) to determine that a particular fault profile candidate is related to the detected error.
[0037] Based on a relevant failure profile, information regarding a specific failure incident indicated in the failure profile is used to determine the root cause failure of an error detected during the operation of a computer application in production. For example, in one embodiment, based on the failure incident indicated in the failure profile, a log file is automatically generated indicating information regarding the detected error and potential root cause failures of the detected error. Alternatively or additionally, the configuration of computer resources in the operating environment can be automatically adjusted based on information provided in the failure profile indicating possible root causes of the detected error to mitigate the detected error. For example, if the root cause is due to a specific computer resource being unavailable and the diagnostic computer service indicates that the specific computer resource is unavailable because it is being used by another computer application, then the diagnostic computer service can issue computer instructions to limit, reschedule, or terminate the other computer application so that the computer resource becomes available.
[0038] Overview of Technical Problem, Technical Solution, and Technical Improvement
[0039] As previously described, the operation of modern computer applications may require the utilization of thousands of computer resources, which can include distributed and / or virtualized computing resources. Problems with any of these computer resources can affect the operation of the computer application, resulting in errors or failures. At the same time, modern computer applications are typically also designed with a certain degree of resiliency to withstand a certain degree of limitations or failures in computer resources. Therefore, when a computer application fails or exhibits an error, determining the root cause of the failure or error (such as identifying which specific component is responsible for the error or failure) can be an extremely difficult task. Moreover, many computer resources are not independent, so problems with one computer resource can cause other problems with other computer resources, further exacerbating the challenge of determining the cause of an error in a computer application.
[0040] Conventional techniques for diagnosing and debugging or alleviating errors in the operation of computer applications are computer resource-intensive and error-prone. Typically, troubleshooting is a manual and time-consuming task that requires specialized computing tools and a broad understanding of the operating state of the computer system on which the computer application operates, as well as an understanding of the computer application and its design. The diagnostic process can also be very complex, especially for large distributed computer applications. Conventional troubleshooting and error alleviation typically use a trial-and-error approach, by attempting to fix or change the configuration of computer resources that are not the root cause of the problem, which can lead to further errors or problems. In particular, not knowing the exact root cause of a particular error can lead to implementing technical solutions, such as software patches for the computer application or modifications to the computer system, that do not actually address the underlying fault and may have a negative impact on the performance of the computer application, such as requiring it to operate in a suboptimal state, and / or may result in additional computer errors. Moreover, depending on the error, the particular application exhibiting the error, other computer applications, or the entire computer system may be offline or unavailable for a significant period of time while diagnostic tests are being performed and various fixes are being evaluated.
[0041] Conventional development of computer applications requires testing the computer application to attempt to anticipate potential problems and considering issues in its design when developing the computer application. However, despite such testing and consideration of some anticipated problems, conventional development and debugging techniques for computer applications are still plagued by the above-described technical limitations. Further, given the complexity of modern computer systems and computer applications, especially considering the number of computer resources involved and the uncertainty of how other computer applications operating on the computer system will affect the particular computer application being developed, it is almost impossible to anticipate or account for every problem. Thus, conventional methods for developing computer applications and diagnosing and debugging or alleviating computer application errors remain deficient. As a result, many computer application errors cannot be effectively or timely diagnosed such that the root cause of the error can be identified and addressed. Instead, users must accept and tolerate computer applications with less resilience than they expect or require.
[0042] Embodiments of the techniques described herein alleviate these and other technical deficiencies of conventional techniques. In particular, among other benefits, embodiments of the present disclosure improve the operation of computer applications by providing improved techniques for efficiently detecting problems in computer applications, including programming error detection and debugging, such that computer applications are more resilient. Moreover, embodiments of the techniques described herein improve conventional diagnostic and debugging techniques by reducing the likelihood of subsequent errors caused by misdiagnosis of the root cause of a fault, reducing troubleshooting time, reducing or eliminating the need for skilled technicians, and reducing the overhead and computing resources required to diagnose and debug computer applications.
[0043] Additionally, as provided by this disclosure, embodiments of computing techniques for programmatically determining and facilitating mitigation of the root causes of errors in the operation of computer applications are beneficial for other reasons. For example, some embodiments of the techniques can not only improve computer applications and their operation by diagnosing and correcting problems more efficiently and accurately, including some problems that would otherwise be unsolvable, but these embodiments can be automated. That is, these embodiments do not require technicians to perform tedious processes using specialized diagnostic techniques that are necessary for some conventional techniques. Further, compared to conventional techniques, some embodiments of this disclosure provide a diagnostic technique that reduces the downtime of a computer system or computer application for diagnosis and problem mitigation.
[0044] Furthermore, by generating and leveraging fault profiles based on telemetry data, as described herein, embodiments of these techniques are capable of detecting the root causes of many more computer problems than a developer could realistically anticipate during application development. Further, by generating and leveraging fault profiles, as described herein, the root causes of errors occurring in large, complex, or distributed computer applications can be efficiently determined and mitigated, making the computer application more resilient. Further, as described herein, generating and leveraging fault profiles improves computational efficiency, increases the utilization of computing resources, and saves diagnostic and debugging time. For example, by eliminating the need to run diagnostic tests on each component of a computer application, some embodiments of these techniques reduce the network bandwidth used and / or reduce other computing resources required for diagnosis and debugging. Further, as described herein, generating and leveraging fault profiles can improve the resiliency of computer application operation by more easily and accurately detecting root cause faults, such that these root cause faults can be mitigated via a revised application design or by modifying the configuration of the computer system on which the computer application operates. Further, as described herein, by generating and leveraging fault profiles based on telemetry data, the embodiments described herein provide a flexible and scalable improvement technique because fault profiles can be added or updated even after the computer application is in production (e.g., adding more fault profiles or updating existing fault profiles). For example, fault profiles can be easily added for new types of errors that may not have existed during initial computer application development. Various descriptions of the embodiments disclosed herein provide other technical advantages and improvements.
[0045] Additional Description of Embodiments
[0046] Now turning to Figure 1, a block diagram of an example computing environment 100 is provided that illustrates some embodiments of the present disclosure. It should be understood that such and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, and function groupings, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be entirely omitted for clarity. Further, many of the elements described herein may be implemented as discrete or distributed components or in combination with other components and in any suitable combination and location as functional entities. The various functions described herein as being performed by one or more entities may be performed by hardware, firmware, and / or software. For example, some functions may be performed by a processor executing instructions stored in a memory.
[0047] Among other components not shown, the example computing environment 100 includes: a number of user computing devices, such as user devices 102a and 102b through 102n; a number of data sources, such as data sources 104a and 104b through 104n; a server 106; sensors 103a and 107; and a network 110. It should be understood that Figure 1 the computing environment 100 shown is an example of a suitable operating environment. Figure 1 Each of the components shown may be implemented via any type of computing device, such as, for example, the computing device 500 described in Figure 5 connection. These components may communicate with each other via the network 110, which may include one or more local area networks (LANs) and / or wide area networks (WANs). In some implementations, the network 110 includes the Internet and / or a cellular network in any one of a variety of possible public and / or private networks.
[0048] It should be understood that within the scope of the present disclosure, any number of user devices, servers, and data sources may be employed within the computing environment 100. Each device may include a single device or multiple devices that cooperate in a distributed environment. For example, the server 106 may be provided via multiple devices arranged in a distributed environment that jointly provide the functionality described herein. Additionally, other components not shown may also be included within the distributed environment.
[0049] User devices 102a and 102b through 102n can be client-side client user devices of the computing environment 100, while the server 106 can be on the server side of the computing environment 100. The server 106 can include server-side software that is designed to work with client-side software on the user devices 102a and 102b through 102n to implement any combination of the features and functionality discussed in this disclosure. This division of the computing environment 100 is provided to illustrate an example of a suitable environment, and for each implementation, it is not required that the server 106 and any combination of the user devices 102a and 102b through 102n remain separate entities.
[0050] User devices 102a and 102b through 102n include any type of computing device that can be used by a user. For example, in one embodiment, the user devices 102a through 102n are of the type of computing device described herein with respect to Figure 5 By way of example, the user device can be implemented as a personal computer (PC), laptop computer, cellular or mobile device, smartphone, smart speaker, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA) device, music player or MP3 player, global positioning system (GPS) or device, video player, handheld communication device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, camera, remote control, appliance, consumer electronic device, workstation, any other suitable computing device, or any combination of the devices depicted.
[0051] Data sources 104a and 104b through 104n can include data sources and / or data systems that are configured to make data available for use in connection with Figure 2 any one of the various components of the computing environment 100 or system 200 described herein. For example, in one embodiment, one or more of the data sources 104a through 104n provide data to Figure 2System 200 provides (or makes available for access) data from computer resources, such as physical or virtual components of a computer system, data about components of a computer application, and / or other data about a computer system, which may include data about another application operating on the computer system. In one embodiment, data sources 104a through 104n provide telemetry data about a computer system or computing environment, such as example computing environment 100. Data sources 104a and 104b through 104n may be separate from user devices 102a and 102b through 102n and server 106, or may be combined and / or integrated into at least one of these components. In one embodiment, one or more of data sources 104a through 104n include one or more sensors, such as sensors 103a and 107, which may be integrated into or associated with one or more of (a) user devices (such as 102a) or server 106.
[0052] Computing environment 100 may be used to implement Figure 2 one or more of the components of system 200 described in Figure 3 and Figure 4 and may also be used to separately implement Figure 2 and Figure 1 each aspect of methods 300 and 400 described in
[0053] Example system 200 includes in connection with Figure 1The described network 110 and communicatively couples the components of system 200, including storage device 285 and various components shown as operating in or associated with a computing environment that includes an evaluation environment 210 and an operating environment 260. In particular and as shown in example system 200, the components operating in or associated with the evaluation environment 210 include an experimental platform 220 and a fault profile generator 250. The components shown as operating in or associated with the operating environment 260 include a subject computer application 265, computer resources 261, an error detector 262, a telemetry data recorder 264, and a diagnostic service 270. For example, the experimental platform 220 (including sub-components 222, 224, 226, 230, and 235), the fault profile generator 250 (including its sub-components 252, 254, and 256), the subject computer application 265, the error detector 262, the telemetry data recorder 264, and the diagnostic service 270 (including its sub-components 272 and 274) may be implemented as a set of compiled computer instructions or functions, program modules, computer software services, or process arrangements executed on one or more computer systems, such as the computing device 500 described in conjunction with Figure 5 or in conjunction with the Figure 6 described cloud computing platform 610.
[0054] In one embodiment, the functions performed or supported by the components of system 200 are associated with one or more computer applications, services, or routines, such as an application development platform, a software debugging application, or a computer application diagnostic tool. As described herein, these functions can operate to facilitate providing error diagnosis or debugging for a subject computer application, or to improve the resiliency of the subject computer application. In particular, such applications, services, or routines can operate on one or more user devices (such as user device 102a) or servers (such as server 106). Moreover, in some embodiments, these components of system 200 are distributed across a network, including one or more servers (such as server 106) and / or client devices (such as user device 102a), and are distributed in the cloud, such as in conjunction with Figure 6described, or may reside on a user device such as user device 102a. Moreover, these components, the functions performed by these components, or the services performed by these components may be implemented at appropriate abstraction layers of one or more computer systems, such as the operating system layer, the application layer, the hardware layer, and the like. Alternatively or additionally, the functionality of these components and / or the embodiments described herein may be performed at least in part by one or more hardware logic components. For example, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like. Additionally, although the functionality has been described herein with respect to specific components shown in example system 200, it is contemplated that in some embodiments, the functionality of these components may be shared or distributed across other components.
[0055] Continue Figure 2 , the evaluation environment 210 and the operating environment 260 each include a computing environment. In some implementations, with respect to the computing environments of the subject computer applications 235 and 265, the operating environment 260 includes a production environment, and the evaluation environment 210 includes a non-production environment, respectively. In particular, the production environment, sometimes referred to as the deployment environment, includes a computing environment in which a computer application (such as the subject computer application 265) is considered user-ready and thus deployed and executed for its intended purpose. In contrast, the non-production environment includes any environment of the software development life cycle other than the production environment, such as analysis and design, development, build, or test environments or combinations thereof. Thus, in some embodiments, a computer application (such as the subject computer application 235) operating in the evaluation environment 210 may be considered to be under development or in a testing, debugging, or refinement state. For example, the computing environments of the evaluation environment 210 and the operating environment 260 include one or more computer systems, which may include computing devices (such as computing device 500) described in conjunction with Figure 5 described, cloud computing systems (such as cloud computing platform 610) described in conjunction with Figure 6 described, or similar computing environments such as the computing environment 100 described in conjunction with Figure 1 .
[0056] An example evaluation environment 210 of system 200 includes an experimental platform 220, which is generally responsible for supporting one or more test environments for computer software, such as test environment 230. In particular, embodiments of the experimental platform 220 can be used for computer software development and can include or support various components to facilitate the testing and evaluation of computer applications under development. For example, the experimental platform 220 can be used to expose computer applications to various test conditions, such as the subject computer application 235, which can be in development and testing before deployment to production. Some embodiments of the experimental platform 220 support chaos testing or chaos engineering, which is a process of deliberately introducing faults into the operating environment of the computer application under test, thereby simulating real-world disruptions. In some embodiments, the experimental platform 220 includes or is associated with software development applications or toolkits, such as Microsoft 's Visual or Azure Chaos Studio TM .
[0057] An example experimental platform 220 includes a test environment 230, one or more test condition controllers 222, a test engine 224, and a telemetry data recorder 226. The test environment 230 is generally responsible for providing an operating environment for the computer application being tested and evaluated, such as the subject computer application 235. Embodiments of the test environment 230 include a computer system on which the subject computer application operates, such as the computing device 500 described in conjunction with Figure 5 or the cloud computing platform 610 described in conjunction with Figure 6 . In some embodiments, the test environment 230 computer system includes multiple computing devices operating in the test environment. The computer system of the test environment 230 can be a physical or virtual computer system, or can include a combination thereof, such as by having one or more physical and virtual components of a computer system. Some implementations of the experimental platform 220 can include multiple test environments 230.
[0058] Embodiments of the test environment 230 can be supported and controlled by the experimental platform 220 for evaluating the computer application being tested, such as the subject computer application 235. For example and as further explained herein, one or more sub-components of the experimental platform 220 or its sub-components can help expose the subject computer application 235 to various test conditions, such as fault injection, and observe the response of the subject computer application 235 and / or the test environment 230 computer system on which the subject computer application operates.
[0059] The test environment 230 includes a subject computer application 235 and computer resources 231. The subject computer application 235 operates on the test environment 230 computer system and may utilize one or more computer resources 231 in its operation. In particular, the subject computer application 235 includes the computer application being evaluated in the test environment 230. For example, the subject computer application 235 under development can be evaluated in the test environment 230 before being released into production. Similarly, the subject computer application 235 that has been operating in production and has experienced an error can be operated in the test environment 230 to evaluate the error, and then the design of the subject computer application 235 or the configuration of the computing system on which it runs can be modified to mitigate the error. When the subject computer application 235 operates in the test environment 230, it may be exposed to various test conditions, such as fault injection. The response of the subject computer application 235 to various test conditions and / or the response of the test environment 230 computer system while operating the subject computer application 235 can be observed or recorded.
[0060] The computer resources 231 include one or more computer resources that support the operation of the subject computer application 235 and / or the test environment 230 computer system, such as computer resources 231a and 231b through 231n. For example, the computer resource 231a can be used by the subject computer application 235 during its operation. Alternatively, the computer resource 231b can be used for the operation of the test environment 230 computer system but not necessarily by the subject computer application 235. The computer resources 231 can include any combination of virtual and / or physical computer resources. Various aspects of one or more computer resources 231 can be controlled by the experimental platform 220 (or a sub-component) to facilitate the testing of the subject computer application 235. For example, the computer resource 231n can be configured by the experimental platform 220 (or a sub-component) to be unavailable or have limited availability for at least a portion of the time when the subject computer application 235 is operating and being observed.
[0061] (Multiple) test condition controllers 222 are generally responsible for determining the test conditions of the environment 230. Embodiments of the (multiple) test condition controllers 222 specify the test conditions of the test environment 230, which may include specifying the state or configuration of one or more computer resources 231. In particular, various aspects of the test environment 230 can be controlled according to various test conditions, which specify the characteristics of the computer system of the test environment 230 on which the subject computer application 235 operates, such as conditions regarding one or more computer resources 231 and / or other computer applications operating on the test environment 230 computing system and that may affect the conditions of one or more computer resources 231. For example, in one embodiment, the specific test conditions specified by the (multiple) test condition controllers 222 include the configuration of the test environment 230 for a test session during which the subject computer application 235 operates on the test environment 230 computer system. The configuration of the test environment 230 specified by the (multiple) test condition controllers 222 may include the (multiple) configurations of the various computer resources 231 of the environment, such as whether a particular computer resource 231n is available or unavailable or the availability is reduced at various times during the test session during the operation of the subject computer application 235. In some instances, the (multiple) test condition controllers 222 may specify the state or configuration of the computer resources 231 before or during the operation of the subject computer application 235 being tested.
[0062] As described herein, the subject computer application 235 can operate on the test environment 230 computer system, in which it is evaluated under typical or ideal conditions and one or more failure incidents occur. In particular, when operating under typical conditions, the (multiple) test condition controllers 222 can specify the typical configuration of the computer resources 231, such as the configuration of the computer resources 231 expected during the operation of the subject computer application 235 in an operating environment (such as the operating environment 260). In embodiments where multiple test sessions are executed, the (multiple) test condition controllers 222 can specify the same configuration of the computer resources 231, or can specify different configurations of the computer resources 231 for at least some of the sessions. These differences in the configurations of the computer resources 231 specified by the (multiple) test condition controllers 222 as typical conditions can account for the various variations or differences in the computer system states that may generally be expected to occur. For example, in this way, the subject computer application 235 can be tested under various conditions of the test environment 230, which are not necessarily the same but are all typical conditions as long as these conditions do not occur abnormally in the operating environment.
[0063] In an embodiment of evaluating the subject computer application 235 under ideal conditions, the (multiple) test condition controllers 222 may specify the configuration of the computer resources 231 such that a particular computer resource 231 is unconstrained during use by the subject computer application 235. For example, under the ideal conditions specified by the (multiple) test condition controllers 222, the subject computer application 235 may always make available the computer resources 231 it requires, and a particular computer resource 231 may operate without errors or unexpected latency.
[0064] In the case where the subject computer application 235 operates in a test environment 230 that has suffered a failure incident, an embodiment of the (multiple) test condition controllers 222 may determine the particular failure incident that is to occur. For example, the (multiple) test condition controllers 222 may specify a failure incident that includes one or more faults to be introduced into the test environment 230. The (multiple) test condition controllers 222 may specify introducing an initial failure incident (referred to herein as a fault source) into the test environment 230, such as during a test session of the subject computer application 235, a particular computer resource 231a fails, becomes unavailable, or has limited availability for at least a portion of the time. In some instances, after being introduced into the test environment 230, the failure incident specified by the (multiple) test condition controllers 222 may cause other resultant errors related to the operation of the subject computer application 235. Thus, the failure incident can be considered the root cause of these other errors resulting from the introduction of the failure incident.
[0065] For different types of failure incidents, the process of operating the subject computer application 235 in a test environment 230 that has suffered a failure incident may be repeated. Thus, some embodiments of the (multiple) test condition controllers 222 specify different sets of failure incidents that will be introduced into the test environment 230 during various test sessions. For example, for each failure incident in the set, at least one test session may be performed. The (multiple) test condition controllers 222 may also specify introducing multiple failure incidents into the same test session. The (multiple) test condition controllers 222 may also specify failure incidents according to a predetermined routine or schedule. In some embodiments, the (multiple) test condition controllers 222 employ chaos engineering to determine at least a portion of the test conditions based on chaos testing and / or fault injection testing.
[0066] The test engine 224 is generally responsible for managing the operation of the test environment 230, including conducting test sessions for the operation of the subject computer application 235 on the test environment 230 computer system. Embodiments of the test engine 224 can operate test sessions on the test environment 230 according to test conditions specified by the (one or more) test condition controllers 222. The test engine 224 can also re-initialize the test environment 230 between test sessions such that the test conditions of a previous test session no longer exist in the test environment 230. In embodiments where multiple test sessions are conducted, the test engine 224 manages each test session.
[0067] Some embodiments of the test engine 224 include a management console to facilitate conducting test sessions. For example, some embodiments of the test engine 224 include a management console having a user interface for receiving user input of test conditions and / or for viewing data or results associated with the test, such as information about the operating state of the test environment 230 when testing the subject computer application 235.
[0068] The telemetry data recorder 226 is generally responsible for recording telemetry data associated with test sessions. In particular, for a given test session of the subject computer application 235, embodiments of the telemetry data recorder 226 can record a telemetry data set about the test environment 230. Thus, the telemetry data set includes information about the test environment 230 computer system during the operation of the subject computer application 235 of the given test session. In embodiments where multiple test sessions are conducted, the telemetry data recorder 226 can be used to obtain multiple telemetry data sets. In some embodiments, the telemetry data recorder 226 captures telemetry data before, during, and / or after the operation of the subject computer application 235 in order to obtain information about the test environment 230 computer system for a period of time before, during, and / or after the operation of the subject computer application 235.
[0069] Some embodiments of the telemetry data recorder 226 utilize one or more sensors, such as the sensors 103a and 107 described in conjunction with Figure 1 to capture or otherwise determine aspects of the telemetry data. A sensor includes functionality, routines, components, or combinations thereof for sensing, detecting, or otherwise obtaining information such as telemetry data, and can be implemented as hardware, software, or both. Thus, the telemetry data can include data sensed or determined from one or more sensors (referred to herein as sensor data). Alternatively or additionally, the telemetry data recorder 226 can obtain telemetry data from a data source, such as in conjunction with Figure 1Receive some aspects of the telemetry data from the described data sources 104a to 104n). For example, in an embodiment, one such data source 104a includes a telemetry data repository that includes one or more instrument programs or agents for providing observational data regarding an aspect of the operation of a computer system. As another example, in some embodiments, the operating system services of the test environment 230 computer system may generate a system event log that can include application and system messages, warnings, errors, or other similar events. The telemetry data recorder 226 may access or otherwise receive information from this event log data source such that it can be included in the telemetry data. In some embodiments, the telemetry data recorder 226 includes a telemetry data capture tool, which may include other tools that facilitate telemetry data capture, such as APIs, embedded instruments, or monitoring agents, and which may be provided as part of a software development kit. For example, as part of a software development kit (such as Microsoft 's Azure Chaos Studio TM ), some embodiments of the experimental platform 220 use a telemetry data capture tool to determine the telemetry data. In some embodiments, a collection of system monitoring services or monitoring agents that are part of the experimental platform 220 are used to determine the telemetry data captured by the telemetry data recorder 226.
[0070] In a test session, where the (multiple) test condition controllers 222 specify typical conditions of the test environment 230 or specify ideal conditions of the test environment 230, the telemetry data sets recorded by the telemetry data recorder 226 can be referred to herein as normal operation telemetry data or ideal operation telemetry data, respectively. In a test session where a failure incident is introduced, the telemetry data set recorded by the telemetry data recorder 226 can be referred to herein as failure state telemetry data. Some examples of failure state telemetry data can include information captured about the introduced failure incident. It is further contemplated that in some cases, the (multiple) test condition controllers 222 can specify specific failure incident test conditions for the test environment 230, such as injecting a specific fault, but the actual result conditions of the test environment 230 may be different from those specified. For example, the test conditions can specify configuring a specific computer resource 231a to experience a 50% increase in latency. However, due to other circumstances in the test environment 230, such as the operation of other applications, the latency of the computer resource 231a can be greater than or less than a 50% increase. In such a case, the telemetry data recorder 226 can capture or determine the actual test conditions (i.e., the actual conditions of the test environment), rather than the specified test conditions (i.e., the conditions of the test environment specified by the (multiple) test condition controllers 222). For example, the telemetry data recorder 226 can use sensors to determine that the latency of a specific computer resource 231a has actually increased by 65%, rather than 50% as specified test conditions.
[0071] In some embodiments, information about the specific test conditions for each test session (which can include the specified and / or actual test conditions for sessions with introduced failures) is associated with the telemetry data sets recorded by the telemetry data recorder 226 for that test session. For embodiments with associated information about specific test conditions, a specific telemetry data set can be identified as including normal operation telemetry data, ideal operation telemetry data, or failure state telemetry data. The telemetry data sets captured by the telemetry data recorder 226 can be stored in a telemetry data repository 282 in a storage device 285, and these telemetry data sets can be accessed by other components of the system 200. In some embodiments, information about the specific test conditions associated with each telemetry data set is stored in the storage device 285.
[0072] In some embodiments, telemetry data sets are organized and / or stored as structured or semi-structured data. In particular, aspects of the telemetry data sets can be captured as structured (or semi-structured) data according to telemetry data patterns (e.g., by the telemetry data recorder 226 for a test session). Alternatively or additionally, aspects of the telemetry data sets can be transformed into structured (or semi-structured) data according to telemetry data patterns (e.g., by the telemetry data recorder 226 or the experimental platform 220). For example, in one embodiment, aspects of the telemetry data in a telemetry data set include one or more n-dimensional vectors representing various characteristics of the test environment during a test session. Aspects of the telemetry data can be structured or organized as one or more feature vectors and / or eigenvalue pairs, which in some aspects can be hierarchical, such as a JSON hierarchy or other hierarchical structure. For example, according to the telemetry data pattern, telemetry data indicating the occurrence of a specific system event can be structured as a feature vector that includes a timestamp of the event (or in some instances its start time, end time, and / or duration), an event type ID indicating the type of event that occurred, and a series of one or more event type attributes, each with a corresponding value, which may be zero or empty in instances where the event type attribute is not detected. For example, the feature vector for the occurrence of this example system event can be structured as: |Timestamp:20220601|EventTypeID:06|EventTypeAttrb_1:0.005|EventTypeAttrb_2:0|…|EventTypeAttrb_n:0|. Certain aspects of the telemetry data may be present in some but not all telemetry data sets for all sessions. For example, a specific event may occur sometimes but not always. However, as described herein, organizing telemetry data sets as structured (or semi-structured) data or organizing according to a pattern facilitates comparison of telemetry data from different telemetry data sets. In some embodiments that use a telemetry data pattern to structure telemetry data, the same or a similar telemetry data pattern is used in the test session of the subject computer application 235 in the evaluation environment 210 and in the operation of the subject computer application 235 in the operating environment 260 to facilitate comparison of telemetry. For example, some embodiments utilize telemetry patterns developed by OpenTelemetry, which provides an open-source observability framework for capturing and processing telemetry data. In some instances where different telemetry data patterns are used, transformation or mapping can be utilized so that telemetry data structured according to different patterns can be compared.
[0073] Continue Figure 2, the system 200 includes a fault profile generator 250. As described herein, the fault profile generator 250 is generally responsible for determining and generating fault profiles. For example, a fault profile may include fault characteristics, such as information about a particular failure incident and various aspects of telemetry data that typically accompany the failure incident, and thus may be considered indicative of the failure incident. Embodiments of the fault profile generator 250 may generate a set of one or more fault profiles for the subject computer application 235 such that each fault profile in the set corresponds to a different failure incident or different test condition of the test environment 230.
[0074] For example and at a high level according to one embodiment, the fault profile generator 250 receives a normal operation telemetry data set and a fault state telemetry data set from the telemetry data repository 282 in the storage device 285 or the telemetry data recorder 226. The fault profile generator 250 (or its sub-components) performs a comparison of the two telemetry data sets to determine specific aspects of the telemetry data that appear in the fault state telemetry data but not in the normal operation telemetry data. Based on these differences, a set of one or more fault characteristics is determined. The fault profile generator 250 associates the determined fault characteristics with information indicating the failure incident from the fault state telemetry data set and stores this information as a fault profile. Fault profiles are typically generated by the fault profile generator 250 during pre-production or non-production development and testing (e.g., within the evaluation environment 210), but in some instances, may be generated as the subject computer application (such as the subject computer application 265) operating in the operating environment 260.
[0075] As shown in the example system 200, the fault profile generator 250 includes a telemetry data comparator 252, a fault characteristic determiner 254, and a fault profile assembler 256. The telemetry data comparator 252 is generally responsible for performing a comparison of two or more telemetry data sets to determine differences in the telemetry data between the sets being compared. For example, in one embodiment, two telemetry data sets are compared, including a normal operation (or ideal operation) telemetry data set and a fault state telemetry data set. A set of differences between the two telemetry data sets is determined. Based on the comparison performed by the telemetry data comparator 252, the differences in the telemetry data sets can be provided to the fault characteristic determiner 254 to determine specific fault characteristics corresponding to the failure incident indicated by the fault state telemetry data set used in the comparison.
[0076] Some embodiments of the telemetry data comparator 252 compare telemetry data features and / or feature values for the same (or similar) feature types to determine differences between different telemetry data sets. In particular, features (which may include features and their corresponding feature values) determined to be different can be used to determine fault features. In some embodiments, a feature difference threshold is utilized to compare a particular feature from each of the telemetry data sets in the telemetry data set such that a small difference in the feature values (e.g., a feature value difference not exceeding the difference threshold) is considered different enough for the feature to be used as a fault feature. Some embodiments of the telemetry data comparator 252 determine the distance (e.g., Euclidean distance or cosine distance) between a first feature vector from a first telemetry data set being compared and a corresponding second feature vector from a second telemetry data set being compared. In some of these embodiments, a distance threshold similar to the difference threshold described previously is used such that small differences between the feature vectors are excluded.
[0077] Some embodiments of the telemetry data comparator 252 (or the fault profile generator 250 or another sub-component) utilize telemetry data comparison logic 290 to facilitate the comparison of telemetry data sets and / or the determination of fault features. The telemetry data comparison logic 290 can include rules, conditions, associations, classification models, or other criteria to determine the feature similarity (and feature differences) of the telemetry data in two or more telemetry data sets. For example, in one embodiment, the telemetry data comparison logic 290 includes a set of rules for comparing feature values for the same feature type and determining the level of similarity (or difference) of the feature values of similar feature types. For example, some embodiments can utilize classification models, such as statistical clustering (e.g., k-means or nearest neighbor) or semantic knowledge graphs, neural networks, data constellations, or other classification techniques, on the proximity of features to each other to determine similarity or difference. In particular, different similarity techniques can be used for different types of features. Thus, the telemetry data comparison logic 290 can take many different forms, depending on the mechanism used to identify similarity and the type of features for which similarity (and thus difference) is determined. For example, in an embodiment, the telemetry data comparison logic 290 includes one or more static rules (which can be predefined or can be set based on the settings of a system administrator or developer), boolean logic, fuzzy logic, neural networks, finite state machines, decision trees (or random forests or gradient boosting), support vector machines, logistic regression, clustering, or machine learning techniques, similar statistical classification processes, other rules, conditions, associations, or combinations thereof, to identify and quantify feature similarity and thus identify feature differences.
[0078] In some embodiments, the telemetry data comparator 252 determines a similarity score or similarity weighting of telemetry data characteristics compared across two or more telemetry data sets. In some embodiments, the telemetry data comparator 252 outputs a feature similarity vector of the similarity scores, which may correspond to two (or more) telemetry data feature vectors, each telemetry data feature vector from a different telemetry data set and being compared. In some instances, the fault feature determiner 254 uses the similarity score (or similarity score vector) to rank them based on the level of difference of the telemetry data characteristics (or effectively rank them based on their ability as fault features). Thus, a score indicating a high degree of similarity means that a particular feature is less effective in indicating a fault and should therefore not be used as a fault feature because it is similar to other features. For example, in the case where a first particular telemetry data feature in a fault state telemetry data set is too similar to a corresponding second telemetry data feature in a normal operation telemetry data set, then the first telemetry data feature may be expected to perform poorly as a fault feature because it is too similar to the features of the normal operation telemetry data. Conversely, another feature (if any) with a score indicating a low degree of similarity is more suitable to be used as a fault feature.
[0079] In some embodiments, in the case of performing multiple test sessions to operate a particular subject computer application 235 to produce multiple data sets of normal operation telemetry data (or multiple data sets of ideal operation telemetry data), the telemetry data comparator 252 can be used to compare at least a portion of these multiple data sets with each other such that outliers can be removed, such as a particular telemetry data feature that exists only in some of these data sets. Some embodiments of the telemetry data comparator 252 use telemetry data comparison logic 290 to perform these comparisons. For example, in these instances, the telemetry data comparison logic 290 can be used to facilitate determining the similarity between the comparison telemetry data characteristics of the normal operation (or ideal operation) telemetry data sets. In this way, a normal operation telemetry data set (or ideal operation telemetry data set) can be generated that includes telemetry data characteristics that are more likely to occur in each instance or session of normal operation (or ideal operation).
[0080] Similarly, in some embodiments, in a case where multiple test sessions are performed on the same type of failure incident while operating a subject computer application 235 to generate multiple failure state telemetry data sets, a telemetry data comparator 252 can be used to compare at least a portion of these multiple data sets to one another such that outliers can be removed, such as specific telemetry data features that exist only in some of these data sets. Some embodiments of the telemetry data comparator 252 use telemetry data comparison logic 290 to perform these comparisons. For example, in these instances, the telemetry data comparison logic 290 can be used to facilitate determining similarities between comparison telemetry data features of the failure state telemetry data sets. In this way, a failure state telemetry data set can be generated that includes telemetry data features that are more likely to be present in each instance of the failure incident.
[0081] In some embodiments, when multiple test sessions generate multiple sets of normal operation telemetry data, ideal operation telemetry data, and / or failure state telemetry data, the telemetry data comparator 252 (or the failure profile generator 250) can use various aspects of the telemetry data from multiple data sets of the same type (e.g., normal operation, ideal operation, or failure state) to generate a composite telemetry data set. For example, for multiple normal operation telemetry data sets, the telemetry data comparator 252 (or the failure profile generator 250) can determine the average feature values of the feature data and use these values to construct a composite telemetry data set of the normal operation telemetry data. Similarly, other statistical determinations can be performed, such as the median, mode, distribution, or similar representations of the feature values can be determined and used as the composite values in the composite telemetry data set.
[0082] Generally, the failure feature determiner 254 is responsible for determining specific telemetry data that is expected to occur when the subject computer application 235 is operating under conditions of a failure incident. The specific telemetry data determined by the failure feature determiner 254 can be considered to indicate the failure incident and thus includes failure features. In particular, based on the comparison of the telemetry data sets performed by the telemetry data comparator 252 that determines differences in the telemetry data, the failure feature determiner 254 determines a set of one or more failure features. In an embodiment, the failure features include various aspects of the telemetry data, such as specific telemetry data features (or feature values) that are typically present in the failure state telemetry data sets but are typically not present in the normal operation telemetry data sets (and generally not present in the ideal operation telemetry data sets, where the ideal operation telemetry data sets are compared to the failure state telemetry data sets).
[0083] Some embodiments of the fault feature determiner 254 utilize a one-class support vector machine (SVM) to determine fault features or represent fault features in a vectorized form, which can be used to facilitate the comparison of telemetry data features. Some embodiments of the fault feature determiner 254 use telemetry data comparison logic 290. For example, as previously described, telemetry data comparison logic 290 is used by some embodiments of the fault feature determiner 254 to rank telemetry data features based on the differences in telemetry data features, which indicates their ability to act as fault features. In some embodiments, the fault feature determiner 254 (which can use telemetry data comparison logic 290) determines a difference score (or weighting) of the telemetry data features, indicating the level or degree to which the feature may be effective as a fault feature. Thus, for example, a particular feature (or feature value) that has a high similarity (which can be determined using telemetry data comparison logic 290) in the normal operation telemetry data set and the fault state telemetry data set will have a lower difference score because these features are similar and thus are less effective as fault features for identifying the fault state. Moreover, in some embodiments, the telemetry data features are ranked as fault features using the difference score, and in some embodiments, a threshold can be applied to the difference score such that only features with a corresponding difference score that meets the threshold can be used as fault features.
[0084] Some embodiments of the fault feature determiner 254 utilize (or some embodiments of the telemetry data comparison logic 290 used by the fault feature determiner 254 include) probability classification logic to classify specific telemetry data features as fault features. In some embodiments, the classification for a particular feature can be binary or can be spectral, such as a score or weight associated with the feature. Thus, some embodiments determine a feature difference score or weighting, indicating how effective a particular feature may be if used as a fault feature. Some embodiments of the telemetry data comparison logic 290 include rules used by the fault feature determiner 254 to determine which data features are more likely to be better fault features, such as features of the fault state telemetry data that are more significant for indicating a failure incident. For example, the rules can specify determining specific fault features based on a particular type of domain of the telemetry data (such as certain metrics or traces). As another example, the rules can specify using a similarity score for the telemetry data features (or specific features or a particular type of features) such that features with a greater difference (such as a lower similarity score) are more likely to be classified as fault features. Similarly, the telemetry data comparison logic 290 rules can specify that certain types of telemetry data features should be favored more than others (and thus have a higher score or weighting).
[0085] The fault profile assembler 256 is generally responsible for assembling or generating fault profiles. As previously described, a fault profile can include information about a set of fault characteristics and an indication of a particular failure incident. Embodiments of the fault profile assembler 256 can generate a fault profile based on a set of one or more fault characteristics determined by the fault characteristic determiner 254 and an indication of a particular failure incident, which can be identified from the fault status telemetry data against which the fault characteristics are compared to determine the fault characteristics. In particular, embodiments of the fault profile assembler 256 can determine an indication of the failure incident from the fault status telemetry data, such as information about the test conditions associated with the fault status telemetry data. Moreover, in some instances where the fault status telemetry data set has associated information about the actual test conditions (or where telemetry data characteristics about the actual test conditions are captured and included in the fault status telemetry data), the information about the actual test conditions can be used to determine an indication of the failure incident. In this way, these embodiments of the fault profiles generated by the fault profile assembler 256 include an indication of the failure incident that actually occurred during the test session, rather than an indication of a specified failure incident that occurred during the test session. In some embodiments, at least a portion of the fault characteristics determined by the fault characteristic determiner 254 are included in the fault profiles generated by the fault profile assembler 256. Alternatively or additionally, an indication of at least a portion of the fault characteristics can be included in the fault profiles generated by the fault profile assembler 256. In some embodiments, a fault profile is generated for each type of failure incident having a corresponding fault status telemetry data set. In this way, embodiments of the fault profile assembler 256 can determine a set of fault profiles associated with a particular subject computer application 235 such that each fault profile corresponds to a different failure incident.
[0086] As described herein, some embodiments of the fault profiles generated by the fault profile assembler 256 include data structures (or portions thereof). Some embodiments of the fault profiles also include or have associated logic or computer instructions to facilitate the detection of particular fault characteristics, determination of a root cause fault, and / or mitigation of the root cause fault after the root cause fault has been determined. For example, some embodiments of the fault profiles include at least a portion of the diagnostic logic 295, or point to an aspect of the diagnostic logic 295 (discussed further in conjunction with the diagnostic service 270) that corresponds to the particular fault or failure incident indicated by the fault profile. The fault profiles generated by the fault profile assembler 256 can be made available to other components of the system 200 and / or stored in a storage device 285, such as stored in the fault profile data repository 284, where they can be accessed by other components of the system 200.
[0087] Continue Figure 2, an example operating environment 260 of system 200 includes a subject computer application 265, computer resources 261, an error detector 262, a telemetry data recorder 264, and a diagnostic service 270. The subject computer application 265 operates on the computer system of the operating environment 260 and may utilize one or more computer resources 261 in its operation. In particular, as shown in the example system 200, as opposed to a test environment, the subject computer application 265 is released and operates in a deployment environment such as an end-user setting. In some embodiments, the subject computer application 265 includes the subject computer application 235 that has been released into the operating environment.
[0088] The computer resources 261 include one or more computer resources that support the operation of the subject computer application 265 and / or the computer system of the operating environment 260 on which the subject computer application 265 operates, such as computer resources 261a and 261b through 261n. For example, the computer resource 261a may be used by the subject computer application 265 during its operation. Alternatively, the computer resource 261b may be used for the operation of the computer system of the operating environment 260 on which the subject computer application 265 operates, but not necessarily by the subject computer application 265. The computer resources 261 may include any combination of virtual and / or physical computer resources and may include the same or similar computer resources as those described in connection with the computer resources 231 in some embodiments.
[0089] The error detector 262 is generally responsible for detecting errors in the operation of the subject computer application 265. For example, an embodiment of the error detector 262 may determine that the subject computer application 265 has crashed, lagged, or hung, or otherwise performed sub-optimally. Alternatively or additionally, some embodiments of the error detector 262 determine that the computer system of the operating environment 260 on which the subject computer application 265 operates experiences an error during or after the operation of the subject computer application 265, or that one or more other computer applications operating on the computer system of the operating environment 260 experience an error during or after the operation of the subject computer application 265. The term "error" is used herein broadly to refer to any unexpected event, condition, or operation of the computer system (e.g., the computer system of the operating environment 260) and / or the subject computer application 265, such as a difference between an expected output and an actual performance output. For example, in some instances, a particular computer service that is executed but takes longer than expected (thereby resulting in longer user wait times than expected) may be considered an error.
[0090] Some embodiments of the error detector 262 include operating system services, such as operating system error handling services, which monitor various aspects of the operating environment 260 of the computer system and check for errors. Alternatively or additionally, some embodiments of the error detector 262 include monitor or observer services, such as monitor agents, which observe activities on the operating environment 260 of the computer system or various aspects of its operation to determine possible errors. For example, the monitor or observer service can detect computer resource conflicts or consider user activities, such as a user accessing the task manager or force-terminating or restarting a program, a computer system restart, system events, or other data from the operating environment 260 of the computer system indicating a possible error. Some embodiments of the fault detector include or utilize a computer system event viewer that provides error detection functionality, such as Microsoft Event Viewer service. When an error is detected, or after an error is detected, the error detector 262 can pass an indication of the error to the telemetry data recorder 264, or can otherwise provide an indication that an error has been detected, such that the telemetry data recorder 264 can determine the telemetry data associated with the error.
[0091] The telemetry data recorder 264 is generally responsible for recording telemetry data associated with the operating environment 260 of the computer system, which can include telemetry data associated with the operation of the subject computer application 265. Telemetry data regarding the operation of the operating environment 260 of the computer system and / or the subject computer application 265 is sometimes referred to herein as operating state telemetry data. Some embodiments of the telemetry data recorder 264 are implemented as the telemetry data recorder 226, or include functionality described in connection with the telemetry data recorder 226, but operate in the operating environment 260 rather than in the evaluation environment 210. For example, some embodiments of the telemetry data recorder 264 are configured to capture and generate telemetry data sets in the same manner as the telemetry data recorder 226, as described in connection with the telemetry data recorder 226. Thus, the operating state telemetry data set can include structured or semi-structured telemetry data sets, such as those described in connection with the telemetry data recorder 226.
[0092] An embodiment of the telemetry data recorder 264 can capture telemetry data based on an indication of an error that can be detected by the error detector 262. Thus, the operational state telemetry data set captured by the telemetry data recorder 264 in such a case includes telemetry data related to the error and / or telemetry data representing aspects of the computer system state of the operating environment 260 during, before, or after the error. In some embodiments, such as in conjunction with operating environment errors, the telemetry data recorder 264 records telemetry data as needed. In some embodiments, the telemetry data recorder 264 captures telemetry data continuously, nearly continuously, or periodically such that upon indication of an error (or after indication of an error), telemetry data regarding the operation of the computer system of the operating environment 260 and / or the subject computer application 265 can be determined within a time window before, during, and / or after the error (or within a portion of the time the error persists during the error occurrence). For example, in one embodiment, the telemetry data recorder 264 records telemetry data into a circular buffer memory or a similar memory structure that is configured to continuously receive and store telemetry data by maintaining the telemetry data within a most recent duration and replacing the oldest telemetry data in the memory with newer telemetry data. The operational state telemetry data set determined by the telemetry data recorder 264 can be made available to other components of the system 200 and / or stored in a storage device 285, such as in the telemetry data repository 282, and can be accessed by other components of the system 200.
[0093] The diagnostic service 270 is generally responsible for determining the root cause of an error, such as a specific failure that results in an error associated with or in conjunction with a computer application or computer system. In particular, embodiments of the diagnostic service 270 can determine the root cause of an error associated with the subject computer application 265 operating on the computing system of the operating environment 260. Some embodiments of the diagnostic service 270 also perform operations to mitigate the failure that is the root cause of the error and / or mitigate the error caused by the root cause failure. For example, these operations can include recommending or automatically implementing changes to the configuration of the computing system of the operating environment 260 and / or one or more computer resources 261 associated with the operating environment 260, modifications to the configuration or topology of the subject computer application 265 (such as the specific computer resources 261 it uses and / or when or how it uses those computer resources 261) and / or other changes to the operation of the subject computer application 265 on the computing system of the operating environment 260. In some embodiments, the diagnostic service 270 includes computer services or routines that run as needed; for example, the diagnostic service 270 can operate automatically after the error detector 262 detects an error. Alternatively or additionally, for example, the diagnostic service 270 can be initiated by a user of the computing system of the operating environment 260.
[0094] As shown in example system 200, diagnostic service 270 includes a related fault profile determiner 272 and a root cause fault determiner 274. The related fault profile determiner 272 is generally responsible for determining one or more fault profiles related to an error (such as the error detected by error detector 262), which is associated with or in combination with a computer application and a computer system. An embodiment of the related fault profile determiner 272 receives (or accesses) operation state telemetry data associated with the error, such as a telemetry data set captured by telemetry data recorder 264 based on the error detected by error detector 262. In some embodiments, the operation state telemetry data is received from the telemetry data recorder 264 or from a storage device 285 (such as telemetry data repository 282). An embodiment of the related fault profile determiner 272 (or diagnostic service 270) also receives (or accesses) a set of one or more fault profiles associated with a specific subject computer application 265 that operates or is installed on the operating environment 260 computing system from which the operation state telemetry data is captured. Some embodiments of the related fault profile determiner 272 include functionality for determining the specific subject computer application 265 that operates or is installed on the operating environment 260 computing system from which the operation state telemetry data is captured. For example, some embodiments of the related fault profile determiner 272 are configured to determine the specific subject computer application 265 based on a process list of currently running computer services or computer applications; the registry or (a) configuration file of the operating system of the operating environment 260 computing system; or by polling, querying, or scanning the operating environment 260 computing system.
[0095] The set of fault profiles received by the related fault profile determiner 272 may include one or more fault profiles that were generated during the evaluation environment 210 testing of the subject computer application 265 and are thus associated with the subject computer application 265. Some embodiments of the fault profile include information about the specific computer application associated with the fault profile, such as the application ID or version information of the subject computer application 265, which the related fault profile determiner 272 may use to determine the set of fault profiles associated with the subject computer application 265. The set of fault profiles associated with the subject computer application 265 may be received from the fault profile data repository 284 of the storage device 285.
[0096] Using the fault profile set and the operational status telemetry data, an embodiment of the associated fault profile determiner 272 performs a comparison of the operational status telemetry data with each fault profile in the set to determine the associated fault profile. In some embodiments, the associated fault profile includes the fault profile having the fault characteristics most similar to the operational status telemetry data (i.e., the closest matching fault profile). Some embodiments of the associated fault profile determiner 272 utilize the telemetry data comparison logic 290 (as previously described) to facilitate the comparison of the operational status telemetry data and the fault characteristics of each fault profile. For example, in one embodiment, for each fault profile, a similarity comparison is performed between the telemetry data characteristics that are the fault characteristics of the fault profile and the telemetry data characteristics of the operational status telemetry data, and the fault profile most similar to the operational status telemetry data is determined as the associated fault profile. In some embodiments, the associated fault profile determiner 272 utilizes a classification model of the telemetry data comparison logic 290, such as a machine learning model, to classify the operational status telemetry data as matching or close to a particular fault profile. For example, in one embodiment, information from the operational status telemetry data is fed into an artificial neural network classifier, which is a machine learning model that has been trained on the fault profiles, such that the operational status telemetry data is classified as most similar to one of the fault profiles in the fault profiles. In some embodiments, the telemetry data comparison logic 290 may specify one or more rules for the associated fault profile determiner 272 to determine the associated fault profile. For example, prior to the comparison, specific rules may be employed to filter out certain fault profiles that may be irrelevant or to determine a subset of fault profiles that are more likely to be relevant. In this way, the computational time for performing the comparison of the fault profiles with the operational status telemetry data is further reduced. In another exemplary embodiment, a decision tree (or random forest or gradient boosting) of the telemetry data comparison logic 290 is employed to classify the fault profiles as relevant and / or eliminate irrelevant fault profiles.
[0097] In some embodiments, the associated fault profile determiner 272 determines (or the telemetry data comparison logic 290 includes instructions for determining) a correlation score for a fault profile to be compared with the operational state telemetry data. For embodiments where the correlation is based on a similarity comparison, the correlation score can include a similarity score or a similar indication of similarity. Based on the correlation score, the fault profile having the highest correlation score (or the correlation score indicating the highest degree of correlation) is determined to be the associated fault profile for the diagnostic service 270 to use in determining the potential root cause fault. In some embodiments, where multiple fault profiles can be determined to be associated (e.g., multiple fault profiles having similarity with respective aspects of the operational state telemetry data), the multiple associated fault profiles can be ranked based on their correlation or similarity (e.g., based on their correlation scores). Then, the diagnostic service 270 can use the highest-ranked fault profile (e.g., the most relevant or the fault profile with the highest correlation score) to determine the potential root cause fault. In some embodiments, a correlation threshold is employed such that those fault profiles that meet the correlation threshold (e.g., fault profiles whose correlation scores exceed the correlation threshold) are considered associated fault profiles. Then, the diagnostic service 270 can utilize each of these associated profiles to determine the potential root cause fault.
[0098] In some embodiments, where two or more fault profiles have approximately the same correlation, such as having approximately the same similarity with the operational state telemetry data, or the correlation score of each fault profile is within a threshold range of the other correlation scores, then the diagnostic service 270 can utilize each of the fault profiles to determine the potential root cause fault. Additionally or alternatively, the diagnostic service 270 (or the diagnostic logic 295 described below) can flag the two or more fault profiles and / or the subject computer application 265 for additional testing, or can specify that additional testing should be performed on the failure incidents associated with the two or more fault profiles. In this way, additional different telemetry data can be captured and used to provide new or additional fault signatures such that a more discriminative comparison can be performed with the operational state telemetry data.
[0099] In some instances, more than one subject computer application 265 is operated or installed on the operating environment 260 computer system. Thus, in some embodiments, the associated fault profile determiner 272 determines at least one subject computer application 265 for which a set of one or more fault profiles is received and which is used to determine an associated fault profile. For example, in some embodiments, a particular subject computer application 265 that is operating at the time of error detection can be determined such that its corresponding set of fault profiles is utilized as compared to another subject computer application 265 that was installed but not operating at the time the error occurred. In some embodiments, a particular subject computer application 265 associated with an error detected by the error detector 262 (e.g., the detected error is related to the operation of the particular subject computer application 265 such as the subject computer application 265 not performing as expected) can be determined such that its corresponding set of fault profiles is used as compared to another subject computer application 265 not associated with the detected error.
[0100] In some embodiments, in the case where multiple subject computer applications 265 are installed or operated on the operating environment 260 computer system, the associated fault profile determiner 272 receives multiple sets of one or more fault profiles such that each received set of fault profiles corresponds to one of the subject computer applications 265. Then, comparison operations, such as those previously described, can be performed using each set of fault profiles. Based on these comparisons, the associated fault profile in each set (or multiple associated profiles in each set, such as those described above) is determined, thereby forming multiple associated profiles. In some embodiments, the diagnostic service 270 uses these multiple associated profiles (or at least a portion of these multiple profiles) to determine at least one potential root cause fault. Alternatively, in some embodiments, the diagnostic service 270 uses the most relevant fault profile from these multiple associated profiles to determine a potential root cause fault.
[0101] The root cause fault determiner 274 is generally responsible for determining the potential root cause of an error associated with or in conjunction with a computer application of the computer system. Embodiments of the root cause fault determiner 274 use an associated fault profile (or for those embodiments where more than one fault profile is determined to be associated, at least one associated fault profile) to determine the potential root cause of an error, such as an error detected by the error detector 262. The associated fault profile used by the root cause fault determiner 274 can be determined by the associated fault profile determiner 272.
[0102] Embodiments of the root cause fault determiner 274 determine a potential root cause of an error based on information from related fault profiles. In particular, based at least in part on the failure incidents indicated in the related fault profiles, the root cause fault determiner 274 determines a potential root cause fault of the error. In some embodiments, the root cause fault is determined to be the failure incident indicated in the related fault profile. In some embodiments where multiple related fault profiles are determined, the set of potential root cause faults is determined as the failure incidents indicated in each of the related fault profiles in the related fault profiles.
[0103] Some embodiments of the root cause fault determiner 274 (or the diagnostic service 270 or another sub-component) utilize diagnostic logic 295 to assist in determining a potential root cause fault. The diagnostic logic 295 includes computer instructions, which may include rules, conditions, associations, classification models, or other criteria for determining a root cause fault, and in some embodiments, may also include mitigation operations. In some embodiments, the root cause fault determiner 274 accesses and utilizes specific aspects of the diagnostic logic 295 based on the indication of the failure incident in the related fault profile. Moreover, as previously described, some embodiments of the fault profile include at least a portion of the diagnostic logic 295 (or point to an aspect of the diagnostic logic 295) corresponding to the specific failure incident indicated by the fault profile. Thus, in some embodiments, the diagnostic logic 295 includes diagnostic logic specific to a particular failure incident. For example, for some failure incidents indicated in the related fault profile, the diagnostic logic 295 may specify that the potential root cause fault is the failure incident. For other failure incidents, the diagnostic logic 295 may include other logic for determining a potential root cause, such as logic for further testing or classification.
[0104] In some embodiments, the diagnostic logic 295 includes logic that includes computer-executable instructions for programmatically determining an indication of a root cause fault by operating on the operational state telemetry data or by obtaining other information about the operating environment 260 of the computer system. For example, some embodiments of the diagnostic logic 295 may include computer-executable instructions for performing diagnostic tests (such as polling or querying) on the operating environment 260 of the computer system. In some embodiments, the diagnostic test includes determining information about the condition of one or more computer resources 261 to determine whether a specific failure incident indicated in the related fault profile is a possible root cause fault.
[0105] Further, some embodiments of the diagnostic logic 295 for programmatically determining root cause failures include computer-executable instructions for querying or processing operational state telemetry data or other data about the operating environment 260 of the computer system to detect specific conditions indicating a specific root cause failure. For example, in some embodiments, the diagnostic logic 295 includes specific logic corresponding to the failure incidents indicated in the associated failure profile and includes computer-executable instructions for querying or processing operational state telemetry data or other data about the operating environment 260 of the computer system to detect specific conditions indicating a specific root cause failure.
[0106] Some embodiments of the diagnostic logic 295 may include logic for inferring possible root cause failures based on information from associated profiles (such as information about the failure incidents indicated in the associated failure profile) and / or based on state telemetry data or other information obtained about the operating environment 260 of the computer system. For example, based on a correlation score (or similarity determination) of the comparison of the associated failure profile with the operational state telemetry data determined by the associated failure profile determiner 272, the failure incidents indicated in the associated failure profile can be inferred as root cause failures.
[0107] Some embodiments of the diagnostic logic 295 include classification logic for classifying or inferring possible root cause failures. The classification can be based in part on information from the operational state telemetry data or other information obtained about the operating environment 260 of the computer system. For example, in some of these embodiments, a particular type of failure incident has one or more corresponding root cause failure classification models. The classification models can be trained using the telemetry data as training data labeled according to root cause failures. The operational state telemetry data and / or other data about the operating environment 260 of the computer system can be input into the (multiple) classification models, which can output a classification of potential root cause failures.
[0108] Thus, the diagnostic logic 295 can take different forms depending on the particular embodiment, the specific error detected by the error detector 262, or the failure incidents indicated in the associated profile, and can include combinations of the logic described herein. For example, the diagnostic logic 295 can include a rule set (which can include static or predefined rules, or can be set based on system administrator settings, boolean logic, decision trees (such as random forests, gradient boosting trees, or similar decision algorithms), conditions, or other logic, classification models, deterministic or probabilistic classifiers, fuzzy logic, neural networks, finite state machines, support vector machines, logistic regression, clustering, machine learning algorithms, similar statistical classification processes, or combinations thereof).
[0109] Based on the root cause failure determined by the root cause failure determiner 274 (or based on at least one potential root cause failure), the diagnostic service 270 (or another component of the system 200) may provide an indication of the root cause failure and / or perform another action, such as a mitigation operation. For example, in some embodiments, the operating environment 260 computer system presents an indication of the root cause failure. In some embodiments, an electronic notification indicating the root cause failure is generated and may be transmitted or otherwise provided to a user of the subject computer application 265 or an application developer or a system event processor associated with the operating environment 260 computer system. In some embodiments, the electronic notification includes an automatically generated information technology (IT) ticket or similar repair request and may be transmitted to an appropriate IT resource or placed in a ticket queue. In some embodiments, the diagnostic service 270 (or another component of the system 200) automatically generates a log (or updates a log in a log file, such as a system event log) indicating information about the root cause failure and, in some instances, further indicating information about the detected error, aspects of the subject computer application 265, and / or the operational state telemetry data. In some embodiments, the system event log including information about the error detected by the error detector 262 may be modified to further include supplementary information about the root cause failure, which represents a potential root cause of the error. For example, the system event log entry for the error may be annotated to further include an indication of the root cause failure (or multiple potential root cause failures).
[0110] In some embodiments, the indication of the root cause failure presented (or otherwise provided) is accompanied by a recommended mitigation operation, such as a recommended modification to the configuration of the operating environment 260 computer system and / or the subject computer application 265, which may mitigate the root cause error and / or the error detected by the error detector 262. In particular and as previously described, some embodiments of the diagnostic logic 295 include a set of mitigation operations that correspond to specific root cause failures and / or detected errors and may be used to mitigate these root cause failures and / or detected errors. For example, these mitigation operations may include recommending or automatically implementing a change to the configuration of the operating environment 260 computing system and / or one or more computer resources 261 associated with the operating environment 260, a modification to the configuration or topology of the subject computer application 265 (such as the specific computer resources 261 it uses and / or when or how it uses these computer resources 261), and / or other changes to the operation of the subject computer application 265 on the operating environment 260 computing system.
[0111] Thus, in some embodiments, access one or more mitigation operations corresponding to the root cause fault determined by the root cause fault determiner 274. The diagnostic service 270 or another component of the system 200 can access the mitigation operations from the diagnostic logic 295. Based on the specific root cause fault determined by the root cause fault determiner 274, a corresponding mitigation operation is determined from the set of mitigation operations and provided along with an indication of the specific root cause fault. Alternatively or additionally, the diagnostic service 270 (or another component of the system 200) can automatically perform the corresponding mitigation operation. For example, if the root cause fault indicates that the fault source is due to the unavailability of a particular computer resource, some embodiments of the mitigation operation can include computer-executable instructions for modifying the configuration of the particular computer resource or controlling its use by other computer applications operating on the computer system in the operating environment 260. For example, the mitigation operation can specify restricting, rescheduling, or terminating another computer application that is using the particular computer resource so that the particular computer resource is available for the subject computer application 265.
[0112] Now turning to Figure 3 and Figure 4 , for some embodiments of the present disclosure, aspects of example process flows 300 and 400 are illustratively depicted. Process flows 300 and 400 each include a method (sometimes referred to herein as method 300 and method 400), which can be executed to implement various example embodiments described herein. For example, process flow 300 or process flow 400 can be executed to programmatically determine the root cause of an error, such as an error associated with a computer application, and facilitate mitigation of the root cause of the error, thereby improving the operation of the computer application.
[0113] Each block or step of process flow 300, process flow 400, and other methods described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory, such as Figure 5 the memory 512 described in Figure 2 and / or Figure 1 the storage device 285 described in Figure 1operating on the server 106) can be distributed across multiple user devices and / or servers or distributed by a distributed computing platform, and / or can be implemented in the cloud, such as in conjunction with Figure 6 described. In some embodiments, the functions performed by the blocks or steps of process flows 300 and 400 are performed by the components of system 200 described in conjunction with Figure 2 described.
[0114] Referring to Figure 3 , various aspects of an example process flow 300 are illustratively depicted for determining the root cause of an error associated with a computer application based on the use of a failure profile. In particular, the example process flow 300 can be executed to generate a failure profile that can include a data structure that includes failure indications, such as failure incidents and telemetry data associated with the occurrence of the failure, such as in conjunction with Figure 2 described.
[0115] At block 310, method 300 includes operating a computer application on a computer system according to a first operating condition. Embodiments of block 310 operate a computer application on a computer system, such as Figure 2 the subject computer application 235 described in. In some embodiments, the computer system includes computer resources that support the operation of the computer application and operates the computer application in an operating environment (such as a test environment). The computer application can operate within a time session according to a first condition related to the operation of the computer application on the computer system. In some embodiments, the first condition includes normal or expected operating conditions. For example, during normal operation, the computer resources are available to support the operation of the computer application, and any calls or inputs received by the computer application are expected. In some embodiments, the first condition includes a condition in which the computer application operates without errors during a time session. Some embodiments of block 310 can be performed using the experimental platform 220 ( Figure 2 ). Embodiments of block 310 or additional details for performing the operation of block 310 are described in conjunction with Figure 2 , particularly the experimental platform 220, and other components of system 200 that operate in conjunction with the experimental platform 220 (such as Figure 2 those components located within the experimental platform 220) described in.
[0116] At block 320, method 300 includes monitoring the computer system during operation to determine a first condition telemetry data set. Embodiments of block 320 monitor the computer system within a time session in which the computer application operates according to the first condition to determine the first condition telemetry data set. For example, telemetry data about the computer system can be captured within a time session (e.g., by one or more telemetry data recorders, such as Figure 2The telemetry data recorder 226). In embodiments where the first condition includes normal operating conditions or where the computer application operates without errors during a time session, the first condition telemetry data can be considered normal operating telemetry data or ideal operating telemetry data, respectively, such as those described herein in connection with Figure 2 described.
[0117] In some embodiments of method 300, the operation of a computer application on a computer system according to a first operating condition is repeated over multiple sessions, and a first condition telemetry data set is captured for each session, thereby forming a plurality of first condition telemetry data sets. In these embodiments, based on the plurality of first condition telemetry data sets, composite first condition telemetry data is determined as the first condition telemetry data. For example, in some embodiments, the telemetry data values of each first condition telemetry data set in the first condition telemetry data sets are averaged, or another statistical operation is performed to determine the composite telemetry data values to be included in the final composite first condition telemetry data set. Alternatively or additionally, in some embodiments, a comparison operation (such as that described in connection with Figure 2 the telemetry data recorder 226 and the telemetry data comparison logic 290) is performed between the plurality of first condition telemetry data sets. Based on this comparison, a composite first condition telemetry data set is determined, which composite first condition telemetry data set includes the telemetry data sets that are present in each first condition telemetry data set of the plurality of first condition telemetry data sets or similar telemetry data sets that are present in each first condition telemetry data set of the plurality of first condition telemetry data sets. For example, similar telemetry data (such as telemetry data having similar characteristics and / or characteristic values, and which can be determined to be similar by using a similarity threshold) that appears in each first condition telemetry data set of the first condition telemetry data sets is used to generate the composite first condition telemetry data set. Further, some embodiments of this composite first condition telemetry data set only include the telemetry data that is present in each first condition telemetry data set of the plurality of first condition telemetry data sets and do not include any telemetry data that is not present in each first condition telemetry data set of the plurality of first condition telemetry data sets. Some embodiments of block 320 use the telemetry data recorder 226 ( Figure 2 ) to perform. Embodiments of block 320 or additional details for performing the operations of block 320 are described in connection with Figure 2 , particularly the telemetry data recorder 226.
[0118] At block 330, method 300 includes operating a computer application on a computer system according to a second operating condition. An embodiment of block 330 operates a computer application on a computer system within a second time session according to a second condition regarding the operation of a computer application on the computer system. In some embodiments, the second condition includes an operating condition different from the first condition, such as a fault state operating condition. In particular, the computer application can go through one or more test conditions, which can be determined or specified by a test condition controller (such as Figure 2 the (multiple) test condition controllers 222). For example, the configuration of various computer resources can be controlled to determine the resilience of a computer application operating on the computer system. Some embodiments of block 330 can be performed using the (multiple) test condition controllers 222 and / or the test engine 224 ( Figure 2 ). In conjunction with Figure 2 , in particular, the (multiple) test condition controllers 222 and the test engine 224 describe embodiments of block 330 or additional details for performing the operations of block 330.
[0119] At block 340, method 300 includes causing a first fault associated with the operation of the computer application. An embodiment of block 340 causes a fault to occur during the operation of the computer application during another second time session, such as a failure incident (described in conjunction with Figure 2 ). For example, causing a fault can include causing a failure incident, such as configuring a computer resource to be unavailable or operating in an unexpected manner. The fault can be specified according to test conditions provided by a test condition controller (such as Figure 2 the (multiple) test condition controllers 222). In particular and as described in block 330, the computer application can go through one or more test conditions, such as a first fault, which can be specified by the test condition controller and introduced into the computer system operation by the test engine (such as Figure 2 the test engine 224). In some embodiments of block 340, the first fault as a failure incident is introduced into the computer system and causes an error in the computer application operation. In some embodiments, the first fault is determined based on chaos testing. For example, an embodiment of the test engine 224 includes a chaos test engine, such as described in conjunction with Figure 2 . Some embodiments of block 340 can be performed using the (multiple) test condition controllers 222 and / or the test engine 224 ( Figure 2 ). In conjunction with Figure 2 , in particular, the (multiple) test condition controllers 222 and the test engine 224 describe embodiments of block 340 or additional details for performing the operations of block 340.
[0120] At block 350, method 300 includes monitoring the computer system during operation to determine a second conditional telemetry data set. Embodiments of block 350 monitor the computer system within a second time session, within which the computer application operates according to a second condition to determine the second conditional telemetry data set. For example, telemetry data about the computer system can be captured within the second time session (e.g., by one or more telemetry data loggers such as Figure 2 the telemetry data logger 226). Due to the first failure induced at block 340, the second conditional telemetry data set determined at block 350 can be regarded as failure state telemetry data, such as described in conjunction with Figure 2 . Thus, the second conditional telemetry data determined at block 350 can include telemetry data associated with the first failure, such as telemetry data reflecting the response of the computer system and / or the computer application to the first failure. Some embodiments of block 350 are performed using the telemetry data logger 226 ( Figure 2 ). Embodiments of block 350 or additional details for performing the operations of block 350 are described in conjunction with Figure 2 , particularly the telemetry data logger 226.
[0121] Some embodiments repeat blocks 340 and 350 for multiple sessions, where for each session, the computer application operates on the computer system according to a second condition, and the first failure is induced during the session. In these embodiments, a second conditional telemetry data set is captured for each of the multiple sessions, thereby forming multiple second conditional telemetry data sets, and composite second conditional telemetry data is determined therefrom as the second conditional telemetry data set. For example, in some embodiments, the telemetry data values of each second conditional telemetry data set in the second conditional telemetry data sets are averaged, or another statistical operation is performed, to determine the composite telemetry data value to be included in the final composite first conditional telemetry data set. Alternatively or additionally, in some embodiments, a comparison operation is performed between the multiple second conditional telemetry data sets (such as described in conjunction with Figure 2as described by the telemetry data recorder 226 and the telemetry data comparison logic 290). Based on this comparison, a composite second-condition telemetry data set is determined, which includes the telemetry data sets present in each of the plurality of second-condition telemetry data sets or similar telemetry data sets present in each of the plurality of second-condition telemetry data sets. For example, similar telemetry data (such as telemetry data having similar characteristics and / or characteristic values, and which can be determined to be similar by utilizing a similarity threshold) that appears in each of the second-condition telemetry data sets in the second-condition telemetry data sets is used to generate the composite second-condition telemetry data set. Further, some embodiments of this composite second-condition telemetry data set only include the telemetry data present in each of the plurality of second-condition telemetry data sets and do not include any telemetry data that is not present in each of the plurality of second-condition telemetry data sets.
[0122] At block 360, method 300 includes determining a fault signature based on a comparison of the first-condition telemetry data set and the second-condition telemetry data set. The fault signature can include aspects of the data that are included in the second-condition telemetry data set and not included in the first-condition telemetry data set. Some embodiments of block 360 include determining a plurality of fault signatures. In particular, embodiments of block 360 can programmatically determine at least one aspect of the data that is included in the second-condition telemetry data set and not included in the first-condition telemetry data set based on a comparison of the first-condition telemetry data set and the second-condition telemetry data set. For example, specific telemetry data characteristics and / or data characteristic values that are determined to be present in the second-condition telemetry data set but not present in the first-condition telemetry data set can be determined as fault signatures. In some embodiments, the specific telemetry data characteristics and / or data characteristic values determined to be fault signatures are sufficiently different from the telemetry data characteristics and / or data characteristic values to serve as an indication of a first fault associated with the second-condition telemetry data set. Some embodiments of block 360 utilize the telemetry data comparison logic 290 to perform the comparison and / or determine the fault signature, as described in conjunction with Figure 2 described. Some embodiments of block 360 utilize a one-class support vector machine (SVM) to determine the fault signature. Some embodiments of block 360 are performed using the fault profile generator 250 and / or using the telemetry data comparator 252 and the fault signature determiner 254 ( Figure 2 ). Embodiments of block 360 or additional details for performing the operations of block 360 are described in conjunction with Figure 2 , particularly the fault profile generator 250, the telemetry data comparator 252, and the fault signature determiner 254.
[0123] At block 370, method 300 includes generating a first fault profile as a data structure, the first fault profile including an indication of a first fault and an indication of fault characteristics. An embodiment of block 370 generates a fault profile, such as described in connection with Figure 2 the fault profile generator 250, the fault profile including an indication of the first fault caused at block 340 and an indication of the fault characteristics (or fault characteristics for embodiments that determine multiple fault characteristics) determined at block 360. In some embodiments of block 370, the fault profile further includes an indication of the computer application and / or the fault profile is associated with the computer application. For example, as described herein, some embodiments of the fault profile include information regarding the application ID or application version information. In some embodiments of block 370, the fault profile further includes logic or computer-executable instructions, such as described in connection with the fault profile generator 250 and the diagnostic service 270 ( Figure 2 ) to facilitate the detection of fault characteristics or to determine that the first fault is a root cause fault, and / or to mitigate the first fault as a root cause fault. Some embodiments of block 370 are performed using the fault profile generator 250 and / or using the fault profile assembler 256 ( Figure 2 ). Embodiments of block 370 or additional details for performing the operations of block 370 are described in connection with Figure 2 , particularly the fault profile generator 250 and the fault profile assembler 256.
[0124] Some embodiments of method 300 may be repeated to generate additional fault profiles based on other different faults. In this way, a set of fault profiles associated with the computer application is generated. Further, some embodiments of method 300 include using at least one fault profile to determine the root cause of an error detected in connection with the operation of the computer application.
[0125] Reference Figure 4 illustrates aspects of an example process flow 400 for determining the root cause of an error associated with a computer application based on using fault profiles. In particular, the example process flow 400 may be performed to determine the root cause of an error associated with the operation of a computer application by using a set of fault profiles, such as described in connection with Figure 2 . At block 410, method 400 includes detecting an error during the operation of a computer application on a computer system. In some embodiments, the error includes an unexpected event associated with the operation of the computer application and may be detected during or after the operation of the computer application on the computer system. Some embodiments of block 410 use an error detector or an error handling computer service to detect the error, such as the error detector 262 described in connection with Figure 2 . Some embodiments of block 410 are performed using the error detector 262 ( Figure 2 ). In connection withFigure 2 , and in particular, the error detector 262 describes an embodiment of block 410 or additional details for performing the operations of block 410.
[0126] At block 420, method 400 includes determining a first telemetry data set within a time frame including the occurrence of an error. An embodiment of block 420 determines a first telemetry data set including telemetry data of a computer system. The telemetry data is captured during a time frame including the occurrence of the error. In some embodiments, the telemetry data is captured after the occurrence of the error. In some embodiments, the captured telemetry data includes telemetry data from before, during, and / or after the error. Some embodiments of block 420 use the telemetry data recorder 264( Figure 2 ) to perform. In conjunction with Figure 2 , and in particular, the telemetry data recorder 264 describes an embodiment of block 420 or additional details for performing the operations of block 420.
[0127] At block 430, method 400 includes determining a relevant failure profile. An embodiment of block 430 determines a relevant failure profile associated with the first telemetry data set. Since the first telemetry data set may include telemetry data associated with the error detected at block 410, the relevant failure profile determined at block 430 can be considered related to the detected error. Some embodiments of block 430 determine the relevant failure profile as described in conjunction with Figure 2 the relevant failure profile determiner 272 in
[0128] According to method 400, determining the relevant failure profile at block 430 includes the following operations. At block 432, a set of one or more failure profiles is accessed. The failure profiles may be associated with computer applications operating on the computer system. In some embodiments, each failure profile in the set includes at least an indication of a failure. For example, each failure profile may include an indication of a failure incident, such as described in conjunction with Figure 2 the failure profile generator 250 and the relevant failure profile determiner 272 in
[0129] In block 434, for each failure profile in the failure profile set, a comparison operation is performed on the first telemetry data set and the failure profile. An embodiment of block 434 performs a comparison operation that compares the first telemetry data set with each failure profile in the failure profile set, thereby forming a set of comparison operations. In some embodiments of block 434, each comparison operation includes a similarity comparison and provides an indication of similarity, such as a similarity score, for the first telemetry data set and the failure profile being compared. For example and as described in conjunction with the telemetry data comparison logic 290( Figure 2)As described, the similarity comparison may include a comparison of data characteristics or determining the difference between the data feature vectors of the first telemetry data and the fault profile. Based on the similarity indication from the comparison, the fault profile of the comparison with the highest similarity indication (such as the highest similarity score) is determined as the relevant fault profile.
[0130] In some embodiments of block 434, each comparison operation includes a classification operation. For example, as described in connection with the relevant fault determiner 272 ( Figure 2 )), a classification model may be utilized, where the model receives an aspect of the first telemetry data set as input. For example, in one embodiment, the classification model includes an artificial neural network (ANN) that is trained on the fault characteristics of fault profiles corresponding to different faults. Aspects of the first telemetry data set (such as one or more telemetry data characteristics) are provided as input to the classification model, and then the classification model classifies the input telemetry data characteristics of the first telemetry data set as being similar to a particular fault profile. At block 436, the particular fault profile classified by the ANN (or classification model) may be determined as the relevant fault profile.
[0131] In block 436 of method 400, based on the set of comparison operations, a relevant fault profile is determined from the set of fault profiles. Embodiments of block 436 determine a relevant fault profile from the set of fault profiles accessed at block 432 based on the comparison operations performed at block 434. For example, in some embodiments where the comparison operation includes a similarity comparison, the relevant fault profile is determined as the most similar fault profile, such as the fault profile with the highest similarity score. Some embodiments of block 430 or blocks 432, 434, and 436 are performed using the diagnostic service 270 and / or using the relevant fault determiner 272 ( Figure 2 )). Embodiments of blocks 430 and 432, 434, and 436 or additional details for performing the operations of blocks 430 or 432, 434, and 436 are described in connection with Figure 2 , in particular the diagnostic service 270 and the relevant fault determiner 272.
[0132] At block 440, method 400 includes determining the root cause of the error based on the relevant fault profile. Embodiments of block 440 may determine the root cause of the error detected at block 410 based on the relevant fault profile determined at block 430. In some embodiments, the root cause of the error determined at block 440 is based on the fault indication in the relevant fault profile. For example, in some instances, the failure indicated in the relevant fault profile at block 440 is determined as the root cause of the error. In some embodiments, the fault profile also includes instructions for performing diagnostic tests (such as in connection with Figure 2as described by the root cause fault determiner 274 in ) to confirm that the fault profile indicated by the relevant fault profile is the root cause of the error. Some embodiments of block 440 use the root cause fault determiner 274 ( Figure 2 ) to perform. In conjunction with Figure 2 , particularly the root cause fault determiner 274 describes embodiments of block 440 or additional details for performing the operations of block 440.
[0133] At block 450, method 400 includes providing an indication of the root cause of the error. Embodiments of block 450 provide an indication of the root cause of the error determined at block 440. In some embodiments of block 450, the indication of the root cause of the error is provided together with the indication of the error detected at block 410. In some embodiments, the indication includes an electronic notification provided to a user, such as an end user of a computer application, a system administrator, or a developer of a computer application. In some embodiments, the indication includes modifying the system event log of the computer system to include supplementary information about the root cause of the error. For example, the system event log entry for the detected error can be annotated to further include an indication of the root cause of the error.
[0134] Some embodiments of method 400 also include performing a mitigation operation corresponding to the root cause of the error. For example, some embodiments of the fault profile also include logic or computer-executable instructions (such as those described in conjunction with Figure 2 the fault profile generator 250 and the diagnostic service 270 in ) to mitigate the fault and / or the detected error that is the root cause of the error. Alternatively or additionally, some embodiments of method 400 utilize the diagnostic logic 295 ( Figure 2 ), which can include a set of mitigation operations that correspond to specific root cause faults and / or detected errors and can be used to mitigate these root cause faults and / or detected errors. Thus, based on the root cause fault, error, and / or fault profile, some embodiments of method 400 also include determining and performing a mitigation operation to mitigate the root cause of the error and / or the error. Some embodiments of method 400 also include determining a mitigation operation, such as previously described, and providing an indication of the mitigation operation as well as an indication of the root cause of the error.
[0135] Accordingly, aspects of improved techniques for diagnosing and mitigating errors in computer applications are described. It is to be understood that the various features, sub-combinations, and modifications of the embodiments described herein are useful and can be used in other embodiments without reference to other features or sub-combinations. Moreover, the order and sequence of the steps shown in the example methods 300 and 400 do not in any way limit the scope of the present disclosure, and in fact, within the embodiments of the present disclosure, these steps can occur in a variety of different sequences. Such variations and their combinations are also contemplated within the scope of the embodiments of the present disclosure.
[0136] Other Embodiments
[0137] In some embodiments, a computerized system is provided for determining a root cause of an error associated with a computer application, such as the computerized (or computer or computing) system described in any of the above embodiments. The computerized system includes at least one processor and a computer memory having computer-readable instructions embodied thereon that, when executed by the at least one processor, perform operations. The operations include operating the subject computer application on a first computer system according to a first operating condition, the subject computer application operating within a first time frame that is a first-condition session. The operations also include monitoring the first computer system during the first-condition session to determine a first-condition telemetry data set. The operations also include operating the subject computer application on the first computer system according to a second operating condition, the subject computer application operating within a second time frame that is a second-condition session. The operations also include, during the second-condition session, causing a first failure associated with the operation of the subject computer application during the second-condition session. The operations also include monitoring the first computer system during the second-condition session to determine a second-condition telemetry data set. The operations also include programmatically determining, based on a comparison of the first-condition telemetry data set and the second-condition telemetry data set, data aspects that are included in the second-condition telemetry data set and not included in the first-condition telemetry data set. The operations also include generating a first failure profile as a data structure, the first failure profile including at least an indication of the first failure and the data aspects included in the second-condition telemetry data set.
[0138] In any combination of the above embodiments of the system, the first operating condition includes a condition in which the subject computer operates without error during the first-condition session.
[0139] In any combination of the above embodiments of the system, the first computer system includes at least one computer resource that supports the operation of the subject computer application, and wherein causing the first failure includes configuring at least one computer resource to cause an error in the operation of the subject computer application.
[0140] In any combination of the above embodiments of the system, a first computer system operates a subject computer application in an evaluation environment on an experimental platform, and wherein the first failure is determined based on chaos testing.
[0141] In any combination of the above embodiments of the system, the first computer system is monitored to determine that a first condition telemetry data set includes telemetry data captured during a first condition session regarding the subject computer application or the first computer system.
[0142] In any combination of the above embodiments of the system, the operation further includes operating the subject computer application on the first computer system according to a first operating condition of a plurality of additional first condition sessions. The operation further includes monitoring each of the plurality of first condition sessions to determine a plurality of additional first condition telemetry data sets. The operation further includes performing a comparison operation on the plurality of additional first condition telemetry data sets and the first condition telemetry data set to determine a set of common telemetry data aspects present in each of the additional first condition telemetry data sets and the first condition telemetry data set. The operation further includes updating the first condition telemetry data set to include the set of common telemetry data aspects. The operation further includes using the updated first condition telemetry data set for comparison with a second condition telemetry data set.
[0143] In any combination of the above embodiments of the system, the comparison of the first condition telemetry data set and the second condition telemetry data set includes one of a data feature similarity comparison operation or using a one-class support vector machine (SVM).
[0144] In any combination of the above embodiments of the system, the first failure profile further includes an indication of the subject computer application.
[0145] In any combination of the above embodiments of the system, the first failure profile further includes an environmental operation for mitigating the first failure.
[0146] In any combination of the above embodiments of the system, the operation further includes using the first failure profile to determine a root cause of an error detected in association with the subject computer application operating on a second computer system.
[0147] In some embodiments, a computerized system is provided for determining a root cause of an error associated with a computer application, such as the computerized (or computer or computing) system described in any of the above embodiments. The computerized system includes at least one processor and a computer memory having computer-readable instructions embodied thereon, which, when executed by the at least one processor, perform operations. The operations include detecting an error during operation of a subject computer application on the computer system. The operations also include determining a first telemetry data set that includes telemetry data of the computer system within a time frame including the time of occurrence of the error. The operations also include determining a first relevant fault profile by accessing a plurality of fault profiles, each of the plurality of fault profiles including at least an indication of a fault different from the faults indicated by the other fault profiles among the plurality of fault profiles; for each of the plurality of fault profiles, performing a comparison operation of the first telemetry data set and the fault profile, thereby performing a plurality of comparison operations; and determining the first relevant fault profile from the plurality of fault profiles based on the plurality of comparison operations. The operations also include determining the root cause of the error based on the first relevant fault profile. The operations also include providing an indication of the root cause of the error.
[0148] In any combination of the above embodiments of the system, the error is detected during operation of the subject computer application, and the error includes an unexpected result associated with the operation of the subject computer application.
[0149] In any combination of the above embodiments of the system, the plurality of fault profiles are associated with the subject computer application; and further includes determining the plurality of fault profiles based on the subject computer application.
[0150] In any combination of the above embodiments of the system, each of the plurality of comparison operations includes a similarity comparison that provides an indication of the similarity between the first telemetry data set and the fault profile, thereby providing a plurality of similarity indications.
[0151] In any combination of the above embodiments of the system, the first relevant fault profile is determined to be the fault profile having the highest similarity indicated by the plurality of similarity indications.
[0152] In any combination of the above embodiments of the system, each of the plurality of comparison operations includes a classification operation using a classification model that receives at least a portion of the first telemetry data set as input.
[0153] In any combination of the above embodiments of the system, the root cause of the error is determined based on the indication of the fault in the first relevant fault profile.
[0154] In any combination of the above embodiments of the system, an indication of the root cause of the error is provided together with an indication of the detected error.
[0155] In any combination of the above embodiments of the system, each fault profile among the plurality of fault profiles further includes a mitigation operation corresponding to the indicated fault.
[0156] In any combination of the above embodiments of the system, the operation further includes receiving a mitigation operation (associated fault profile mitigation operation) included in the associated fault profile.
[0157] In any combination of the above embodiments of the system, the operation further includes performing the associated fault profile mitigation operation to mitigate the root cause of the error.
[0158] In any combination of the above embodiments of the system, the operation further includes providing an indication of the associated fault profile mitigation operation and an indication of the root cause of the error.
[0159] In some embodiments, a computer-implemented method for determining the root cause of an error associated with a computer application is provided. The method includes operating a subject computer application on a first computer system according to a first operating condition, the subject computer application operating within a first time frame as a first condition session. The method further includes monitoring the first computer system during the first condition session to determine a first condition telemetry data set. The method further includes operating the subject computer application on the first computer system according to a second operating condition, the subject computer application operating within a second time frame as a second condition session. The operation further includes, during the second condition session, causing a fault associated with the operation of the subject computer application during the second condition session. The method further includes monitoring the first computer system during the second condition session to determine a second condition telemetry data set. The method further includes programmatically determining, based on a comparison of the first condition telemetry data set and the second condition telemetry data set, data aspects included in the second condition telemetry data set that are not included in the first condition telemetry data set. The method further includes generating a fault profile as a data structure, the fault profile including at least an indication of the fault and the data aspects included in the second condition telemetry data set. The method further includes providing the fault profile for determining the root cause of the error based on an indication of an error detected in conjunction with the subject computer application.
[0160] In any combination of the above embodiments of the method, the first computer system includes at least one computer resource that supports the operation of the subject computer application.
[0161] In any combination of the above embodiments of the method, causing the fault includes configuring at least one computer resource to cause an induced error in the operation of the subject computer application.
[0162] In any combination of the above embodiments of the method, the comparison of the first conditional telemetry dataset and the second conditional telemetry dataset includes a data feature similarity comparison operation.
[0163] In any combination of the above embodiments of the method, the fault profile further includes an indication of the subject computer application and a mitigation operation for mitigating the first fault.
[0164] Example Computing Environment
[0165] After describing various implementations, several example computing environments suitable for implementing the embodiments of the present disclosure are now described, including example computing devices and example distributed computing environments in Figure 5 and Figure 6 respectively. Referring to Figure 5 , an example computing device is provided and is generally referred to as computing device 500. Computing device 500 is only one example of a suitable computing environment and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the present disclosure. Nor should computing device 500 be construed as having any dependency or requirement with respect to any one or combination of the illustrated components.
[0166] Embodiments of the present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-usable or computer-executable instructions executed by a computer or other machine, such as a smartphone, tablet PC, or other mobile device, a server, or a client device, such as program modules. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs a particular task or implements a particular abstract data type. Embodiments of the present disclosure may be practiced in various system configurations, including mobile devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. Embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.
[0167] Some embodiments include an end-to-end software-based system that can operate within the system components described herein to operate computer hardware to provide system functionality. At a low level, a hardware processor can execute instructions selected from a machine language (also referred to as machine code or native) instruction set for a given processor. The processor recognizes native instructions and performs corresponding low-level functions related to, for example, logic, control, and memory operations. Low-level software written in machine code can provide more complex functionality for higher-level software. Thus, in some embodiments, computer-executable instructions can include any software, including low-level software written in machine code, high-level software such as application software, and any combination thereof. In this regard, the system components can manage resources and provide services for system functionality. Any other variations and combinations are contemplated by the embodiments of the present disclosure.
[0168] Reference Figure 5 , computing device 500 includes a bus 510 that directly or indirectly couples the following devices: a memory 512, one or more processors 514, one or more presentation components 516, one or more input / output (I / O) ports 518, one or more I / O components 520, and an illustrative power supply 522. Bus 510 represents one or more buses (such as an address bus, a data bus, or a combination thereof). Although Figure 5 the various blocks are shown as lines for clarity, in reality, these blocks represent logical components and not necessarily actual components. For example, a presentation component (such as a display device) can be considered an I / O component. Also, a processor has memory. The inventors recognize this as being in the nature of the art and reiterate Figure 5 that the figures herein only illustrate example computing devices that can be used in conjunction with one or more embodiments of the present disclosure. There is no distinction between categories such as "workstation", "server", "laptop computer", or "handheld device" because all of these are contemplated to be within Figure 5 the scope of and are referred to as "computing device".
[0169] Computing device 500 generally includes or uses a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 500 and includes volatile and non-volatile media, removable and non-removable media. By way of example, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile media, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other hardware media that can be used to store the desired information and can be accessed by computing device 500. Computer storage media does not itself include signals. Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" means a signal whose one or more characteristics are set or changed in such a manner as to encode information in the signal. By way of example, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0170] Memory 512 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Example hardware devices include, for example, solid state memory, hard disk drives, and optical disk drives. Computing device 500 includes one or more processors 514 that read data from various entities such as memory 512 or I / O components 520. As described herein, the term processor or "processor" can refer to more than one computer processor. For example, the term processor (or "processor") can refer to at least one processor, which can be a physical or virtual processor, such as a computer processor on a virtual machine. The term processor (or "processor") can also refer to multiple processors, each of which can be physical or virtual, such as a multiprocessor system, distributed processing, or distributed computing architecture, a cloud computing system, or parallel processing by more than a single processor. Further, various operations described herein as being implemented or performed by a processor can be executed by more than one processor.
[0171] (Multiple) presentation components 516 present data indications to a user or other devices. Example presentation components include display devices, speakers, printing components, vibration components, etc. The I / O port 518 allows the computing device 500 to be logically coupled to other devices including I / O components 520, some of which may be built-in. Illustrative components include keyboards, touchscreens or touch-sensitive surfaces, microphones, cameras, mice, joysticks, gamepads, satellite dishes, scanners, printers, or wireless peripherals. The I / O components 520 may provide a natural user interface (NUI) that processes air gestures, sounds, or other physiological inputs generated by the user. In some instances, the inputs may be sent to appropriate network elements for further processing. The NUI may implement any combination of speech recognition, touch and stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures associated with the display on the computing device 500, head and eye tracking, and touch recognition. The computing device 500 may be equipped with a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 500 may be equipped with an accelerometer or gyroscope capable of implementing motion detection. The output of the accelerometer or gyroscope may be provided to the display of the computing device 500 to render immersive augmented reality or virtual reality.
[0172] Some embodiments of the computing device 500 may include one or more (multiple) radios 524 (or similar wireless communication components). The radios transmit and receive radio or wireless communications. The computing device 500 may be a wireless terminal suitable for receiving communications and media via various wireless networks. The computing device 500 may communicate via wireless protocols such as Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), or Time Division Multiple Access (TDMA), and other wireless protocols, to communicate with other devices. The radio communication may be a short-range connection, a long-range connection, or a combination of short-range and long-range wireless communication connections. The short-range and long-range connection types do not refer to the spatial relationship between two devices, but rather to short-range and long-range as different categories or types of connections (e.g., primary connection and secondary connection). By way of example, a short-range connection may include a connection to a device that provides access to a wireless communication network (e.g., a mobile hotspot), such as a WLAN connection using the 802.11 protocol; a Bluetooth connection to another computing device is a second example of a short-range connection or a near-field communication connection. By way of example, a long-range connection may include a connection using one or more of CDMA, General Packet Radio Service (GPRS), GSM, TDMA, and 802.16 protocols, or other long-range communication protocols used by mobile devices.
[0173] Now refer to Figure 6, an example distributed computing environment 600 in which implementations of the present disclosure can be employed is illustratively provided. In particular, Figure 6 shows a high-level architecture of an example cloud computing platform 610, which can host a technology solution environment or a part thereof (such as a data trustee environment). It should be understood that such and other arrangements described herein are presented only as examples. Further, as described above, many of the elements described herein can be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and location. Other arrangements and elements (such as machines, interfaces, functions, sequences, and function groupings) can be used in addition to or instead of those shown.
[0174] A data center can support a distributed computing environment 600, which includes a cloud computing platform 610, racks 620, and nodes 630 (such as computing devices, processing units, or blades) in the racks 620. A technology solution environment can be implemented using a cloud computing platform 610 that runs cloud services (such as cloud computing applications) across different data centers and geographical regions. The cloud computing platform 610 can implement a structure controller 640 component for resource allocation, deployment, upgrade, and management for providing and managing cloud services. Generally, the cloud computing platform 610 is used to store data or run service applications in a distributed manner. The cloud computing platform 610 in the data center can be configured to host and support the operation of endpoints of a specific service application. The cloud computing platform 610 can be a public cloud, a private cloud, or a dedicated cloud.
[0175] The nodes 630 can be provided with a host 650 (such as an operating system or a runtime environment) that runs a defined software stack on the nodes 630. The nodes 630 can also be configured to perform specialized functionality (such as computing nodes or storage nodes) within the cloud computing platform 610. The nodes 630 are assigned to run one or more parts of a tenant's service application. A tenant can refer to a customer who uses the resources of the cloud computing platform 610. The service application components of the cloud computing platform 610 that support a specific tenant can be referred to as a multi-tenant infrastructure or a lease. Cloud services can include any software or software part that runs on top of a data center or accesses storage and computing device locations within a data center.
[0176] When more than one separate service application is supported by node 630, node 630 can be partitioned into virtual machines (such as virtual machines 652 and 654). A physical machine can also concurrently run separate service applications. A virtual machine or a physical machine can be configured as a personalized computing environment supported by resources 660 (such as computing resources, which can include hardware resources and / or software resources) in cloud computing platform 610. It is contemplated that resources 660 can be configured for a specific service application. Further, each service application can be divided into functional parts such that each functional part can run on a separate virtual machine. In cloud computing platform 610, multiple servers can be used to run service applications and perform data storage operations in a cluster. In particular, the servers can perform data operations independently but are exposed as a single device called a cluster. Each server in the cluster can be implemented as a node.
[0177] Client device 680 can be linked to a service application in cloud computing platform 610. Client device 680 can be any type of computing device, such as user device 102n described with reference to Figure 1 FIG. 1, and client device 680 can be configured to issue commands to cloud computing platform 610. In an embodiment, client device 680 can communicate with the service application via a virtual Internet Protocol (IP) and a load balancer or other means that direct communication requests to a specified endpoint in cloud computing platform 610. Components of cloud computing platform 610 can communicate with each other via a network (not shown), which can include one or more LANs and / or WANs.
[0178] Additional Structural and Functional Features of Embodiments of the Technical Solution
[0179] Having identified the various components used herein, it should be understood that any number of components and arrangements can be employed to achieve the desired functionality within the scope of the present disclosure. For example, for clarity of concept, the components in the embodiments depicted in the figures are shown connected by lines. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as single components, many of the elements described herein can be implemented as discrete or distributed components or in combination with other components and in any suitable combination and location. Some elements can be entirely omitted. Moreover, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, the various functions can be performed by a processor executing instructions stored in a memory. Thus, other arrangements and elements (such as machines, interfaces, functions, orders, and function groupings) can be used in addition to or in place of the arrangements and elements shown.
[0180] The embodiments described in the following paragraphs may be combined with one or more of the specifically described alternatives. In particular, the claimed embodiments may incorporate references to more than one other embodiment in the alternatives. The claimed embodiments may specify further limitations to the claimed subject matter.
[0181] The subject matter of various aspects of the present disclosure is described specifically herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. Instead, it is contemplated that the claimed subject matter may also be implemented in other ways, such as in combination with other existing or future technologies, including different steps or combinations of steps similar to those described herein. Moreover, although the terms "step" and / or "block" may be used herein to denote different elements of the methods employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of individual steps is explicitly described and otherwise. Each method described herein may include a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory. These methods may also be implemented as computer-usable instructions stored on a computer storage medium. These methods may be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service) or as a plug-in to another product, to name but a few.
[0182] For the purposes of this disclosure, the word "comprising" has the same broad meaning as the word "including", and the word "access" includes "receiving", "referencing", or "obtaining". Additionally, the word "communicating" has the same broad meaning as the words "receiving" or "transmitting", which are facilitated by software- or hardware-based buses, receivers, or transmitters using the communication media described herein. Further, unless otherwise indicated to the contrary, words such as "a" and "an" include the plural as well as the singular. Thus, for example, the recitation of "a feature" is satisfied in the presence of one or more features. Moreover, the term "or" includes conjunctive, disjunctive, and both (a or b, thus including a or b as well as a and b).
[0183] As used herein, the term "set" may be used to refer to an ordered (i.e., sequential) or unordered (i.e., non-sequential) collection of objects (or elements), such as machines (e.g., computer devices), physical and / or logical addresses, graphical nodes, graphical edges, functionality, etc. As used herein, a set may include N elements, where N is any positive integer. That is, a set may include 1, 2, 3, … N objects and / or elements, where N is a positive integer with no upper limit. Thus, as used herein, a set does not include the null set (i.e., the empty set), which does not include any elements (e.g., for the null set, N = 0). A set may include only a single element. In other embodiments, a set may include multiple elements significantly greater than one, two, or three or billions of elements. A set may be an infinite set or a finite set. The objects included in some sets may be discrete objects (e.g., a set of natural numbers . The objects included in other sets may be continuous objects (e.g., a set of real numbers ). In some embodiments, a "set of objects" that is not the null set of objects may be interchangeably referred to as "one or more objects" or "at least one object", where the term "object" may represent any object or element that may be included in a set. Thus, the phrases "one or more objects" and "at least one object" may be used interchangeably to refer to a set of objects that is not the null set or empty set of objects. A set of objects that includes at least two objects may be referred to as "multiple objects".
[0184] As used herein, the term "subset" is a set that is included in another set. A subset may or may not be a proper or strict subset of the other set in which the subset is included. That is, if set B is a subset of set A, then in some embodiments, set B is a proper or strict subset of set A. In other embodiments, set B is a subset of set A but not a proper or strict subset of set A. For example, set A and set B may be equal sets, and set B may be referred to as a subset of set A. In such an embodiment, set A may also be referred to as a subset of set B. If the intersection between two sets is the null set, the two sets may be disjoint sets.
[0185] As used herein, the terms "source code" and "code" may be used interchangeably to refer to human-readable instructions that at least partially implement the execution of a computer application. The source code may be encoded in one or more programming languages, such as Fortran, C, C++, Python, Ruby, Julia, R, Octave, Java, JavaScript, etc. In some embodiments, a compilation and / or linking process may be performed on the source code before implementing the execution of the computer application. As used herein, the term "executable file" may refer to any set of machine instructions that instantiates a copy of a computer application and enables one or more computing machines (e.g., a physical machine or a virtual machine) to execute, run, or otherwise implement the instantiated application. A computer application may include a set of executable files. The executable file may be a binary executable file, such as an executable set of machine instructions generated by compiling human-readable source code (in a programming language) and linking the binary objects generated by the compilation. That is, an executable file of a computer application may be generated by compiling the source code of the computer application. Although the embodiments are not limited thereto, a computer application may include human-readable source code, such as an application generated by an interpreted programming language. For example, the executable file of a computer application may include the source code of the computer application. The executable file may include one or more binary executable files, one or more source-code-based executable files, or any combination thereof. The executable file may include and depend on one or more function libraries, object libraries, etc. The executable file may be encoded in a single file, or the encoding may be distributed across multiple files. That is, the encoding of the executable file may be distributed across multiple files. The encoding may include one or more data files, where the execution of the computer application may depend on the reading and / or writing of one or more data files.
[0186] For the purposes of the detailed discussion above, embodiments of the present invention are described with reference to a computing device or a distributed computing environment; however, the computing devices and distributed computing environments depicted herein are merely illustrative. Moreover, the terms computer system and computing system may be used interchangeably herein, such that a computer system is not limited to a single computing device, and a computing system does not require multiple computing devices. Instead, as described herein, various aspects of embodiments of the present disclosure may be performed on a single computing device or multiple computing devices. Additionally, components may be configured to perform novel aspects of the embodiments, where the term "configured to" may refer to "programmed to" perform a particular task or implement a particular abstract data type using code. Further, although embodiments of the present invention may generally refer to a technical solution environment and the schematic diagrams described herein, it should be understood that the described technology may be extended to other implementation contexts.
[0187] Many different arrangements of the various components depicted, as well as components not shown, are possible without departing from the scope of the following claims. Embodiments of the present disclosure have been described, which are intended to be illustrative rather than restrictive. Alternative embodiments will become apparent to the reader of the present disclosure after reading the present disclosure and as a result of reading the present disclosure. Alternative means for achieving the above can be accomplished without departing from the scope of the following claims. Certain features and sub-combinations are useful and can be employed without reference to other features and sub-combinations and are contemplated within the scope of the claims.
Claims
1. A system for determining the root cause of an error associated with a computer application, the system comprising: a processor; and a computer memory having computer-readable instructions embodied thereon, which, when executed by the processor, perform operations including: operating a subject computer application on a first computer system according to a first operating condition, the subject computer application operating within a first time frame as a first condition session; monitoring the first computer system during the first condition session to determine a first condition telemetry data set; operating the subject computer application on the first computer system according to a second operating condition, the subject computer application operating within a second time frame as a second condition session; during the second condition session, causing a first fault associated with the operation of the subject computer application within the second condition session; monitoring the first computer system during the second condition session to determine a second condition telemetry data set; programmatically determining, by comparing the first condition telemetry data set and the second condition telemetry data set, data aspects included in the second condition telemetry data set that are not included in the first condition telemetry data set; and generating a first fault profile as a data structure, the first fault profile including at least an indication of the first fault and the data aspects included in the second condition telemetry data set.
2. The system according to claim 1, wherein the first operating condition includes the subject computer application operating without error during the first condition session.
3. The system according to claim 1, wherein the first computer system includes computer resources that support the operation of the subject computer application, and wherein causing the first fault includes configuring the computer resources to cause an error in the operation of the subject computer application.
4. The system according to claim 1, wherein the first computer system operates the subject computer application in an evaluation environment on an experimental platform, and wherein the first fault is determined based on chaos testing.
5. The system according to claim 1, wherein monitoring the first computer system to determine a first condition telemetry data set includes capturing telemetry data regarding the subject computer application or the first computer system during the first condition session.
6. The system according to claim 1, further comprising: operating the subject computer application on the first computer system according to the first operating condition for a plurality of additional first condition sessions; monitoring each of the plurality of additional first condition sessions to determine a plurality of additional first condition telemetry data sets; performing a comparison operation on the plurality of additional first condition telemetry data sets and the first condition telemetry data set to determine a set of common telemetry data aspects present in each of the additional first condition telemetry data sets and the first condition telemetry data set; updating the first condition telemetry data set to include the set of common telemetry data aspects; and and Utilize the updated first conditional telemetry dataset for comparison with the second conditional telemetry dataset.
7. The system according to claim 1, wherein the comparison of the first conditional telemetry dataset and the second conditional telemetry dataset includes one of the following: a data feature similarity comparison operation or using a one-class support vector machine (SVM).
8. The system according to claim 1, wherein the first fault profile further includes an indication of the subject computer application.
9. The system according to claim 1, wherein the first fault profile further includes a mitigation operation for mitigating the first fault.
10. The system according to claim 1, wherein the operation further includes utilizing the first fault profile to determine a root cause of an error detected in conjunction with the subject computer application operating on a second computer system.
11. A computer system for determining a root cause of an error associated with a computer application, the system comprising: a processor; and a computer memory having computer-readable instructions embodied thereon, which, when executed by the processor, perform operations including: detecting an error during operation of a subject computer application on the computer system; determining a first telemetry dataset that includes telemetry data of the computer system within a time frame including the occurrence of the error; accessing a plurality of fault profiles, each fault profile in the plurality of fault profiles including at least an indication of a fault different from the faults indicated by other fault profiles in the plurality of fault profiles; for each fault profile in the plurality of fault profiles, performing a comparison operation between the first telemetry dataset and the fault profile, thereby performing a plurality of comparison operations; and based on the plurality of comparison operations, determining a first relevant fault profile from the plurality of fault profiles; based on the first relevant fault profile, determining the root cause of the error; and providing an indication of the root cause of the error.
12. The system according to claim 11, wherein the error is detected during operation of the subject computer application, and the error includes an unexpected result associated with the operation of the subject computer application.
13. The system according to claim 11, wherein the plurality of fault profiles are associated with the subject computer application; and further includes determining the plurality of fault profiles based on the subject computer application.
14. The system according to claim 11: wherein each comparison operation in the plurality of comparison operations includes a similarity comparison that provides an indication of similarity between the first telemetry dataset and the fault profile, thereby providing a plurality of similarity indications; and wherein the first relevant fault profile is determined to be the fault profile having the highest similarity indicated by the plurality of similarity indications.
15. The system according to claim 11, wherein each of the plurality of comparison operations includes a classification operation using a classification model that receives at least a portion of the first telemetry data set as input.
16. The system according to claim 11, wherein the root cause of the error is determined based on the indication of the fault in the first correlation fault profile.
17. The system according to claim 11, wherein an indication of the root cause of the error is provided with an indication of the detected error.
18. The system according to claim 11, wherein each of the plurality of fault profiles further includes a mitigation operation corresponding to the indicated fault; and further includes: receiving the mitigation operation included in the correlation fault profile (the correlation fault profile mitigation operation); and performing the correlation fault profile mitigation operation to mitigate the root cause of the error; or providing an indication of the correlation fault profile mitigation operation and an indication of the root cause of the error.
19. A computer-implemented method for determining a root cause of an error associated with a subject computer application, the method comprising: operating the subject computer application on a first computer system according to a first operating condition, the subject computer application operating within a first time frame as a first condition session; determining a first condition telemetry data set by monitoring the first computer system during the first condition session; operating the subject computer application on the first computer system according to a second operating condition, the subject computer application operating within a second time frame as a second condition session; during the second condition session, causing a fault associated with the operation of the subject computer application during the second condition session; determining a second condition telemetry data set by monitoring the first computer system during the second condition session; programmatically determining, based on a comparison of the first condition telemetry data set and the second condition telemetry data set, data aspects included in the second condition telemetry data set that are not included in the first condition telemetry data set; generating a fault profile as a data structure, the fault profile including at least an indication of the fault and the data aspects included in the second condition telemetry data set; and providing the fault profile based on an indication of an error detected in conjunction with the subject computer application.
20. The computer-implemented method according to claim 19: wherein the first computer system includes computer resources that support the operation of the subject computer application, and wherein causing the fault includes configuring the computer resources to cause an induced error in the operation of the subject computer application; wherein the comparison of the first condition telemetry data set and the second condition telemetry data set includes a data feature similarity comparison operation; and wherein the fault profile further includes an indication of the subject computer application and a mitigation operation for mitigating the first fault.