Method and apparatus for continuous in-situ telemetry monitoring

The telemetry monitoring system addresses the inefficiencies of reactive client device diagnostics by predicting and proactively addressing issues using machine learning, enhancing the speed and efficiency of problem resolution.

JP7786923B2Active Publication Date: 2025-12-16INTEL CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021187446
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-23
Filing Date
2021-11-18
Publication Date
2025-12-16
Estimated Expiration
2041-11-18

Smart Images

  • Figure 0007786923000001
    Figure 0007786923000001
  • Figure 0007786923000002
    Figure 0007786923000002
  • Figure 0007786923000003
    Figure 0007786923000003
Patent Text Reader

Abstract

To provide continuous monitoring of telemetry in a site.SOLUTION: Disclosed are a method, a device, a system and a product for continuous monitoring of telemetry in a site. An example device includes a failure prediction unit that predicts a result of one or more execution paths. A solution handler determines one or more solution strategies for the execution path to apply the solution strategy. An influence training unit determines whether a prediction result of the execution path has changed to store influence data of the one or more solution strategies applied.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates generally to monitoring telemetry data, and more particularly to methods and apparatus for continuous monitoring of telemetry in the field. [Background technology]

[0002] To monitor client devices (e.g., computing devices) deployed in the field, telemetry data (e.g., information regarding characteristics, operating conditions, resource utilization, location, etc.) may be collected. For example, telemetry data may be pulled (e.g., requested from a central or distributed location) or pushed (e.g., sent to a central or distributed location). Such telemetry data may be analyzed to detect problems, to diagnose problems, or for any other desired purpose. [Brief explanation of the drawings]

[0003] [Figure 1] FIG. 1 is a block diagram illustrating an example environment in which telemetry data of one or more computing devices may be monitored, in accordance with the teachings of the present disclosure. [Figure 2] FIG. 2 is a block diagram of an example implementation of the telemetry monitor of FIG. 1. [Figure 3] FIG. 3 is a block diagram of an example implementation of the fault predictor of FIG. 2. [Figure 4] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 5] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 6] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 7A]4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 7B] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 7C] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 7D] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 7E] 4 is a flowchart representing machine-readable instructions that may be executed to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 8] FIG. 7B is a block diagram of an exemplary database encoding scheme according to the machine readable instructions of FIGS. 7A-7E. [Figure 9] FIG. 7B is a block diagram of an exemplary fractal similarity search query according to the machine-readable instructions of FIGS. 7A-7E. [Figure 10] FIG. 7B is a block diagram of an exemplary processing platform configured to execute the instructions of FIGS. 4-7E to implement the telemetry monitor of FIGS. 1, 2, and / or 3. [Figure 11] FIG. 7B is a block diagram of an exemplary software distribution platform for distributing software (e.g., software corresponding to the exemplary computer-readable instructions of FIGS. 4-7E) to client devices, such as consumers (e.g., for license, sale, and / or use), retailers (e.g., for sale, resale, license, and / or sublicense), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products distributed to retailers and / or direct purchase customers). DETAILED DESCRIPTION OF THE INVENTION

[0004] The drawings are not to scale. Generally, the same reference numbers are used throughout the drawings and accompanying written description to refer to the same or similar parts. As used herein, connection references (e.g., attached, coupled, connected, and joined) may include intermediate members between and / or relative movement between the elements referenced by the connection reference, unless otherwise indicated. Thus, a connection reference does not necessarily infer that two elements are directly connected and / or in a fixed relationship to each other.

[0005] Unless specifically indicated otherwise, descriptors such as "first," "second," "third," etc. are used herein without imputing or otherwise indicating a sense of priority, physical order, placement within a list, and / or ordering in any way, but are used merely as labels and / or arbitrary names to distinguish elements to facilitate understanding of the disclosed examples. In some instances, the descriptor "first" may be used to refer to an element in the detailed description, while the same element may be referred to in the claims using a different descriptor, such as "second" or "third." In such instances, such descriptors are used merely to distinguishably identify those elements, for example, that may otherwise share the same name.

[0006] When a problem occurs within a client device, organizations often take a reactive approach by attempting to identify the root cause of the problem and fix the root cause. With such an approach, organizations often lack access to data that would allow them to quickly resolve the problem. In most cases, recreating a customer's problem in-house is the only path to identifying the root cause, which is quite resource-intensive. The methods and apparatus disclosed herein facilitate lightweight profiling and / or monitoring of telemetry data in the field.

[0007] In some previous solutions, when an exemplary client device 102 experiences a performance problem or fault, the client device reports a signal over an exemplary network to an exemplary backend database, where the fault or performance problem is later recreated and a root cause analysis is performed to determine how and why the problem occurred in the first place.

[0008] When a problem occurs within a client device, organizations often take a reactive approach by attempting to identify the root cause of the problem and fix the root cause. With such an approach, organizations often lack access to data that would allow them to quickly resolve the problem. In most cases, recreating a customer's problem in-house is the only path to identifying the root cause, which is quite resource-intensive. The methods and apparatus disclosed herein facilitate lightweight profiling and / or monitoring of telemetry data in the field.

[0009] In an exemplary approach disclosed herein, a telemetry monitor predicts the outcome of an execution path and determines a resolution strategy (e.g., the best resolution strategy) to apply in an attempt to change the predicted outcome of the execution path. An execution path represents a hardware, firmware, or software module from a source data input to a data output measurement. An execution path in software is represented by a control and data flow graph (CFG-DFG) of colored executable code and hardware modules. Module scoping allows artificial intelligence and machine learning networks to partition dependent and independent variables to provide variable tuning scope. Variable tuning scope enables learned corrective reconfiguration and procedure sequences that can explore the solution space for an optimal solution adapted based on device state. Tuning can be scoped using software or hardware paths that self-change reconfigurable behavior. For example, if the predicted outcome of an execution path is a failure (e.g., a negative result, an anomaly, etc.), the telemetry monitor can apply one or more resolution strategies in an attempt to change the predicted outcome to be non-faulty. For example, a failure may include any result of an execution path that reduces or inhibits the performance of a client device or has negative consequences (e.g., an execution path that leads to overheating of a CPU, GPU, FPGA, computer accelerator, or storage solid-state drive (SSD), or that causes a particular application to stop responding, etc.).

[0010] Exemplary approaches disclosed herein enable host systems to take proactive action to mitigate current and / or future anomalous behavior. Some exemplary approaches enable machine learning of data signatures within collected telemetry data, allowing the system to discover and / or store unique signatures, cluster behaviors, predict sequences, and learn the best intervention strategy for a given set of control parameters. Systems using the inventions disclosed herein consume fewer resources to reproduce customer problems and access customer data. Forward prediction sequences are used bijectively on metadata to backward predict dependent and independent variables that label software, firmware, and hardware modules, resulting in anomalous CFG-CFG meta being labeled within a set scope.

[0011] 1 is a block diagram of an example environment 100 that includes several example client devices 102 connected to an example network 105 and an example telemetry analyzer 110 that analyzes telemetry data from the example client devices. To record the telemetry data, one or more of the example client or enterprise devices 102 includes an example telemetry monitor 114.

[0012] 1 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable logic devices (FPLDs), digital signal processors (DSPs), coarse grained reconfigurable architectures (CGRAs), image signal processors (ISPs), graphics processing units (GPUs), and central processing units (CPUs), microprocessors, etc.

[0013] An exemplary client device 102 generates telemetry data (e.g., application data, system data, etc.) that can be monitored and collected. In this example, the client device 102 communicates with an exemplary backend server 110 over an exemplary network 105. However, the client device 102 may communicate directly with the backend server 110, with an aggregation node, and / or in any other manner.

[0014] The exemplary network 105 communicatively couples the exemplary client devices 102 to the exemplary telemetry analyzer 110. The exemplary network 105 in the illustrated example of FIG. 1 may be implemented by one or more web services, cloud services, virtual private networks (VPNs), local area networks (LANs), Ethernet connections, 1G-5G cellular or satellite networks, the Internet, or any other means of communicating or relaying data. In the illustrated environment 100, multiple exemplary client devices 102 communicate with the telemetry analyzer 110 using the same exemplary network 105. In some examples, multiple client devices may use any combination of networks to communicate data to the exemplary telemetry analyzer 110.

[0015] The exemplary telemetry analyzer 110 in the illustrated example of FIG. 1 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable logic devices (FPLDs), digital signal processors (DSPs), coarse-grained reconfigurable architectures (CGRAs), image signal processors (ISPs), etc. In the exemplary telemetry monitoring system 100, one or more client devices, such as the exemplary client device 102, communicate telemetry data to the telemetry analyzer 110 to predict and intervene in negative execution paths. In the illustrated example of FIG. 1, the client devices deliver telemetry data to the telemetry analyzer 110 via the exemplary network 105. In some examples, the exemplary telemetry analyzer 110 can receive data directly from the exemplary client device 102.

[0016] 1 is implemented with logic circuitry, such as, for example, a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as, for example, one or more analog or digital circuits, logic circuits, programmable processors, application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable logic devices (FPLDs), digital signal processors (DSPs), coarse-grained reconfigurable architectures (CGRAs), image signal processors (ISPs), etc. Furthermore, data communicated by the example telemetry monitor 115 may be in any data format, such as, for example, binary data, comma-separated data, tab-separated data, structured query language (SQL) structures, dynamic self-describing structures, etc.

[0017] In operation, the example client device 102 generates system and application data that is reported by the example telemetry monitor 115 and collected and monitored by the example telemetry analyzer 110. Upon predicting a failure, the example telemetry analyzer 110 applies an intervention strategy to interrupt the execution path of the example client device 102 and change the predicted outcome of the execution path. In this example, the example client device is connected to the example network 105 and thereby communicatively coupled to the telemetry analyzer 110. In the examples disclosed herein, upon a failure, data is communicated (e.g., by the telemetry monitor 115) from the example client device 102 over the example network 105 to the example telemetry analyzer 110 for triage analysis. In some examples, the triage analysis may occur within the example client device.

[0018] A flowchart representing example hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the example telemetry analyzer 110 of FIGS. 1-2 is shown in FIG. 4. The machine-readable instructions may be one or more executable programs or portions of executable programs for execution by a computer processor and / or processor circuitry, such as the processor 1012 shown in the example processor platform 1000 discussed below in connection with FIG. 10. The programs may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processor 1012, although the entire program and / or portions thereof may alternatively be executed by a device other than the processor 1012 and / or embodied in firmware or dedicated hardware. Furthermore, although the example program is described with reference to the flowchart shown in FIG. 4, many other ways of implementing the example telemetry analyzer 110 may instead be used. For example, the order of execution of the blocks may be changed, and / or some of the described blocks may be modified, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (opamps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware. Processor circuits may be distributed across different network locations and / or local to one or more devices (e.g., multi-core processors in a single machine, multiple processors distributed across a server rack, etc.).

[0019] The machine-readable instructions described herein may be stored in one or more of a compressed format, an encrypted format, a fragmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions described herein may also be stored as data or data structures (e.g., portions of instructions, code, representations of code, etc.) that can be utilized to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be fragmented and stored on one or more storage devices and / or computing devices (e.g., servers) located in the same or different locations of a network or collection of networks (e.g., in the cloud, in edge devices, etc.). The machine-readable instructions may require one or more of installing, modifying, adapting, updating, combining, supplementing, configuring, decrypting, decompressing, unpacking, distributing, reassigning, compiling, etc. to make them directly readable, interpretable, and / or executable by computing devices and / or other machines. For example, machine-readable instructions may be stored in multiple portions that are individually compressed, encrypted, and stored on separate computing devices, and that when decoded, decompressed, and combined form a set of executable instructions that implement one or more functions that may together form a program as described herein.

[0020] In another example, machine-readable instructions may be stored in a state where they can be read by a processor circuit, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), a configuration hardware bitstream, an application programming interface (API), etc., in order to execute the instructions on a particular computing device or other device. In another example, the machine-readable instructions may need to be configured (e.g., settings stored, data entered, network addresses recorded, etc.) before the machine-readable instructions and / or corresponding program can be executed in whole or in part. Thus, machine-readable media, as used herein, may include machine-readable instructions and / or programs regardless of the particular format or state of the machine-readable instructions and / or programs when stored or otherwise stationary or in transit.

[0021] The machine-readable instructions described herein may be expressed in any past, present, or future command language, scripting language, programming language, etc. For example, the machine-readable instructions may be expressed using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0022] 4 may be implemented using executable instructions (e.g., computer and / or machine readable instructions) stored on a non-transitory computer and / or machine readable medium, such as a hard disk drive, flash memory, read-only memory, compact disc, digital versatile disc, cache, random access memory, and / or any other storage device or storage disk on which information is stored for any period of time (e.g., for an extended period of time, permanently, for a short period of time, for temporary buffering, and / or for caching of information). As used herein, the term non-transitory computer readable medium is expressly defined to include any type of computer readable storage device and / or storage disk, to exclude propagating signals, and to exclude transmission media.

[0023] The terms "including" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, whenever a claim employs any form of "including" or "comprises" (e.g., includes, includes, including, including, having, etc.) as a preamble or within any type of claim recitation, it is to be understood that additional elements, terms, etc. may be present without exceeding the scope of the corresponding claim or recitation. As used herein, the phrase "at least," when used as a transitional term, for example, in a claim preamble, is open-ended in the same way that the terms "including" and "comprising" are open-ended. The term "and / or," when used in the form, for example, A, B, and / or C, refers to any combination or subset of A, B, and C, for example, (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, (7) A, B, and C, etc. As used herein, when used in the context of describing a structure, component, item, object, and / or thing, the phrase “at least one of A and B” is intended to refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase “at least one of A or B” is intended to refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein, when used in the context of describing the performance or execution of a process, instruction, action, activity, and / or step, the phrase “at least one of A and B” is intended to refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.Similarly, herein, when used in the context of describing the performance or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to refer to implementations that include any of: (1) at least one A; (2) at least one B; and (3) at least one A and at least one B.

[0024] As used herein, singular references (e.g., "a," "an," "first," "second," etc.) do not exclude a plurality. The term "a" or "an" entity, when used herein, refers to one or more of that entity. The terms "a," "one or more," and "at least one" may be used interchangeably herein. Furthermore, although individually listed, a plurality of means, elements, or method actions may be implemented by, for example, a single unit or processor. Furthermore, although individual features may be included in different examples or claims, these may also be combined, and inclusion in different examples or claims does not imply that a combination of features is not feasible and / or advantageous.

[0025] Figure 2 is a block diagram of one example implementation of the example telemetry analyzer 110 of Figure 1. The example telemetry analyzer 110 of Figure 2 includes an example fault predictor 205, an example resolution handler 210, and an example impact trainer 215.

[0026] The example failure predictor 205 in the illustrated example of FIG. 2 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable logic devices (FPLDs), digital signal processors (DSPs), coarse-grained reconfigurable architectures (CGRAs), image signal processors (ISPs), etc. The example failure predictor 205 characterizes and predicts execution path outcomes and intervenes when a negative outcome or failure is predicted. In this example, data characterized and monitored by the example failure predictor 205 is stored in the example database 334 and later referenced by the example failure predictor 205 to determine execution path outcomes or device states. In some embodiments, the example failure predictor 205 references instructions 1032 stored in the example database 334 to reference parameters indicative of a faulty execution path outcome or parameters describing a device state. If a fault is predicted, the example fault predictor 205 outputs an interrupt to interrupt the execution path and sends the predicted fault metadata and associated parameters, including the profile, to the example resolution handler 210. In this example, the interrupt is a signal, but may also or instead be a flag, a data value, a register setting, etc.

[0027] 2 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, ASICs, PLDs, FPLDs, programmable controllers, GPUs, DSPs, CGRAs, ISPs, etc. The example resolution handler 210 determines the best resolution strategy for a predicted fault. If a given fault has a resolution strategy, the resolution handler 210 applies that resolution strategy. If the predicted fault does not have a resolution strategy, the resolution handler 210 assembles a list of resolution strategies from the example database 334, listed in ascending cost relative to system performance. The example resolution handler 210 applies the lowest-cost resolution strategy first; if that strategy is unsuccessful in changing the path's outcome from the fault, the example resolution handler 210 applies the next lowest-cost resolution strategy. The goal of the exemplary resolution handler 210 is to try all possible resolution strategies to change the prediction of an execution path from faulty to non-faulty.

[0028] The example impact trainer 215 in the illustrated example of FIG. 2 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, ASICs, PLDs, FPLDs, programmable controllers, GPUs, DSPs, CGRAs, ISPs, etc. The example impact trainer 215 notifies the client device of the results of one or more attempted resolution policies, whether or not they were successful in changing the predicted outcome, and stores impact data from the attempted policies in the example database 334. In the example disclosed herein, the impact data includes the results of each resolution strategy, how each strategy affected the predicted outcome of the execution path, metadata and profiles associated with each resolution strategy, and methods for integrating the results of applying the resolution strategies to future execution paths. Generally, the impact data is stored in the example database 334 to improve the system's prediction and intervention capabilities for future similar execution paths. After the impact data is reported and stored in the sample database 334, the example fault predictor clears the interrupt, and the system continues execution.

[0029] Although one example method for implementing the telemetry monitor of Figure 1 is illustrated in Figure 2, one or more of the elements, processes, and / or devices illustrated in Figure 2 may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other manner. Furthermore, the example fault predictor 205, the example resolution handler 210, the example impact trainer 215, and / or more generally the example telemetry monitor of Figure 1 may be implemented in hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, any of the example fault predictor 205, the example resolution handler 210, the example impact trainer 215, and / or more generally the telemetry analyzer 110 may be implemented in one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs). When reading any of the apparatus or system claims of this patent to cover purely software and / or firmware implementations, at least one of the example fault predictor 205, example solution handler 210, and example impact trainer 215 are expressly defined herein to include a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile disk (DVD), compact disk (CD), Blu-ray disk, etc., containing software and / or firmware. Furthermore, the example telemetry analyzer 110 of FIG. 1 may include one or more elements, processes, and / or devices in addition to or instead of those illustrated in FIG. 2, and / or may include more than one of any or all of the illustrated elements, processes, and devices.As used herein, the phrase "in communication," including variations thereof, encompasses direct communication and / or indirect communication via one or more intermediary components, and further includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events that do not require direct physical (e.g., wired) communication and / or constant communication, but rather.

[0030] Figure 3 is a block diagram of an example implementation of the example fault predictor 205 of Figure 2. The example fault predictor 205 of Figure 3 includes an example sampling tuner 305, an example profile extractor 310, and an example database 334.

[0031] 3 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, ASICs, PLDs, FPLDs, programmable controllers, GPUs, DSPs, CGRAs, ISPs, etc. The example sampling tuner 305 determines an optimized rate for data collection and characterization to ensure all telemetry variables are observable. By analyzing the changing variables and the velocity of the sampling data, the example sampling tuner 305 generates a sampling rate used by the example fault predictor 205 to further sample and profile the telemetry data.

[0032] The example profile extractor 310 in the illustrated example of FIG. 3 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, ASICs, PLDs, FPLDs, programmable controllers, GPUs, DSPs, CGRAs, ISPs, etc. The example profile extractor 310 profiles the metadata through the use of fractal similarity searches and extracts data profiles to predict the outcome of execution paths. The example profile extractor searches the example database 334 for similar profiles and subsequences that match the recorded telemetry data points, creating a growing database of profiles and subsequences. This growing database improves the predictive capabilities of the example fault predictor 205. In this example, the example profile extractor 310 uses fractal similarity searches to match existing profiles to the extracted profiles and refine the existing profiles. The example profile extractor 310 continuously extracts data profiles from the metadata that are used to characterize the data performed by the overall process of the example failure predictor 205 .

[0033] 3 is implemented with logic circuitry, such as a hardware processor. However, any other type of circuitry may additionally or alternatively be used, such as one or more analog or digital circuits, logic circuits, programmable processors, ASICs, PLDs, FPLDs, programmable controllers, GPUs, DSPs, CGRAs, ISPs, etc. The example fault interface 315 determines whether a system fault has occurred and, if a fault has occurred, initiates triage to generate parameters for the fault condition. In some examples, the fault interface 315 reports the generated parameters to a backend server for triage analysis.

[0034] The exemplary database 334 in the illustrated example of FIG. 3 is implemented by any memory, storage device, and / or storage disk for storing data, such as, for example, flash memory, magnetic media, optical media, solid-state memory, hard drives, thumb drives, etc. Furthermore, the data stored in the exemplary database 334 may be in any data format, such as, for example, binary data, comma-separated data, tab-separated data, Structured Query Language (SQL) structures, binary hardware aging / health characterization bitstreams, etc. In the illustrated example, the database 334 is shown as a single device, but the exemplary database 334 and / or any other data storage devices described herein may be embodied with any number and / or type of memory. In the illustrated example of FIG. 3, the exemplary database 334 stores metadata profiles, signatures, collected telemetry data, resolution strategies, and / or impact data.

[0035] 2 is illustrated in FIG. 3, one or more of the elements, processes, and / or devices illustrated in FIG. 3 may be combined, divided, rearranged, omitted, deleted, and / or implemented in any other manner. Furthermore, the example sampling tuner 305, the example profile extractor 310, the example database 334, and / or more generally the example failure predictor 205 of FIG. 2 may be implemented in hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, any of the example sampling tuner 305, the example profile extractor 310, the example database 334, and / or more generally the example failure predictor 205 may be implemented in one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs). When reading any of the apparatus or system claims of this patent to cover purely software and / or firmware implementations, at least one of the example sampling tuner 305, the example profile extractor 310, and the example database 334 are expressly defined herein to include a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile disk (DVD), compact disk (CD), Blu-ray disk, etc., containing software and / or firmware. Furthermore, the example failure predictor of FIG. 2 may include one or more elements, processes, and / or devices in addition to or instead of those illustrated in FIG. 3, and / or may include more than one of any or all of the illustrated elements, processes, and devices.As used herein, the phrase "in communication," including variations thereof, encompasses direct communication and / or indirect communication via one or more intermediary components, and further includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events that do not require direct physical (e.g., wired) communication and / or constant communication, but rather.

[0036] Flowcharts representing example hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the example telemetry analyzer 110 of FIGS. 1-3 are shown in FIGS. 4-7E. The machine-readable instructions may be one or more executable programs or portions of executable programs for execution by a computer processor and / or processor circuitry, such as the processor 1012 shown in the example processor platform 1000 discussed below in connection with FIGS. 10 and 11. The programs may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processor 1012, although the entire program and / or portions thereof may alternatively be executed by a device other than the processor 1012 and / or embodied in firmware or dedicated hardware. Additionally, although the example programs are described with reference to the flowcharts shown in FIGS. 4-7E, many other methods of implementing the example telemetry analyzer 110 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the described blocks may be modified, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware. Processor circuits may be distributed across different network locations and / or local to one or more devices (e.g., multi-core processors in a single machine, multiple processors distributed across a server rack, etc.).

[0037] The machine-readable instructions described herein may be stored in one or more of a compressed format, an encrypted format, a fragmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions described herein may also be stored as data or data structures (e.g., portions of instructions, code, representations of code, etc.) that can be utilized to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be fragmented and stored on one or more storage devices and / or computing devices (e.g., servers) located in the same or different locations of a network or collection of networks (e.g., in the cloud, in edge devices, etc.). The machine-readable instructions may require one or more of installing, modifying, adapting, updating, combining, supplementing, configuring, decrypting, decompressing, unpacking, distributing, reassigning, compiling, etc. to make them directly readable, interpretable, and / or executable by computing devices and / or other machines. For example, machine-readable instructions may be stored in multiple portions that are individually compressed, encrypted, and stored on separate computing devices, and that when decoded, decompressed, and combined form a set of executable instructions that implement one or more functions that may together form a program as described herein.

[0038] In another example, machine-readable instructions may be stored in a state where they can be read by a processor circuit, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., in order to execute the instructions on a particular computing device or other device. In another example, the machine-readable instructions may need to be configured (e.g., settings stored, data entered, network addresses recorded, etc.) before the machine-readable instructions and / or corresponding program can be executed in whole or in part. Thus, as used herein, machine-readable media may include machine-readable instructions and / or programs regardless of the particular format or state of the machine-readable instructions and / or programs when stored or otherwise stationary or in transit.

[0039] The machine-readable instructions described herein may be expressed in any past, present, or future command language, scripting language, programming language, etc. For example, the machine-readable instructions may be expressed using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0040] 4-7E may be implemented using executable instructions (e.g., computer- and / or machine-readable instructions) stored on a non-transitory computer- and / or machine-readable medium, such as a hard disk drive, flash memory, read-only memory, compact disc, digital versatile disc, cache, random access memory, and / or any other storage device or disk on which information is stored for any period of time (e.g., for an extended period of time, permanently, for a short period of time, for temporary buffering, and / or for caching of information). As used herein, the term non-transitory computer-readable medium is expressly defined to include any type of computer-readable storage device and / or disk, to exclude propagating signals, and to exclude transmission media.

[0041] The terms "including" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, whenever a claim employs any form of "including" or "comprises" (e.g., includes, includes, including, including, having, etc.) as a preamble or within any type of claim recitation, it is to be understood that additional elements, terms, etc. may be present without exceeding the scope of the corresponding claim or recitation. As used herein, the phrase "at least," when used as a transitional term, for example, in a claim preamble, is open-ended in the same way that the terms "including" and "comprising" are open-ended. The term "and / or," when used in the form, for example, A, B, and / or C, refers to any combination or subset of A, B, and C, for example, (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, (7) A, B, and C, etc. As used herein, when used in the context of describing a structure, component, item, object, and / or thing, the phrase “at least one of A and B” is intended to refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase “at least one of A or B” is intended to refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein, when used in the context of describing the performance or execution of a process, instruction, action, activity, and / or step, the phrase “at least one of A and B” is intended to refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.Similarly, herein, when used in the context of describing the performance or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to refer to implementations that include any of: (1) at least one A; (2) at least one B; and (3) at least one A and at least one B.

[0042] As used herein, singular references (e.g., "a," "an," "first," "second," etc.) do not exclude a plurality. The term "a" or "an" entity, when used herein, refers to one or more of that entity. The terms "a," "one or more," and "at least one" may be used interchangeably herein. Furthermore, although individually listed, a plurality of means, elements, or method actions may be implemented by, for example, a single unit or processor. Furthermore, although individual features may be included in different examples or claims, these may also be combined, and inclusion in different examples or claims does not imply that a combination of features is not feasible and / or advantageous.

[0043] 4 is a flowchart representing example machine-readable instructions that may be executed to implement the example telemetry analyzer 110 to continuously monitor telemetry. The example process 400 of FIG. 4 begins when the example failure predictor 205 predicts a failure and drives an interrupt (block 405). In this example, the failure predictor 205 predicts the outcome of an execution path by referencing instructions in the example database 334 that indicate a trace of previous failures or negative outcomes. These instructions may include time-series system data, application data, or any other type of data about the client device that can be monitored and referenced to match a current execution path with an execution path that led to a failure or negative outcome.

[0044] Next, the example fault predictor 205 provides a set of control parameters to the example resolution handler 210 based on the predicted outcome of the execution path (block 410). The parameters include, but are not limited to, a resolution policy and a vector of predicted fault metadata and signatures. More generally, these parameters are tailored to the current execution path of the system. The example resolution handler 210 receives these parameters and determines whether a resolution strategy exists for the provided set of parameters (block 415). If the resolution handler 210 determines that a resolution strategy does not exist for the provided set of parameters (e.g., block 415 returns a NO result), the resolution handler 210 creates a ranked list of resolution strategies in order of ascending performance cost (block 420). In this example, the ranked list is in order of ascending performance cost. However, any other method of ranking the resolution strategies may also or instead be used. After ranking the resolution strategies, the resolution handler 210 applies the first resolution strategy in the ranked list (block 425). If the resolution handler 210 determines that a resolution strategy exists for the set of parameters (eg, block 415 returns a YES result), the resolution handler 210 applies the existing resolution strategy (block 425).

[0045] After the resolution handler 210 applies the corresponding resolution strategy, the example impact trainer 215 determines whether the predicted outcome of the execution path has changed (block 430). If the predicted outcome of the execution path has changed (e.g., block 430 returns a YES result), the impact trainer 215 notifies the client of the state change and saves the impact data (block 440). If the predicted outcome of the execution path has not changed (e.g., block 430 returns a NO result), the impact trainer 215 determines whether all resolution strategies have been applied (block 435). If there are resolution strategies that have not yet been applied (e.g., block 435 returns a NO result), the example resolution handler 210 selects and applies the next resolution strategy in the ranked list. If all resolution strategies have been applied (e.g., block 435 returns a YES result), the impact trainer 215 notifies the client of the state change and saves the impact data (block 440). In some examples, the impact data links a resolution strategy to a signature match. In response to the example impact trainer 215 notifying the client of the state change and saving the impact data, the example failure predictor 205 clears the interrupt (block 445).

[0046] 5 is a flowchart illustrating example machine-readable instructions 500 that, when executed, implement the failure predictor 205 to characterize telemetry data. The example process 500 of the illustrated example of FIG. 5 begins when the example failure predictor 205 is initialized to begin collecting data (block 505).

[0047] The example fault predictor 205 then waits for activity (block 510). Upon detecting telemetry activity, the example sampling tuner 305 is activated (block 515). In the examples disclosed herein, activating the example sampling tuner 305 initiates an observation phase to understand changing variables, speeds, and appropriate frequencies for data collection. In some examples, the sampling tuner 305 is activated upon request from a client device.

[0048] The example profile extractor 310 is then launched (block 520). In the examples disclosed herein, launching the example profile extractor 310 initiates execution of a metadata profile to profile the collected data. In some examples, the profile extractor 310 is launched by a request from a client device. The example profile extractor 310 begins sampling and recording telemetry data at a sampling frequency determined by the sampling tuner 305 (block 525). In the examples disclosed herein, the sampling frequency is determined by the sampling tuner 305. However, any other approach for determining the sampling frequency may also or instead be used. In the examples disclosed herein, the sampled telemetry data may be a given data object or multiple data objects.

[0049] The example profile extractor 310 extracts a telemetry profile, which includes a metadata profile of the sampled telemetry data (block 530). After the metadata profile is extracted, execution of the profile extractor 310 is terminated (block 535). The example fault interface 315 then determines whether a system fault has occurred (block 540). In some examples, a system header check of the telemetry payload is used to understand if the device is in a failure mode. In some examples, the telemetry may also send a controller-initiated event with the failure mode that will be used to determine the fault. If a fault has occurred (e.g., block 540 returns a YES result), the fault interface 315 initiates triage and generates parameters for the fault condition (block 545). The fault predictor 205 then returns to block 510 and waits for activity. If a fault has not occurred (e.g., block 540 returns a NO result), the fault predictor 205 returns to block 510 and waits for activity.

[0050] Figure 6 is a flowchart depicting example machine-readable instructions 600 that, when executed, implement the example sampling tuner 305 to initiate an observation phase to understand changing variables, speeds, and / or appropriate frequencies for data collection. Figure 6 is an example process for implementing block 515 of Figure 5. The example process 600 of the illustrated example of Figure 6 begins when the example sampling tuner 305 is activated (block 605).

[0051] The sampling tuner 305 then reviews the device configuration and capabilities and determines the samples required for the reliable population (block 610). The sampling tuner 305 selects a sampling rate based on the device configuration and capabilities and the number of samples required for the reliable population (block 615).

[0052] The example sampling tuner 305 then samples the telemetry metadata (block 620). In this example, the extracted telemetry data can be for either a given object or multiple objects. The sampling tuner 305 then waits for a sample cadence (block 625). In this example, the example sampling tuner 305 enters a thread idle sleep for the event time. The example sampling tuner then determines whether there are enough data points for each object (block 630). To evaluate all data points for data extraction as shown in FIGS. 7A-7E, the example sampling tuner 305 requires enough unique sample data points for each object. If the sampling tuner 305 determines there are enough unique data points for each object (e.g., block 630 returns a YES result), the sampling tuner 305 measures the Nyquist frequency distance of the object or set of objects (block 635). If the sampling tuner 305 determines that there are not enough unique data points for each object (e.g., block 630 returns a NO result), the sampling tuner 305 returns to block 620 to extract more telemetry data.

[0053] Once the sampling tuner 305 calculates the Nyquist frequency distance, the sampling tuner 305 determines the required sampling rate change (block 640). If the sampling tuner 305 determines that the distance between the objects is zero (e.g., block 640 returns a YES result), the sampling tuner 305 records the sampling frequency for the objects (block 650). The example sampling tuner 305 then doubles the frequency to decrease the sampling rate of the data object observations (block 655). If the sampling tuner 305 determines that the distance between the objects is not zero (e.g., block 640 returns a NO result), the sampling tuner 305 halves the sampling frequency to increase the sample rate of the data object observations (block 645).

[0054] The example sampling tuner 305 determines whether the sampling frequency has been recorded (block 660). If the sampling tuner 305 has not recorded the sampling frequency (e.g., block 660 returns a NO result), the sampling tuner 305 returns to block 620 and continues sampling telemetry metadata. If the example sampling tuner 305 has recorded the sampling frequency (e.g., block 660 returns a YES result), the sampling tuner 305 records the sample in a population list (block 665).

[0055] The example sampling tuner 305 then determines whether there are enough samples to determine a sampling rate within the confidence interval (block 670). If the sampling tuner has enough samples for confidence (e.g., block 670 returns a YES result), the sampling tuner 305 aggregates the rate required to observe the data object, and the appropriate sampling frequency is relayed to the example fault predictor 205 (block 675). If the sampling tuner 305 does not have enough samples for confidence (e.g., block 670 returns a NO result), the sampling tuner 305 returns to block 615 and selects another random sampling cadence.

[0056] 7A-7E are flowcharts illustrating example machine-readable instructions 700 that, when executed, implement the profile extractor 310 to initiate metadata profiling and device state machine learning. Figures 7A-7E are an example process for implementing block 520 of Figure 5. The example process 700 of the illustrated example of Figures 7A-7E begins when the example failure predictor 205 invokes the example profile extractor 310 (block 702).

[0057] The profile extractor 310 then loads requirement thresholds for observation of one or more sets of object containers (block 704). In the examples disclosed herein, the object containers contain one or more data objects in an encapsulated format. The requirement thresholds are calculated based on a sampling rate determined by the example sampling tuner 305. The profile extractor 310 then marks a sample start window indicating the period during which sampling should begin (block 706).

[0058] The profile extractor 310 then begins recording telemetry data (block 708) that is stored in the example database 334. In this example, the telemetry data includes snapshots of system metadata and time-series telemetry data objects. These data points are collected in linked object containers within the example database 334. The example profile extractor 310 then determines if a change in condition or distance has occurred (block 710).

[0059] If the example profile extractor 310 determines that no state change has occurred (e.g., block 710 returns a NO result), the profile extractor 310 compresses the window sample range (block 712). Compressing the window sample range period indicates the beginning and end of a range of consecutive values. For a given time-series data stream, data processing is optimized by using a matrix profile to extract unique sequence-to-sequence signatures. Because the velocity of data structures varies, data is collected at a frequency adjusted to the highest rate data and then resampled to a lower rate data. For example, thermal element data changes at a rate of 1 / 32 of a second, while device telemetry snapshot (NVMe) queues operate at 1 / 1.6 millionth of a second. Due to the dramatic difference in frequency, metadata is collected on an on-demand basis, and then the data range of recurring values ​​is compressed to reduce the storage footprint of telemetry data collection. Compressions of these signatures are encoded into a fractal database, so that events can be agnostic compared to these repeating sequence patterns through dynamic time warping. These compressed window representations mean that minimal sampled data is obtained for each data signature referencing a late- or early-occurring event, so that predictions can be projected to precise time event indices or intervals into the future. These predictions and projections are refined and accurate as more data is collected, allowing the artificial neural network to learn statistical variances so that the dimensions of a given projection are within the range of the magnitude of the observed events. The profile extractor 310 then waits for data to be recorded (block 714). The profile extractor 310 then returns to block 708 to record more telemetry data.

[0060] If the example profile extractor 310 determines that a state change has occurred (e.g., block 710 returns a YES result), the profile extractor 310 extracts a matrix profile (block 716). In doing so, the example profile extractor 310 uniquely identifies the current time series signature of the linked container. To match events, a distance matrix of all subsequence pairs of length is constructed, and the pairs are projected onto a vector down the smallest off-diagonal value. In some examples, the matrix profile is this vector. The profile extractor 310 then searches the example database 334 for similar profiles (block 718). In this example, the example database 334 contains a ranked set of profile (or query) matches, each with a quantified similarity measure to the extracted profile.

[0061] The example profile extractor 310 then determines whether the example database 334 contains any similar profiles (block 720). If the example database 334 does not contain similar profiles (e.g., block 720 returns a NO result), the profile extractor 310 determines whether the extracted profile is a subsequence (block 722). If the profile extractor 310 determines that the extracted profile is not a subsequence (e.g., block 722 returns a NO result), the extracted profile is added to the example database 334 (block 724). If the profile extractor 310 determines that the extracted profile is a subsequence (e.g., block 722 returns a YES result), the profile extractor 310 determines whether a state path exists for the profile (block 726). If a state path exists for the extracted profile (e.g., block 726 returns a YES result), the example profile extractor 310 performs windowed fractal expansion (block 738). Window fractal expansion involves expanding or compressing the profile set to a desired dimension and representing the chaotic factors as a compression or expansion of the self-similarity dimension. If the chaotic factors (e.g., roughness, irregularity, etc.) were not represented, the expansion or compression of the profile set would repeat infinitely. Therefore, the chaotic factors are an essential part of the window fractal expansion. Furthermore, statistical data facilitates identifying compressions and expansions of the metadata window so that the relativity of observed events to the rate of change in the metadata is maintained. These statistically characterized events create a series of fitting equations with various coefficient factors so that the basis is maintained in the core algorithm characterization. The profile extractor 310 then returns to block 714 and waits to record telemetry data. If no state path exists for the extracted profile (e.g., block 726 returns a NO result), the profile extractor 310 records the entry path of the extracted profile (block 728).The profile extractor 310 then continues to block 738 and performs a window fractal expansion.

[0062] If the example database 334 contains similar profiles (e.g., block 720 returns a YES result), the example profile extractor 310 determines whether the similar profiles have a common subsequence (block 730). If the extracted profile and the similar profile have a common subsequence (e.g., block 730 returns a YES result), the profile extractor 310 adds the extracted profile to the database 334 in the previous state tree (block 732). Additionally, the example profile extractor 310 determines whether a state path exists (block 726).

[0063] If the extracted profile and the similar profiles do not have a common subsequence (e.g., block 730 returns a NO result), the example profile extractor 310 determines whether there are enough samples in the extracted profile to ensure an adequate representation of the variance (block 740). If there are enough samples in the extracted profile (e.g., block 740 returns a YES result), the profile extractor 310 returns to block 706 to indicate a new period of sampling. If there are not enough samples in the extracted profile (e.g., block 740 returns a NO result), the profile extractor 310 adds additional profiles to the extracted profile (block 742). The example profile extractor 310 continues to block 732 and adds the extracted profile, along with the additional profiles, to the database 334 in the previous state tree.

[0064] The example profile extractor 310 then re-indexes the interval profiles to introduce the new data points (block 734). The profile extractor 310 then re-clusters the database 334 to balance the data structure for performance access based on the new data points (block 736). The reclustering technique implements a machine learning technique that clusters states and entry paths for similar profiles. The profile extractor 310 then returns to block 738 to perform window fractal expansion. According to the illustrated example, the process 700 of FIGS. 7A-7E is then continuously repeated. Alternatively, the process 700 may terminate after a singleton set or multiple sets of interactions.

[0065] 8 is a flowchart depicting example machine-readable instructions 800 that, when executed, implement the profile extractor 310 to encode profiles in the example database 334. The profiles in the database are encoded in a manner that reduces the necessary storage required for the database. The example process 800 of the illustrated example FIG. 8 begins when the profile extractor 310 extracts an example profile (block 802).

[0066] The profile extractor 310 converts the example profile into a list indicating the frequency of each item in the example profile (block 804). The example profile extractor 310 then combines two items in the profile to form a string of the two items (block 806). In this example, the profile extractor combines the two items with the lowest frequency of occurrence in the example profile. In this example, the string generated by the profile extractor 310 is "CB." The string "CB" has a state tree containing two branches (or fractals) indicating the frequency of each item in the profile (e.g., C:2, B:6). Each branch further has a binary digit to distinguish between the two branches.

[0067] The example profile extractor 310 then generates a new list that includes the remaining items in the profile and the string generated in block 804. The profile extractor 310 generates a new string and updates the state tree to include a new branch associated with the new string (block 808). The profile extractor 310 repeats this string generation process (blocks 810-814) until all items in the profile are included in the string.

[0068] The example profile extractor 310 then retrieves a binary representation of each item in the string from the state tree (block 816). The example profile extractor then stores the example profile in the example database 334. Once stored in the database, the state tree of the profile is easily searchable through a fractal similarity search.

[0069] FIG. 9 is a block diagram of an exemplary fractal similarity search query 900 according to the machine-readable instructions of FIGS. 7A-7E. FIG. 9 further illustrates the training of a time-series recurrent neural network (RNN). The RNN layer inputs an exemplary input sequence 902 to an exemplary long short-term memory (LSTM) 904 at each time step and outputs exemplary variables 908 as input to the LSTM for the next time step. The variables 908 are also fed to an exemplary softmax 906, which returns a vector representing a probability distribution of potential outcomes. This vector is fed to the LSTM for the next time step. In this example, the LSTM is functioning as an RNN. Using a data input stream from the exemplary client device 102, a database query, and a coding tree such as that mentioned in FIG. 8, the RNN is trained to reproduce the fractal map. The RNN is trained on the full stream of data (e.g., Q = Q(1) + Q(2) + Q(3) + Q(4) + Q(5)), and the time series fractal is removed (e.g., now Q = Q(1) + Q(2) + Q(3) + Q(4), and Q(5) is removed), completing the RNN training. The predictive RNN is trained by selecting a random sequence input as the starting point and predicting the next step using the steps shown in Figures 7A-7E to determine the full profile of the input. An LSTM and max Softmax are used to bound the variance of the sequence, and a generative fractal similarity search query based on the input fractal vector is created. In general, the system can predict a full data profile using only partial data input from any point within that profile. Furthermore, for focused event identification, an autoencoder and decoder are used to reduce the latent space of vectorized variables (approximately 1.8 million in vector width) on a large time series metadata set.

[0070] Once the RNN contains a baseline fractal (a state-space fractal of positive outcomes), it is trained to understand the entire state space of possibilities by forcing failures in the baseline states at various rates. In this example, Euclidian Distance, Pearson's Correlation, and Dynamic Time Warping were used as similarity search engines. After full RNN training, the system was able to predict failures from a random test dataset with 92% accuracy. Based on the size and quality of the metadata, there was up to 98.3% accuracy in component characterization.

[0071] 10 is a block diagram of an exemplary processor platform 1000 configured to execute the instructions of FIGS. 4-7E to implement the telemetry analyzer 110 of FIGS. 1, 2, and / or 3. The processor platform 1000 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., neural network), a mobile device (e.g., cell phone, smartphone, iPad®), or the like. TM The mobile device may be a mobile computer, a tablet, a personal digital assistant (PDA), an internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a gaming console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.

[0072] The processor platform 1000 of the illustrated example includes a processor 1012. The processor 1012 of the illustrated example is hardware. For example, the processor 1012 can be implemented with one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor implements the example fault predictor 205, the example resolution handler 210, and the example impact trainer 215.

[0073] The processor 1012 of the illustrated example includes local memory 1013 (e.g., cache). The processor 1012 of the illustrated example communicates with main memory, including volatile memory 1014 and non-volatile memory 1016, via a bus 1018. The volatile memory 1014 may be implemented with synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The non-volatile memory 1016 may be implemented with flash memory and / or any other desired type of memory device. Access to the main memory 1014, 1016 is controlled by a memory controller.

[0074] The processor platform 1000 of the illustrated example further includes an interface circuit 1020. The interface circuit 1020 may be implemented with any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth® interface, a near field communication (NFC) interface, and / or a PCI Express interface.

[0075] In the depicted example, one or more input devices 1022 are coupled to the interface circuit 1020. The input devices 1022 allow a user to enter data and / or commands into the processor 1012. The input devices may be implemented by, for example, a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.

[0076] Additionally, one or more output devices 1024 are connected to the interface circuitry 1020 of the illustrated example. The output device(s) 1024 may be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Accordingly, the interface circuitry 1020 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.

[0077] The interface circuitry 1020 of the depicted example further includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate the exchange of data with external machines (e.g., computing devices of any type) over the network 1005. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-site wireless system, a cellular telephone system, etc.

[0078] The processor platform 1000 of the depicted example further includes one or more mass storage devices 1028 for storing software and / or data. Examples of such mass storage devices 1028 include a floppy disk drive, a hard drive disk, a compact disk drive, a Blu-ray disk drive, a redundant array of independent disks (RAID) system, and a digital versatile disk (DVD) drive. In this example, the exemplary mass storage device 1028 includes the exemplary database 334. However, the exemplary database 334 may also be included in the exemplary volatile memory 1014, the exemplary non-volatile memory 1016, and / or on a removable, non-transitory computer-readable storage medium such as a CD or DVD.

[0079] 4-7E may be stored in mass storage device 1028, in volatile memory 1014, in non-volatile memory 1016, in local memory 1013, in database 33, and / or on a removable non-transitory computer-readable storage medium such as a CD or DVD. Additionally, coded instructions 1032 may correspond to one or more elements for implementing the example telemetry monitor tool 115 described above.

[0080] FIG. 11 shows a block diagram illustrating an example software distribution platform 1105 for distributing software, such as the example computer-readable instructions 1032 of FIG. 10, to third parties. The example software distribution platform 1105 can be implemented by any computer server, data facility, cloud service, etc., capable of storing and transmitting software to other computing devices. The third parties may be customers of the entity that owns and / or operates the software distribution platform. For example, the entity that owns and / or operates the software distribution platform may be a developer, seller, and / or licensor of software, such as the example computer-readable instructions 1032 of FIG. 10. The third parties may also be consumers, users, retailers, OEMs, etc., that purchase and / or license software for use and / or resale and / or sublicense. In the illustrated example, the software distribution platform 205 includes one or more servers and one or more storage devices. The storage devices store computer-readable instructions 1032, which may correspond to the example computer-readable instructions 115 of FIG. 1 and / or FIG. 2, as described above. One or more servers of the example software distribution platform 1105 communicate with a network 1110, which may correspond to the Internet and / or any one or more of the example networks 105 and / or 1005 described above. In some examples, the one or more servers respond to requests to send software to requesters as part of a commercial transaction. Payment for the distribution, sale, and / or license of the software can be handled by one or more servers of the software distribution platform and / or through a third-party payment entity. The servers enable purchasers and / or licensors to download computer-readable instructions 1032 from the software distribution platform 1105.For example, software that may correspond to the example computer-readable instructions 115 of Figures 1 and / or 2 may be downloaded to the example processor platform 1000, which executes the computer-readable instructions 1032 to implement the example telemetry analyzer 110. In some examples, one or more servers of the software distribution platform 1105 periodically provide, transmit, and / or force updates to the software (e.g., the example computer-readable instructions 1032 of Figure 10) to ensure that improvements, patches, updates, etc. are distributed and applied to the software at end-user devices.

[0081] From the foregoing, it will be appreciated that exemplary methods, apparatus, and articles of manufacture are disclosed that enable continuous characterization of execution paths, prediction of their outcomes, and intervention methods to prevent negative outcomes. The disclosed methods, apparatus, and articles of manufacture improve the efficiency of using computing devices by proactively predicting and intervening on paths of execution that lead to negative outcomes. Furthermore, systems that deploy this tool increase their overall efficiency through machine learning of new or improved intervention techniques to prevent these negative outcomes. Thus, the disclosed methods, apparatus, and articles of manufacture are directed to one or more improvements in computer functioning.

[0082] Example 1 includes an apparatus for monitoring telemetry in a computing environment, the apparatus including a failure predictor that predicts an outcome of an execution path, a resolution handler that determines a resolution strategy for the execution path and applies the resolution strategy, and an impact trainer that determines whether the predicted outcome of the execution path has changed and stores impact data of the applied resolution strategy.

[0083] Example 2 includes the apparatus of Example 1, wherein the failure predictor is further configured to: drive an interrupt in response to predicting the outcome of the execution path to be faulty; provide a control parameter to the resolution handler; and clear the interrupt in response to the predicted outcome of the execution path being no longer faulty.

[0084] Example 3 includes the apparatus of Example 1, wherein the resolution handler is further configured to: in response to determining that the resolution strategy exists for the execution path, apply the resolution strategy to the execution path; and in response to determining that the resolution strategy does not exist for the execution path, create a resolution strategy list containing resolution strategies in ascending order of system performance cost and apply a first resolution strategy from the list.

[0085] Example 4 includes the apparatus of Example 3, wherein the impact trainer is further configured to: apply a next solution strategy in the solution strategy list in response to determining that the predicted outcome of the execution path has not changed and all solution strategies from the solution strategy list have not been tried; relay impact data of the solution strategy to the fault predictor in response to determining that the predicted outcome of the execution path has not changed and all solution strategies from the solution strategy list have been tried; and relay impact data of the solution strategy to the fault predictor in response to determining that the predicted outcome of the execution path has changed.

[0086] Example 5 includes the apparatus of example 1, wherein the failure predictor further includes a sampling tuner that determines an appropriate frequency for data collection.

[0087] Example 6 includes the apparatus of example 1, wherein the fault predictor further includes a profile extractor that extracts and refines profiles to predict outcomes of the execution paths.

[0088] Example 7 includes the apparatus of example 6, wherein the profile extractor extracts and refines profiles using a fractal similarity search.

[0089] Example 8 includes a non-transitory computer-readable medium including instructions that, when executed, cause at least one processor to at least predict an outcome of an execution path, determine a resolution strategy for the execution path, apply the resolution strategy, determine whether the predicted outcome of the execution path has changed, and store impact data of the applied resolution strategy.

[0090] Example 9 includes the non-transitory computer-readable medium of Example 8, wherein the instructions, when executed, cause the at least one processor to: drive an interrupt and provide a control parameter in response to predicting an outcome of the execution path to be faulty; and clear the interrupt in response to the predicted outcome of the execution path being no longer faulty.

[0091] Example 10 includes the non-transitory computer-readable medium of Example 8, wherein the instructions, when executed, cause the at least one processor to, in response to determining that the resolution strategy exists for the execution path, apply the resolution strategy to the execution path, and, in response to determining that the resolution strategy does not exist for the execution path, create a resolution strategy list that includes resolution strategies in ascending order of system performance cost and apply a first resolution strategy from the list.

[0092] Example 11 includes the non-transitory computer-readable medium of Example 10, wherein the instructions, when executed, cause the at least one processor to: apply a next resolution strategy in the resolution strategy list in response to determining that the predicted outcome of the execution path has not changed and all resolution strategies from the resolution strategy list have not been tried; relay impact data of the resolution strategy in response to determining that the predicted outcome of the execution path has not changed and all resolution strategies from the resolution strategy list have been tried; and relay impact data of the resolution strategy in response to determining that the predicted outcome of the execution path has changed.

[0093] Example 12 includes the non-transitory computer-readable medium of example 8, wherein the instructions, when executed, cause the at least one processor to determine an appropriate frequency for data collection.

[0094] Example 13 includes the non-transitory computer-readable medium of example 8, wherein the instructions, when executed, cause the at least one processor to extract and refine a profile to predict outcomes of the execution paths.

[0095] Example 14 includes the non-transitory computer-readable medium of example 13, wherein the instructions, when executed, cause the at least one processor to extract and refine profiles using fractal similarity searching.

[0096] Example 15 includes a method including predicting an outcome of an execution path, determining a resolution strategy for the execution path, applying the resolution strategy, determining whether the predicted outcome of the execution path has changed, and storing impact data of the applied resolution strategy.

[0097] Example 16 includes the method of example 15, further including: in response to predicting an outcome of the execution path to be faulty, driving an interrupt and providing a control parameter; and in response to the predicted outcome of the execution path being no longer faulty, clearing the interrupt.

[0098] Example 17 includes the method of example 15, further including: in response to determining that the resolution strategy exists for the execution path, applying the resolution strategy to the execution path; and in response to determining that the resolution strategy does not exist for the execution path, creating a resolution strategy list that includes resolution strategies in ascending order of system performance cost and applying a first resolution strategy from the list.

[0099] Example 18 includes the method of Example 17, further including: in response to determining that the predicted outcome of the execution path has not changed and all resolution strategies from the resolution strategy list have not been tried, applying a next resolution strategy in the resolution strategy list; in response to determining that the predicted outcome of the execution path has not changed and all resolution strategies from the resolution strategy list have been tried, relaying impact data of the resolution strategy; and in response to determining that the predicted outcome of the execution path has changed, relaying impact data of the resolution strategy.

[0100] Example 19 includes the method of example 15, further including determining an appropriate frequency for data collection.

[0101] Example 20 includes the method of example 15, further including extracting and refining profiles to predict outcomes of the execution paths using a fractal similarity search.

[0102] Although certain exemplary methods, apparatus, and articles of manufacture are disclosed herein, the scope of this patent coverage is not limited thereto. On the contrary, this patent covers all methods, apparatus, and articles of manufacture that fall within the scope of the claims of this patent.

[0103] The following claims are hereby incorporated by reference into this detailed description, with each claim standing on its own as a separate embodiment of this disclosure.

Claims

1. 1. An apparatus for monitoring telemetry in a computing environment, comprising: a failure predictor that predicts the outcome of an execution path from the monitored telemetry; determining a resolution strategy for said execution path; Apply the resolution strategy A resolution handler; determining whether the predicted outcome of the execution path has changed; storing impact data of the applied resolution strategy; an impact trainer; An apparatus comprising:

2. 2. The apparatus of claim 1, wherein the fault predictor is further configured to drive an interrupt and provide control parameters to the resolution handler in response to predicting the outcome of the execution path to be faulty.

3. The apparatus of claim 2 , wherein the fault predictor is further configured to clear the interrupt in response to the predicted outcome of the execution path being no longer faulty.

4. The apparatus of claim 1 , wherein the resolution handler is further configured to, in response to determining that the resolution strategy exists for the execution path, apply the resolution strategy to the execution path.

5. 5. The apparatus of claim 1, wherein the resolution handler is further configured to, in response to determining that the resolution strategy does not exist for the execution path, create a resolution strategy list containing resolution strategies in ascending order of system performance cost and apply a first resolution strategy from the list.

6. 6. The apparatus of claim 5, wherein the influence trainer is further configured to apply a next solving strategy in the solving strategy list in response to determining that the predicted outcome of the execution path has not changed and all solving strategies from the solving strategy list have not been tried.

7. The apparatus of claim 5 or claim 6, wherein the impact trainer further relays impact data of the resolution strategies to the fault predictor in response to determining that the predicted outcome of the execution path has not changed and that all resolution strategies from the resolution strategy list have been tried.

8. 8. The apparatus of claim 1, wherein the impact trainer is further configured to relay impact data of the resolution strategy to the fault predictor in response to determining that the predicted outcome of the execution path has changed.

9. 9. The apparatus of claim 1, wherein the failure predictor further comprises a sampling tuner that determines an appropriate frequency for data collection.

10. The apparatus of claim 1 , wherein the fault predictor further comprises a profile extractor for extracting and refining profiles to predict the outcome of the execution paths.

11. The apparatus of claim 10 , wherein the profile extractor extracts and refines profiles using a fractal similarity search.

12. 1. A computer program comprising instructions that, when executed, cause at least one processor to: Predict the outcome of execution paths from monitored telemetry, determining a resolution strategy for said execution path; applying said resolution strategy; determining whether the predicted outcome of the execution path has changed; storing impact data of the applied resolution strategy; A computer program that makes things happen.

13. The instructions, when executed, cause the at least one processor to: driving an interrupt and providing a control parameter in response to predicting the outcome of the execution path to be faulty; clearing the interrupt in response to the predicted outcome of the execution path being no longer faulty.

13. The computer program of claim 12,

14. The instructions, when executed, cause the at least one processor to: In response to determining that the resolution strategy exists for the execution path, applying the resolution strategy to the execution path; In response to determining that the resolution strategy does not exist for the execution path, creating a resolution strategy list containing resolution strategies in ascending order of system performance cost and applying a first resolution strategy from the list.

14. A computer program product according to claim 12 or 13, which causes a computer to:

15. The instructions, when executed, cause the at least one processor to: applying a next resolution strategy in the resolution strategy list in response to determining that the predicted outcome of the execution path has not changed and that all resolution strategies from the resolution strategy list have not been tried; relaying impact data of the resolution strategies in response to determining that the predicted outcome of the execution path has not changed and that all resolution strategies from the resolution strategy list have been tried; relaying impact data of the resolution strategy in response to determining that the predicted outcome of the execution path has changed; 15. A computer program product as claimed in claim 14, which causes the computer to:

16. 16. A computer program product according to any one of claims 12 to 15, wherein the instructions, when executed, cause the at least one processor to determine an appropriate frequency for data collection.

17. 17. A computer program product as claimed in any one of claims 12 to 16, wherein the instructions, when executed, cause the at least one processor to extract and refine a profile to predict the outcome of the execution path.

18. 18. A computer program product according to any one of claims 12 to 17, wherein the instructions, when executed, cause the at least one processor to extract and refine profiles using fractal similarity searching.

19. The method of claim 18, further comprising: predicting an outcome of an execution path from monitored telemetry; determining a resolution strategy for said execution path; applying said resolution strategy; determining whether the predicted outcome of the execution path has changed; storing impact data of the applied resolution strategies; A method comprising:

20. driving an interrupt and providing a control parameter in response to predicting the outcome of the execution path to be faulty; clearing the interrupt in response to the predicted outcome of the execution path being no longer a fault; 20. The method of claim 19 further comprising:

21. responsive to determining that the resolution strategy exists for the execution path, applying the resolution strategy to the execution path; responsive to determining that the resolution strategy does not exist for the execution path, creating a resolution strategy list containing resolution strategies in ascending order of system performance cost and applying a first resolution strategy from the list; 21. The method of claim 19 or claim 20, further comprising:

22. applying a next resolution strategy in the resolution strategy list in response to determining that the predicted outcome of the execution path has not changed and that all resolution strategies from the resolution strategy list have not been tried; relaying impact data of the resolution strategies in response to determining that the predicted outcome of the execution path has not changed and that all resolution strategies from the resolution strategy list have been tried; relaying impact data of the resolution strategy in response to determining that the predicted outcome of the execution path has changed; 22. The method of claim 21 further comprising:

23. 23. The method of any one of claims 19 to 22, further comprising determining an appropriate frequency for data collection.

24. 24. The method of any one of claims 19 to 23, further comprising extracting a profile to predict the outcome of the execution path using a fractal similarity search.

25. 25. The method of any one of claims 19 to 24, further comprising extracting a profile to predict the outcome of the execution path using a fractal similarity search.

26. A computer-readable storage medium storing a computer program according to any one of claims 12 to 18.

Citation Information

Patent Citations

  • Trouble prediction operation control device

    JP1995002200A

  • Computer fault detection system

    JP1997251402A

  • Fault solution prediction system and method

    JP2019153306A

  • Failure-Model-Driven Repair and Backup

    US20100318837A1