System and method for supervised event monitoring

By adopting supervised machine learning model and detection thresholds in the operation impact event detection of business services, the problem of low detection accuracy of operation data in the prior art is solved, and more accurate detection of business services operation impact event is achieved.

CN120153364APending Publication Date: 2025-06-13MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380077568.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-11-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is prone to errors when using operation data to detect operational impact events of service services. Especially in the case of noisy operation data, it is difficult to effectively distinguish data indicating operational impact events from data not indicating operational impact events.

Method used

The supervised machine learning model is adopted and the model is trained using marked event information to distinguish between operational data indicating the impact of business service operations and operational data that does not indicate the impact of operational operations. By generating detection thresholds, the detection accuracy of operational impact events is improved.

Benefits of technology

Reduces noisy operation data, improves the detection accuracy of events affecting business service operations, and ensures more accurate event detection and monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153364A_ABST
    Figure CN120153364A_ABST
Patent Text Reader

Abstract

A method, computer program product and computing system for processing event data associated with a plurality of known operational impact events on a business service and operational data associated with the business service using a supervised machine learning model conditioned on operational impact parameters associated with the business service. The detection threshold is generated using a supervised machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] During the operation of a business service (i.e., infrastructure, platform, and / or software hosted by a provider and available to users), events that affect the performance of the business service often occur. Thus, operation data from the service can indicate the operational impact of the events on the business service. For example, during a natural disaster or storm, a server loses power, and the service performance changes (e.g., the server outage causes a performance degradation). In this example, the event (e.g., the server losing power) and the operation data showing the performance degradation of the server define an operational impact event of the service (e.g., the server losing power causes the service performance to degrade). However, not all operation data can indicate or at least meaningfully indicate the operational impact on the business service. Therefore, using statistical methods to determine whether an event is an operational impact event from noisy operation data can be inefficient or error-prone. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Figure 1 is a schematic diagram of an example computer system and an event detection process coupled to a distributed computing network;

[0003] Figure 2 is Figure 1 a flowchart of an example implementation of the event detection process of

[0004] Figures 3 to 5 is Figure 1 a schematic diagram of an example event detection process of

[0005] Figure 6 is Figure 1 a flowchart of an example implementation of the event detection process of; and

[0006] Figure 7 is Figure 1 a schematic diagram of an example event detection process of

[0007] In the various figures, like reference numerals represent like elements. DETAILED DESCRIPTION

[0008] Traditional methods for detecting operational impact events (i.e., events that impact the operation of a business service) using operational data (e.g., telemetry data and / or service level indicator (SLI) data) are error-prone. For example, traditional methods for detecting operational impact events are unsupervised (i.e., a machine learning paradigm for dealing with problems where the available data consists of unlabeled examples). Examples of unsupervised learning tasks include clustering, dimensionality reduction, density estimation, and other anomaly-based methods. For example, assume that the normal latency range for a business service is from 800 milliseconds to 1200 milliseconds. Assume that the operational data for the business service reports a latency higher than the normal latency (e.g., a latency of 1600 milliseconds). In this example, traditional unsupervised methods will detect an operational impact event for the business service because the latency of 1600 milliseconds is outside the normal latency range.

[0009] However, due to the increased latency, this traditional method does not consider whether an operational impact event is actually observed. For example, assume that no operational impact (e.g., performance degradation) on the business service will be observed unless the latency exceeds a higher threshold, such as 3000 milliseconds. In this example, according to various implementations of the present disclosure, labeled event information indicating that no operational impact event will be observed unless the latency exceeds, for example, 3000 milliseconds can be used to supervise the training of a machine learning model to distinguish operational data into operational data indicating an operational impact on the business service (i.e., an operational impact event) and operational data not indicating an operational impact on the business service. Thus, the implementations of the present disclosure reduce noisy operational data and improve the detection accuracy of operational data indicating operational impact events.

[0010] Details of example implementations are set forth in the accompanying drawings and the following description. Other features and advantages will become apparent from the specification, drawings, and claims.

[0011] System Overview:

[0012] Reference Figure 1 , the event detection process 10 is shown as residing on and being executed by a storage system 12, which is connected to a network 14 (e.g., the Internet or a local area network). Examples of the storage system 12 include: network-attached storage (NAS) systems, storage area networks (SAN), personal computers with memory systems, server computers with memory systems, and cloud-based devices with memory systems. A SAN includes one or more of personal computers, server computers, a series of server computers, microcomputers, mainframes, RAID devices, and NAS systems.

[0013] Various components of the storage system 12 execute one or more operating systems, examples of which include: OS Red Mobile, Chrome OS, Blackberry OS, Fire OS, or a custom operating system (Microsoft and Windows are registered trademarks of Microsoft Corporation in the United States, other countries, or both; Mac and OS X are registered trademarks of Apple Inc. in the United States, other countries, or both; Red Hat is a registered trademark of Red Hat Corporation in the United States, other countries, or both; Linux is a registered trademark of Linus Torvalds in the United States, other countries, or both).

[0014] The instruction set and subroutines of the event detection process 10 stored on the storage device 16 included in the storage system 12 are executed by one or more processors (not shown) and one or more memory architectures (not shown) included in the storage system 12. The storage device 16 may include: a hard disk drive; an optical disk drive; a RAID device; random access memory (RAM); read only memory (ROM); and all forms of flash storage devices. Additionally or alternatively, some portions of the instruction set and subroutines of the event detection process 10 are stored on a storage device external to the storage system 12 (and / or executed by a processor and memory architecture).

[0015] In some implementations, the network 14 is connected to one or more auxiliary networks (e.g., network 18), examples of which include: a local area network; a wide area network; or an intranet.

[0016] Various input / output (I / O) requests (e.g., I / O request 20) are sent from the client applications 22, 24, 26, 28 to the storage system 12. Examples of the I / O request 20 include a data write request (e.g., a request to write content to the storage system 12) and a data read request (e.g., a request to read content from the storage system 12).

[0017] The instruction sets and subroutines of client applications 22, 24, 26, 28 (which may be stored on storage devices 30, 32, 34, 36 respectively coupled to client electronic devices 38, 40, 42, 44) may be executed by one or more processors (not shown) and one or more memory architectures (not shown) incorporated in client electronic devices 38, 40, 42, 44 respectively. Storage devices 30, 32, 34, 36 may include: hard disk drives; tape drives; optical disk drives; RAID devices; random access memory (RAM); read only memory (ROM); and all forms of flash memory storage devices. Examples of client electronic devices 38, 40, 42, 44 include personal computer 38, laptop computer 40, smart phone 42, laptop computer 44, server (not shown), data-enabled devices, and dedicated network devices (not shown). Each of client electronic devices 38, 40, 42, 44 executes an operating system.

[0018] Users 46, 48, 50, 52 may access storage system 12 directly through network 14 or through secondary network 18. Additionally, storage system 12 may be connected to network 14 through secondary network 18, as shown by link 54.

[0019] Various client electronic devices may be directly or indirectly coupled to network 14 (or network 18). For example, personal computer 38 is shown as being directly coupled to network 14 via a hardwired network connection. Additionally, laptop computer 44 is shown as being directly coupled to network 18 via a hardwired network connection. Laptop computer 40 is shown as being wirelessly coupled to network 14 via a wireless communication channel 56 established between laptop computer 40 and a wireless access point (e.g., WAP) 58, which wireless communication channel is shown as being directly coupled to network 14. WAP 58 may be, for example, an IEEE 802.11a, 802.11b, 802.11g, 802.11n, and / or device. Smart phone 42 is shown as being wirelessly coupled to network 14 via a wireless communication channel 60 established between smart phone 42 and a cellular network / bridge 62, which wireless communication channel is shown as being directly coupled to network 14.

[0020] Event detection process:

[0021] At least with reference to Figures 2 to 7, the event detection process 10 processes 200 event data associated with multiple known operation impact events on the business service and operation data associated with the business service using a supervised machine learning model conditioned on operation impact parameters associated with the business service. A detection threshold is generated 202 using the supervised machine learning model.

[0022] As will be discussed in more detail below, implementations of the present disclosure allow for supervised monitoring and detection of operation impact events from known operation impact events and operation data associated with the events. During the operation of a business service (e.g., a cloud service), various events occur (i.e., natural disasters, routine maintenance, unplanned system upgrades, technical failures, unexpected user bandwidth, etc.), which may or may not have an operation impact on the performance of the business service. In some implementations, events can be detected or indicated by event data (e.g., alerts, communications, or other indications associated with the business service). In one example, assume that a storm interrupts the power of a server associated with a business service, and event data indicating that the storm interrupted the power of the server is received. In this example, due to the power interruption, the operation data shows a degradation in the performance of the business service. Thus, this operation data (e.g., degradation in the performance of the business service) is an operation impact event because the operation of the business service is affected (e.g., degradation in the performance of the business service). In another example, normal business service functionality is observed because even though the operation data indicates an increase in latency, no event alert is reported. Thus, known operation impact events and corresponding operation data allow for the detection of operation impact events. Specifically, by using known operation impact events and operation data for supervised training, noisy operation data can be filtered, and the detection of operation impact events from operation data can be enhanced. In this way, during the operation of a business service, operation impact events are automatically monitored and accurately detected using various types of operation data.

[0023] In some implementations, a business service generally includes infrastructure, platforms, and / or software hosted by a provider and available for users for electronic data processing, data management, and / or data storage. In each of these examples, access to the business service can be implemented using an electronic device. In this way, a commercial service can include a cloud computing service that utilizes electronic computing devices (e.g., mobile phones, desktop computers, laptop computers, tablet computers).

[0024] In one example, the business service is a data storage service. As Figure 1As shown, a data storage service (e.g., business service 64) allows data from various computing devices (e.g., client electronic devices 38, 40, 42, 44) to independently access data stored in computer system 12. In this way, computer system 12 provides "cloud-based" storage access to client electronic devices 38, 40, 42, 44. During the operation of business service 64, various events may occur that affect the operation of the business service. For example, assume that in addition to computer system 12, business service 64 includes various computer systems across different locations. In one example, access to one of these computer systems is temporarily disabled (e.g., during routine maintenance). In another example, a particular computer system may experience a natural disaster, thereby interrupting communication between the computer system and the client electronic devices.

[0025] At any point in time before, during, or after such an operation impact, event data (e.g., alerts, communications, or other indications associated with the business service) related to the event affecting the operation of the business service can be observed. For example, an electronic communication or email can be sent to the users of business service 64, e.g., to alert them of upcoming or planned operation impacts in business service 64. In another example, an outage mapping system maintained by business service 64 or others can provide a list of operation impacts. In another example, an electronic feed or social media platform can post messages from users and / or system administrators regarding operation impacts in the business service. In each of these cases, event data (e.g., electronic communication, outage mapping system, list of operation impacts) is generated that indicates an event associated with the business service. In one example, the event data includes individual "tickets" of a ticket management system associated with the business service. In another example, the event data includes alerts or other messages associated with impacts (e.g., outages) observed by a monitoring system associated with the business service. In another example, the event data includes communications regarding the status of the business service (e.g., email or social media posts). Although several examples of event data types have been provided, it should be understood that any type of event data can be used within the scope of the present disclosure.

[0026] In some implementations, the event detection process 10 optionally receives multiple event data respectively associated with multiple known operation impact events on a business service. In one example, it is assumed that a specific operation impact event (e.g., interruption of a corresponding business service) is observed through event data (e.g., a ticket from a ticket management system). In this example, the ticket from the ticket management system is event data associated with a known operation impact event (e.g., business service interruption). In another example, it is assumed that event data (e.g., a social media post regarding the unavailability of a business service) occurs when a known operation impact event (e.g., storage system failure) occurs. In this example, the social media post is event data associated with a known operation impact event (e.g., storage system failure).

[0027] In some implementations, receiving multiple event data includes requesting or polling for specific event data (e.g., once or periodically) from various sources (e.g., a ticket management system, an internal or external operation status management tracking system, a specific social media or social network feed or user, and / or service level indicator information generated by an internal or external system). In some implementations, the multiple events include time information, location information, and information related to operation impact events (e.g., business service loss, business service reduction). As described above, the multiple event data is associated with known operation impacts on a business service.

[0028] In some implementations, the multiple event data associated with multiple known operation impact events on a business service includes multiple event data associated with multiple known operation impact events on a cloud computing service. For example, a cloud computing service provides on-demand availability of computer system resources such as data storage and computing power. Large cloud computing services typically have functions distributed across multiple locations, each of which is a data center. Examples of cloud computing services include those from Microsoft Corporation in the United States Amazon Web Services from Amazon.com, Inc. TM (AWS TM ) and Google Cloud Platform from Google LLC TM . In one example, the event detection process 10 receives multiple event data associated with multiple known operation impact events on a cloud computing service. However, it should be understood that within the scope of the present disclosure, various other event data associated with known operation impact events on any type of business service can be received.

[0029] In some implementations, the event detection process 10 processes 200 event data associated with multiple known operation impact events on the cloud computing service and operation data associated with the cloud computing service using a supervised machine learning model conditioned on operation impact parameters associated with the cloud computing service. As will be discussed in more detail below, the operation impact parameters are parameters used to determine the least sensitive or lowest detection threshold for achieving a specific performance metric. In some implementations, the operation data includes telemetry data and performance data indicating the performance of the business service (e.g., latency, throughput, reliability) associated with the business service. For example, assume that the business service includes an application installed on a client electronic device that interacts with a server. In this example, operation data indicating the performance of the business service (e.g., latency, throughput) on the client electronic device is generated and transmitted to a predefined resource for processing. For example, depending on the application, the operation data is aggregated or processed for individual client devices.

[0030] In some implementations, the operation data includes service level indicator (SLI) data. SLI data is an indicator of the service level of a business service. For example, many business services include a service level agreement (SLA) that defines measurable metrics that describe the performance of the business service for a specific customer. SLI data is an indicator of the metrics defined by the SLA. Examples of SLI data include: latency, throughput, availability, and error rate; persistence (in a storage system), end-to-end latency (for complex data processing systems, especially pipelines), and correctness. It should be understood that within the scope of the present disclosure, various types of SLI data can be used as operation data associated with a specific business service.

[0031] In some implementations, the event detection process 10 utilizes a supervised machine learning model to process multiple event data and operation data associated with a business service to generate a detection threshold. A machine learning model typically includes an algorithm or combination of algorithms that are trained to identify specific types of patterns. For example, depending on the nature of the training data, machine learning methods are generally divided into three categories: supervised learning, unsupervised learning, and reinforcement learning. Supervised learning involves presenting a computing device with example inputs and their desired outputs given by a "teacher", with the goal of learning the general rule that maps the input to the output. In unsupervised learning, the learning algorithm has no labels and can only find structure in the input on its own. Unsupervised learning can itself be a goal (discovering hidden patterns in the data) or a means to an end (feature learning). Reinforcement learning typically involves a computing device interacting in a dynamic environment where the device must perform a certain goal (such as driving a vehicle or playing a game against an opponent). As the program navigates the problem space, it receives feedback similar to a reward and attempts to maximize it. As described above, a machine learning model is supervised if labeled training data is available to "teach" the model. If labeled training data is not available, the machine learning model is unsupervised. If the machine learning model receives feedback to influence the model, the machine learning model is a reinforcement machine learning model. Although three examples of machine learning methods have been provided, it should be understood that other machine learning methods are possible within the scope of the present disclosure.

[0032] In some implementations, the supervised machine learning model is conditioned on an operation impact parameter (i.e., trained with an operation impact parameter). An operation impact parameter is a parameter used to determine the least sensitive or lowest detection threshold for achieving a specific performance metric. For example, assume that the business service 64 is a cloud computing service. In this example, the operation impact parameter is the minimum performance metric of the cloud computing service. In one example, the operation impact parameter is the minimum percentage or number of client devices that experience an operation impact when using the cloud computing service. The operation impact parameter can be a default value or user-defined. Although an example of a single operation impact parameter has been described above, it should be understood that within the scope of the present disclosure, any number of operation impact parameters related to various metrics of operation impact can be used to adjust the supervised machine learning model.

[0033] In some implementations, the event detection process 10 processes 200 multiple event data and operation data associated with a business service using a supervised machine learning model conditioned on operation impact parameters associated with the business service to identify characteristics of the multiple events. For example, in addition to a detection threshold, the event detection process 10 also utilizes a supervised machine learning model to identify characteristics of event data associated with known operation impact events. In one example, assume an event involves the unavailability of a particular business service (e.g., a cloud-based storage service is unavailable to many client devices, as reported via a feedback ticket system). In this example, the event detection process 10 identifies (e.g., via a supervised machine learning model) characteristics or factors that contribute to a known operation impact event. In this example, assume the SLI data indicates a significant increase in the latency experienced by the cloud-based storage service, resulting in unavailability. Thus, the event detection process 10 identifies this characteristic (e.g., a significant increase in latency) associated with the known operation impact event. In this way, the event detection process 10 can associate specific characteristics in the operation data with specific operation impact events of the business service. In some implementations, the event detection process 10 utilizes this mapping of specific characteristics in the operation data to enhance the monitoring and detection of operation impact events.

[0034] In some implementations, the event detection process 10 generates 202 a detection threshold using a supervised machine learning model. For example, the event detection process 10 generates 202 a detection threshold by processing multiple known operation events and operation data associated with a business service using a supervised machine learning model conditioned on operation impact parameters associated with the business service. The detection threshold is a criterion for operation data that uses multiple known operation impact events to separate operation data indicating an operation impact event from operation data not indicating an operation impact event. In some implementations, the event detection process 10 generates 202 a detection threshold by training a supervised machine learning model with known operation impact events and operation data associated with the known operation impact events. For example, using labeled training data (i.e., known operation impact events and corresponding operation data), the event detection process 10 generates 202 a detection threshold using a supervised machine learning model that uses multiple known operation impact events to separate operation data indicating an operation impact event from operation data not indicating an operation impact event.

[0035] For example, also refer to Figure 3, assume that the event detection process 10 processes multiple event data (e.g., multiple event data 300, 302, 304) associated with a known operation impact event to indicate, for example, a failure of a storage system associated with a business service. Further assume that operation data (e.g., operation data 306, 308, 310) indicates that during the operation impact event associated with event data 300, 302, 304, the client latency exceeds, for example, 3000 milliseconds. In this example, the operation impact parameter (e.g., operation impact parameter 312) is defined as, for example, 10%, which means that a supervised machine learning model (e.g., supervised machine learning model 314) is conditioned to generate a detection threshold (e.g., detection threshold 316), which is the least sensitive or minimum criterion for detecting at least, for example, a 10% deviation in client devices. In other words, the detection threshold 316 represents the minimum criterion (e.g., latency in this example) of operation data for detecting operation impacts in at least 10% of the client devices of the business service 64.

[0036] In some implementations, generating 202 a detection threshold using a supervised machine learning model includes determining 204 the number of false positive detections and generating 206 multiple detection thresholds when the number of false positive detections exceeds a false positive detection threshold. For example, there may be cases where a single global detection threshold is less effective than multiple discrete detection thresholds. Assume that a single detection threshold is too low such that a large amount of noisy operation data results in false alarm detections to cover all multiple business events associated with a known operation impact event. In this example, the event detection process 10 determines whether using multiple detection thresholds is more efficient than a single detection threshold by determining 204 the number of false positive detections compared to the number of detection thresholds. If the number of false alarm detections exceeds a predefined threshold, the event detection process 10 generates 206 multiple detection thresholds.

[0037] In some implementations, a baseline calculation algorithm generates a prediction and a "boundary" or detection threshold around the prediction to apply a specific level of sensitivity. For example, "3-Sigma" is based on choosing the mean as the prediction and building around it (e.g., 3 times the standard deviation of a metric). In this way, everything within the 3-sigma boundary is considered "normal" and everything outside this boundary is "deviant". However, for a more sensitive model, a "2-sigma" method (e.g., 2 times the standard deviation) can be used, and for a very insensitive model, a "4.5-sigma" method (i.e., 4.5 times the standard deviation) can be used. Although various examples of granularity have been discussed, it should be understood that any sigma value can be used within the scope of the present disclosure.

[0038] In some implementations, the event detection process 10 generates 206 different detection thresholds for each "split" or grouping of the operational data. For example, the event detection process 10 generates different detection thresholds based on different metrics associated with the operational data. In one example, the event detection process 10 generates 206 multiple detection thresholds for different groups of operational data based on, for example, the location or geographical region of the operational data. In another example, the event detection process 10 generates different detection thresholds based on the respective layers or quality of service layers of the business service. Although several examples have been provided for various groupings of operational data and detection thresholds, it should be understood that these examples are for illustrative purposes only, and any number or type of detection thresholds may be used within the scope of the present disclosure.

[0039] In some implementations, the event detection process 10 uses the detection thresholds to determine 208 the coverage of multiple known operation impact events, and identifies 210 gaps in the coverage of multiple known operation impact events. For example, the event detection process 10 evaluates the blind spots in the coverage of the metrics of the collected operational data. If a dynamic supervised machine learning model or algorithm for differentiating marked business events is not found, the event detection process 10 indicates that the coverage of the collected operational data is incomplete. In one example, the dynamic supervised machine learning model is a supervised learning tool that is used to identify hyperspheres in an N-dimensional space to differentiate known operation impact events and / or events without operation impact. In another example, the dynamic supervised machine learning model is a supervised learning tool that identifies hyperplanes in an N-dimensional space to classify known operation impact events. In another example, the dynamic supervised machine learning model is a supervised learning tool that identifies hyperrectangles in an N-dimensional space to distinguish known operation impact events into multiple classes of known operation impact events. In another example, the dynamic supervised machine learning model is a supervised learning tool that identifies hypercubes in an N-dimensional space to differentiate known operation impact events.

[0040] In some implementations, the quality of most of the operational data (e.g., SLI signals) is poor and the correlation with events (e.g., outages) is poor. In such cases, attempting to cover all known operation impact events will result in a high level of noise. In this way, the event detection process 10 identifies 210 gaps in the coverage of multiple known operation impact events. In some implementations, the event detection process 10 uses clustering analysis to understand which operational data is covering which operation impact events.

[0041] In some implementations, in addition to identifying gaps in the coverage of more than 210 known operation impact events, the event detection process 10 also verifies the coverage of more than 210 known operation impact events and the monitoring quality. For example, as described above, if a supervised machine learning model (e.g., a hypersphere) for differentiating marked business events cannot be found or determined, the event detection process 10 indicates that the coverage of the collected operation data is incomplete. In some implementations, the event detection process 10 generates a report of undetected operation data and / or undetected known operation impacts. In one example, the event detection process 10 determines that the monitoring configuration is ineffective for actual operation impact events based on the number of undetected known operation impact events and / or undetected operation data. In this way, the monitoring effectiveness can be determined and updated in real time by identifying gaps in the coverage in terms of uncovered operation data and uncovered known operation impact events.

[0042] Also refer to Figure 4 and assume that the event detection process 10 receives multiple event data associated with known operation impact events and processes more than 200 event data and operation data to generate a detection threshold 400. In this example, assume that the supervised machine learning model 314 determines that the detection threshold 400 will cover events 402, 404, 406, 408, 410. However, further assume that the supervised machine learning model 314 determines that another detection threshold (e.g., detection threshold 412) will cover events 408, 410 without detecting events 402, 404, 406. In this example, the event detection process 10 can determine that using the detection threshold 400 to detect events 408, 410 will introduce too much "noise" because events 402, 404, 406 are different from events 408, 410. In this way, the event detection process 10 can generate multiple detection thresholds (e.g., detection thresholds 400, 412) to account for various clusters or concentrations of operation impact events.

[0043] Also refer to Figure 6 and the event detection process 10 uses a detection threshold to process operation data 600 associated with a business service, where the detection threshold is generated using a supervised machine learning model conditioned on operation impact parameters associated with the business service. In this way, the event detection process 10 can use the trained supervised machine learning model (e.g., the supervised machine learning model 314) as described above to detect operation impact events 212 associated with the business service. Thus, Figure 2 and Figure 6 illustrate training a supervised machine learning model to generate a detection threshold and using the generated detection threshold to detect operation impact events.

[0044] In some implementations, the event detection process 10 detects an operation impact event associated with a business service 212 by determining that an operation impact event exceeds a detection threshold. For example, the event detection process 10 uses the detection threshold to determine whether the operation data indicates an operation impact event. In one example, assume that the normal latency range of a business service includes 800 milliseconds to 1200 milliseconds. Assume that the operation data of the business service indicates that one or more users are experiencing a latency of 1600 milliseconds. In this example, traditional unsupervised methods would detect a problem with the business service because the 1600 - millisecond latency exceeds the acceptable or normal latency range.

[0045] Now, assume that the event detection process 10 processes multiple event data and operation data associated with multiple known operation impact events to generate a detection threshold for a latency of 3000 milliseconds before detecting an operation impact event. In this example, the event detection process 10 monitors the operation data in real - time to determine whether the latency within the operation data exceeds the detection threshold (e.g., 3000 milliseconds). In response to monitoring operation data that exceeds the detection threshold, the event detection process 10 detects 202 an operation impact event. Detecting an operation impact event can include generating an alert, providing a notification to a business service administrator, updating a log of operation impact events, and / or performing various other remedial measures.

[0046] In some implementations, detecting 212 an operation impact event includes processing 214 the operation data associated with a business service with a detection threshold to identify the operation data that indicates an operation impact event. As described above, the operation data that indicates an operation impact event includes operation data that meets the criteria of the detection threshold of the operation impact parameter.

[0047] Also refer to Figure 5 , assume that the event detection process 10 receives operation data (e.g., operation data 500, 502, 504, 506, 508, 510) associated with various customers of a business service 64. In this example, the event detection process 10 uses the detection threshold 316 to filter out the operation data that indicates an operation impact event from the operation data that does not indicate an operation impact event. In this example, assume that the operation data 500, 508 include latency values (e.g., 3200 milliseconds and 3300 milliseconds, respectively) for latency metrics that exceed the detection threshold 316. In this example, the event detection process 10 detects operation impact events for the operation data 500 and the operation data 508 because these two portions of the operation data both indicate that an operation impact event is occurring. As Figure 5 shown, the dashed line extending from the detection threshold 316 represents the event detection process that separates the operation data 500, 508 from the other operation data. In this example, the event detection process 10 generates an alert or takes some other remedial measure in response to each of the operation data 500, 508.

[0048] In some implementations, the detecting 212 operation impact event includes processing 216 the operation data associated with the service with a detection threshold to identify operation data that does not indicate an operation impact event. As described above, the operation data that does not indicate an operation impact event includes operation data that does not meet the criteria of the detection threshold of the operation impact parameter.

[0049] Referring again to Figure 5 , assume that the event detection process 10 receives operation data (e.g., operation data 500, 502, 504, 506, 508, 510) associated with various customers of the service 64, and filters out the operation data indicating an operation impact event from the operation data that does not indicate an operation impact event by using the detection threshold 316. In this example, assume that the operation data 502, 504, 506, 510 include latency values of latency metrics that are below the detection threshold 316. In this example, the event detection process 10 does not detect an operation impact event for any of the operation data 502, 504, 506, 510 because no part of the operation data indicates that an operation impact event is occurring. As Figure 5 shown, the dashed line extending from the detection threshold 316 represents the event detection process that separates the operation data 502, 504, 506, 510 from the operation data 500, 508. In this example, the event detection process 10 can continue to process the operation data to monitor the operation data that exceeds the detection threshold 316.

[0050] In some implementations, the event detection process 10 updates 218 a supervised machine learning model with the operation data indicating an operation impact event to generate an updated detection threshold. For example, the event detection process 10 can update 218 the supervised machine learning model to consider subsequent (i.e., at any point in time after the previous training of the supervised machine learning model) operation data indicating an operation impact event. Also refer to Figure 7 , in some implementations, assume that the event detection process 10 processes subsequent operation impact events (e.g., operation impact events 700, 702, 704) and subsequent operation data (e.g., operation data 706, 708, 710). In this example, the event detection process 10 updates 218 the detection threshold (e.g., detection threshold 712) by processing the operation impact events 700, 702, 704 and the operation data 706, 708, 710 in the above manner. For example, the event detection process 10 conditions the supervised machine learning model 314 on the operation impact parameter (e.g., operation impact parameter 312). In some implementations, the operation impact parameter can be adjusted over time to limit or relax the detection threshold of various parameters.

[0051] In some implementations, updating 218 a supervised machine learning model with operational data indicating an operation's impact on an event includes periodically updating 220 the supervised machine learning model. For example, event detection process 10 may update 220 the supervised machine learning model 314 at various intervals. In one example, event detection process 10 updates 220 the supervised machine learning model 314 after a threshold amount of time and / or periodically. In this example, event detection process 10 updates 220 the supervised machine learning model 314 daily, for example. In another example, event detection process 10 updates 220 the supervised machine learning model 314 hourly, for example. While various time periods for updating the supervised machine learning model have been described, it should be understood that these are for example purposes only, and event detection process 10 may update 220 the supervised machine learning model at any interval within the scope of the present disclosure.

[0052] Overview:

[0053] As will be understood by those skilled in the art, the present disclosure may be embodied as a method, system, or computer program product. Accordingly, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, generally referred to herein as a "circuit," "module," or "system." Furthermore, the present disclosure may also take the form of a computer program product on a computer-usable storage medium having computer-usable program code embodied therein.

[0054] Any suitable computer-usable or computer-readable medium can be utilized. The computer-usable or computer-readable medium can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. More specific examples of the computer-readable medium include the following: an electrical connection having one or more wires, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable ROM (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a transmission medium such as those supporting the Internet or an intranet, or a magnetic storage device. The computer-usable or computer-readable medium can also be a paper or other suitable medium on which a program is printed, as the program can be electronically captured via, for example, optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner and then stored in a computer memory. In the context of this document, the computer-usable or computer-readable medium can be any medium that can contain, store, transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The computer-usable medium can include a propagated data signal having the computer-usable program code embodied therein (either in baseband or as part of a carrier wave). The computer-usable program code can be transmitted using any appropriate medium, including the Internet, wireline, optical fiber cable, RF, etc.

[0055] The computer program code for carrying out operations of the present disclosure can be written in an object-oriented programming language such as Java, Smalltalk, C++, etc. However, the computer program code for carrying out operations of the present disclosure can also be written in a conventional procedural programming language such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer can be connected to the user's computer through a local area network, a wide area network, or the Internet (e.g., network 14).

[0056] The present disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the instructions, which execute via the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block.

[0057] These computer program instructions can also be stored in a computer-readable memory, which can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer programmable memory produce an article of manufacture including instruction means that implement the functions / acts specified in the flowchart and / or block diagram block(s).

[0058] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions / acts specified in the flowchart and / or block diagram block(s).

[0059] The flowcharts and block diagrams in the figures may illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, in fact, two blocks shown in succession may be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, not at all, or in any other flowchart combination depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a system based on dedicated hardware performing the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0060] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. Unless the context clearly dictates otherwise, the singular forms “a,” “an,” and “the” as used herein also include the plural forms. It should be further understood that when used in this specification, the terms “comprise” and / or “comprising” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0061] All structural, material, acts, and equivalents of the means or step plus function elements in the following claims are intended to include any structure, material, or act that performs the function in combination with other elements specifically claimed. The description of the present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. The embodiment was chosen and described in order to best explain the principles of the present disclosure and the practical application, and to enable others of ordinary skill in the art to understand the present disclosure of various embodiments with various modifications that are suited to the particular use contemplated.

[0062] A variety of implementations have been described. Thus, the disclosure of the present application has been described in detail by reference to embodiments of the present application, and it is apparent that modifications and variations can be made without departing from the scope of the disclosure as defined in the appended claims.

Claims

1. A computer-implemented method executed on a computing device, comprising: processing event data associated with a plurality of known operation impact events on the cloud computing service and operation data associated with the cloud computing service using a supervised machine learning model conditioned on operation impact parameters associated with the cloud computing service; and generating a detection threshold using the supervised machine learning model.

2. The computer-implemented method according to claim 1, wherein generating the detection threshold using the supervised machine learning model comprises: determining the number of false positive detections; and generating a plurality of detection thresholds when the number of false positive detections exceeds a false positive detection threshold.

3. The computer-implemented method according to claim 1, further comprising: determining a coverage of the plurality of known operation impact events using the detection threshold; and identifying gaps in the coverage of the plurality of known operation impact events.

4. The computer-implemented method according to claim 1, further comprising: detecting an operation impact event associated with the business service by determining that the operation data exceeds the detection threshold.

5. The computer-implemented method according to claim 4, wherein detecting the operation impact event comprises processing the operation data associated with the business service using the detection threshold to identify operation data indicative of the operation impact event.

6. The computer-implemented method according to claim 5, wherein detecting the operation impact event comprises processing the operation data associated with the business service using the detection threshold to identify operation data not indicative of the operation impact event.

7. The computer-implemented method according to claim 5, further comprising: updating the supervised machine learning model using the operation data indicative of the operation impact event to generate an updated detection threshold.

8. A computing system, comprising: a processing system including a processor; and a memory storing instructions that, when executed by the processing system, cause the system to perform operations, the operations including: processing operation data associated with a business service using a detection threshold generated using a supervised machine learning model conditioned on operation impact parameters associated with the business service, and detecting an operation impact event associated with the business service by determining that the operation data exceeds the detection threshold.

9. The computing system according to claim 8, wherein the operations further comprise: generating the detection threshold using the supervised machine learning model.

10. The computing system according to claim 9, wherein generating the detection threshold by processing event data associated with a plurality of known operation impact events on the business service and operation data associated with the business service using the supervised machine learning model comprises: determining the number of false positive detections; and generating a plurality of detection thresholds when the number of false positive detections exceeds a false positive detection threshold.

11. The computing system according to claim 9, wherein the operations further comprise: Determine the coverage of the multiple known operation impact events using the detection threshold; And Identify gaps in the coverage of the multiple known operation impact events.

12. The computing system according to claim 8, wherein detecting the operation impact event includes processing subsequent operation data associated with the cloud computing service using the detection threshold to identify operation data indicating the operation impact event.

13. The computing system according to claim 12, wherein detecting the operation impact event includes processing the operation data associated with the cloud computing service using the detection threshold to identify operation data that does not indicate the operation impact event.

14. The computing system according to claim 12, wherein the operation further Includes: Updating the supervised machine learning model using the operation data indicating the operation impact event to generate an updated detection threshold.

15. A computer program product residing on a computer-readable medium, the computer-readable medium having a plurality of instructions stored thereon, the plurality of instructions causing the processor to perform operations when executed by the processor, the operations Includes: Generating a detection threshold by separately processing multiple event data associated with multiple known operation impact events of the cloud computing service and service level indicator (SLI) data associated with the cloud computing service using a supervised machine learning model conditioned on operation impact parameters associated with the cloud computing service; And Detecting an operation impact event associated with the cloud computing service by determining that the operation data exceeds the detection threshold.