Customized Anomaly Detection for Virtual Machines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing monitoring systems often generate false positives, leading to unnecessary alarms and resource reallocation, as they rely on generic technical metrics that do not account for the specific operational characteristics of different systems, potentially causing inconvenience in non-critical systems but posing risks in critical environments like hospitals.

Innovation Solution

A machine learning engine is used to create customized anomaly detectors for specific system resources, selecting operational and technical parameters to define normal operation, thereby reducing false positives by analyzing historical data and determining weights for parameter influence, allowing for a blended function to detect anomalies effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If generic technical metrics are used for monitoring, then the monitoring system can be applied to various systems universally, but false positive alarms increase and system-specific operational characteristics are not accounted for

Engineering Contradiction:
Improvemonitoring system applicabilityVSAvoidanomaly detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by creating customized anomaly detectors for different system resources with system-specific operational parameters. Each detector is tailored to the specific operational characteristics of its target system (e.g., hospital power system vs. gaming system), using locally relevant parameters and weights rather than generic uniform monitoring across all systems.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements parameter changes by dynamically selecting and weighting operational parameters based on system-specific characteristics. The machine learning engine adjusts parameter importance weights and selects relevant operational parameters according to the specific system being monitored, transforming the fixed generic parameter set into an adaptive system-specific parameter configuration.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If generic threshold-based monitoring is used, then the system is simple to implement, but many false positive alarms are generated requiring downstream resource allocation

Engineering Contradiction:
Improvemonitoring system complexityVSAvoidalarm accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces an intermediary machine learning engine that sits between the raw metric collection and the alarm generation process. This intermediary layer processes operational parameters through learned models to generate anomaly scores, acting as a mediator that translates simple metric changes into reliable anomaly detections without requiring complex manual threshold configuration.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements self-service by automatically training and updating anomaly detectors using historical operational data without requiring manual intervention for threshold tuning. The machine learning engine autonomously learns system-specific patterns and adjusts detection parameters, eliminating the need for continuous manual calibration while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If system-specific customized parameters are used, then false positives are reduced, but the monitoring system becomes more complex to implement

Engineering Contradiction:
Improveanomaly detection accuracyVSAvoidcustomized detector implementation
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system implements self-service by automatically training and updating anomaly detectors using historical operational data without requiring manual intervention for threshold tuning. The machine learning engine autonomously learns system-specific patterns and adjusts detection parameters, eliminating the need for continuous manual calibration while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies universality by using a unified machine learning framework that handles diverse system types through a common architecture. The same core anomaly detection engine serves multiple different system resources (power systems, gaming systems, etc.) by adapting to their specific parameters, providing a universal solution that manages complexity through standardization rather than custom implementations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If false positive alarms are generated, then potential issues may be caught early, but unnecessary resource reallocation and human interaction are required

Engineering Contradiction:
Improvesystem monitoring coverageVSAvoidcomputing resource waste
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent applies partial action by generating alarms only when anomaly scores exceed dynamically determined thresholds based on system-specific patterns. Rather than alerting on any deviation, the system selectively triggers alarms only for genuine anomalies, performing partial monitoring action that focuses resources on actual problems rather than exhaustive alerting on all variations.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10102056B1Anomaly detection using machine learning
Publication Date: 2018.10.16 AMAZON TECH INC
  • US10102056B1 patent drawing
  • US10102056B1 patent drawing
  • US10102056B1 patent drawing

AI summary

A machine learning engine is configured to create a customized anomaly detector for use by a system resource, such as a virtual machine instance running specific operations. Anomalies are determined by comparing aspects of current data to “normal” baseline data indicating a normal range of performance and operation of a system resource. Operation by a system resource that is outside of this normal range of performance and operation may be considered an anomaly. The machine learning engine may be used to create customized monitoring that detects anomalies in operation and/or performance of a system resource based on at least some custom parameters that are unique to the particular system that is to be monitored. The parameters may include operational parameters, which may be selected specifically for the system to be monitored. The parameters may also include some technical parameters which are used by many other systems to monitor hardware performance.