Composite System Health Indicator for Alert Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing monitoring systems generate false positive alerts in complex systems, wasting technician time and obscuring critical issues, due to challenges in setting single-measurement thresholds and reliance on subjective user evaluations for machine learning models.

Innovation Solution

A system health indicator is developed, measuring the percentage of time active user sessions spend waiting on events, combined with user impact weighting, to accurately predict and prioritize alerts, reducing false positives and labor costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If single-measurement thresholds are used for alerting, then the system can detect critical problems, but false positive alerts are generated that waste technician time and obscure critical issues

Engineering Contradiction:
Improvealert accuracyVSAvoidtechnician time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent combines multiple measurements into a single composite health indicator that represents overall system health. Instead of monitoring individual measurements separately with multiple thresholds, the invention aggregates them into one unified metric that reduces false positives while maintaining detection of critical problems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The composite health indicator acts as an intermediary between raw measurements and alert generation. This intermediate metric synthesizes multiple measurements and applies intelligent weighting to determine when alerts should be generated, filtering out false positives before they reach technicians.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple measurements are monitored to identify user operations failing or executing inefficiently, then the system can detect root causes, but the complexity of analyzing multiple measurements increases

Engineering Contradiction:
Improveproblem detection accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple measurements into a single composite health indicator that maintains the diagnostic information of individual measurements while presenting a unified view. This reduces the complexity of analyzing multiple separate measurements while preserving the ability to identify root causes.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The composite health indicator can be segmented into contributions from different measurement categories, allowing technicians to drill down into specific problem areas when needed, while maintaining an overall simplified view for routine monitoring.

Inventive Principle:
Principle #1Segmentation

3Reliability

If thresholds are set to predict critical problems before they occur, then sub-critical alerts can be generated, but it is difficult to set thresholds that accurately predict problems without generating false positives

Engineering Contradiction:
Improveproblem prediction accuracyVSAvoidthreshold setting complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the threshold-setting problem by changing from multiple individual thresholds to a single composite health indicator threshold. This parameter transformation simplifies the prediction mechanism while maintaining accuracy through the intelligent aggregation and weighting of multiple measurements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8458530B2Continuous system health indicator for managing computer system alerts
Publication Date: 2013.06.04 ORACLE INT CORP
  • US8458530B2 patent drawing
  • US8458530B2 patent drawing
  • US8458530B2 patent drawing

AI summary

A method is provided for detecting when users are being adversely impacted by poor system performance. A system health indicator is determined that is based on the amount of work that is blocked waiting for each of a set of an external events and combined with a heuristic that is based on the number of users waiting for the work to complete. The system health indicator is compared to a threshold such that an alert is generated when the system health indicator crosses the threshold. However, the system health indicator is designed so that an alert is only generated when a significant user base is or will in the near future experience a problem with the system. Furthermore, the system health indicator is designed to vary smoothly to maintain its suitability for the application of predictive technology.