Hierarchical Fault Detection in Application Performance Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monitoring and managing application performance in large, distributed computer systems is complex due to dynamic configurations and the difficulty in identifying operational faults across multiple hierarchical levels, making it challenging to provide timely and accurate alerts and automated repairs.
Innovation Solution
An application performance management system with a dynamic discovery agent that monitors hierarchical operational elements, processes metrics from different levels, and issues status reports to users, enabling identification of operational faults and related elements, along with suggested fixes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If monitoring is extended to cover all applications across distributed processors, then measurement precision improves, but device complexity increases
Solution Approach 1:
The monitoring system is segmented into hierarchical levels (system level, application level, component level) with separate monitoring agents at each level. This allows comprehensive monitoring of all applications across distributed processors while managing complexity through modular organization of monitoring functions at different hierarchical strata.
Solution Approach 2:
Monitoring agents serve as intermediaries between the monitored applications and the central management system. These agents collect metrics locally and transmit them upward through the hierarchy, enabling comprehensive monitoring without requiring direct complex interactions between all system components.
2Adaptability or versatility
If dynamic configuration is implemented to move applications between processors, then adaptability improves, but difficulty of detecting and measuring increases
Solution Approach 1:
The system pre-establishes hierarchical monitoring relationships and metric collection templates before applications are moved. When dynamic configuration occurs, the monitoring framework already has the structure in place to track applications across processors, eliminating the need to reconfigure monitoring relationships during runtime.
Solution Approach 2:
The monitoring agents are designed with universal functionality to monitor any application regardless of which processor it resides on. The hierarchical metric structure allows the same monitoring mechanisms to track applications dynamically moved between processors without requiring application-specific monitoring configurations.
3Reliability
If comprehensive monitoring of all hierarchical levels is implemented, then reliability improves, but loss of information increases
Solution Approach 1:
Metrics are segmented by hierarchical level (system-level metrics, application-level metrics, component-level metrics) and only relevant metrics are aggregated and transmitted at each level. This maintains comprehensive monitoring for reliability while filtering out redundant information to prevent information overload in the central system.
Solution Approach 2:
The system extracts only the essential diagnostic information from comprehensive metric collections at each hierarchical level. By selecting and transmitting only the most relevant metrics upward through the hierarchy, the system maintains high reliability through comprehensive monitoring while minimizing information loss through intelligent filtering of non-essential data.
Data Source
AI summary
An application performance management system is disclosed. Operational elements are dynamically discovered and extended when changes occur. Programmatic knowledge is captured. Particular instances of operational elements are recognized after changes have been made using a fingerprint/signature process. Metrics and metadata associated with a monitored operational element are sent in a compressed form to a backend for analysis. Metrics and metadata from multiple similar systems may be used to adjust/create expert rules to be used in the analysis of the state of an operational element. A 3-D user interface with both physical and logical representations may be used to display the results of the performance management system.


