Containerized Application Health Monitoring for Critical Event Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current containerized application management systems (CAMS) lack effective methods to monitor and report operational health of enterprise applications, as existing tools provide unfiltered and overwhelming event data, failing to identify and filter events that impact application performance, and do not notify administrators of critical resource failures or changes that jeopardize application health.
Innovation Solution
A CAMS-application health monitoring (CAMS-AHM) system is introduced to map resource operational states to application states, determine the impact on operational health, and generate notifications or recommended actions based on defined thresholds and filters, enabling targeted alerts for significant or catastrophic issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing monitoring tools are used to track CAMS cluster events, then comprehensive event data is collected, but the data becomes unfiltered and overwhelming, making it difficult to identify critical issues
Solution Approach 1:
The patent extracts and filters only the critical event data from the overwhelming volume of CAMS cluster events. The monitoring system identifies and separates significant events (such as resource failures, performance degradation, and critical alerts) from routine operational data, presenting only the essential information to administrators for actionable insights.
Solution Approach 2:
The patent introduces an intermediary filtering layer between the CAMS cluster events and the administrator. This intermediary system processes, analyzes, and prioritizes events before presentation, acting as a mediator that transforms raw event data into meaningful, actionable alerts while maintaining the comprehensive monitoring capability.
2Reliability
If all CAMS cluster events are monitored and reported, then complete system visibility is achieved, but administrators cannot distinguish critical failures from routine operations
Solution Approach 1:
The patent applies local quality by assigning different levels of importance and filtering characteristics to different types of events. Critical events such as resource failures, performance degradation, and system alerts receive prioritized handling and immediate notification, while routine operational events are filtered or aggregated, ensuring that critical information is not lost in the volume of data.
Solution Approach 2:
The patent implements feedback mechanisms that continuously monitor CAMS cluster events and dynamically adjust filtering and alerting based on system state. The system provides feedback to administrators about critical events that require attention while maintaining comprehensive monitoring capabilities, ensuring that reliability is preserved without information loss.
3Measurement precision
If resource operational state changes are tracked in detail, then application performance impact can be determined, but the complexity of mapping resources to applications increases
Solution Approach 1:
The patent implements a universal mapping mechanism that handles multiple resource types (nodes, pods, containers, services) and application components through a common framework. This universal approach simplifies the complexity of resource-application mapping by providing standardized methods for tracking and correlating operational state changes across the entire CAMS ecosystem, maintaining precise application performance monitoring.
Data Source
AI summary
The described technology is generally directed towards monitoring the operational health of one or more applications deployed on a containerized application management system (CAMS) cluster. Various application resources can be mapped to one or more applications, in conjunction with various events and filters. Events can be monitored and reviewed to determine which application resource(s) is associated with the event, and further, the event's effect on the application(s). The events can be filtered based upon the effect on the operational health of the application. For example, if an event is identified as insignificant, the event does not have to be reported. In another example, if the event is identified as being potentially critical to the operational health of the application, a notification can be generated informing an operator that an application may have experienced potentially catastrophic damage. Based thereon, the administrator can then undertake an action(s) necessary to mitigate the damage.


