Application Health Scoring from Correctable Network Errors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network monitoring systems fail to capture, analyze, and score correctable errors in application traffic, lacking visibility of network-wide application health and failing to provide alerts before catastrophic failures occur.
Innovation Solution
A system and method for monitoring network-wide application health by detecting and correcting correctable errors, using network devices to increment error counters, calculate scores, and telemeter data to a controller for generating graphical interfaces and alerts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If network monitoring focuses only on uncorrectable errors, then device complexity is reduced and ease of operation is improved, but measurement precision of application health is insufficient and reliability is compromised
Solution Approach 1:
The patent segments error monitoring into distinct categories: correctable errors (detected and corrected by ECC) and uncorrectable errors (dropped packets, checksum errors). This segmentation allows the system to track both types separately, providing comprehensive application health visibility without overwhelming complexity. Each error type is counted and reported independently, enabling precise measurement while maintaining manageable system structure.
Solution Approach 2:
The patent introduces an intermediary error counter mechanism that sits between the network device and the monitoring system. This intermediary component captures correctable error information from the network device's error logs and forwards it to the monitoring system, which then generates application health reports. This intermediary layer simplifies the overall system architecture by centralizing the complexity of error analysis while providing comprehensive measurement precision.
2Reliability
If correctable errors are captured and analyzed, then reliability is improved through early warning, but loss of information increases due to additional error data collection
Solution Approach 1:
The patent extracts only the essential error information from network traffic - specifically, the count of correctable errors detected by ECC mechanisms. Rather than capturing and analyzing complete error packets or detailed error sequences, the system extracts simplified error counts and transmits only these aggregated metrics to the monitoring system. This extraction approach maintains high reliability through comprehensive error tracking while minimizing bandwidth consumption by transmitting only critical summarized data.
3Measurement precision
If comprehensive error tracking is implemented, then measurement precision of application health is improved, but productivity of network operations decreases due to additional monitoring overhead
Solution Approach 1:
The patent implements partial monitoring by focusing only on correctable errors that are detected and corrected by ECC mechanisms, rather than attempting to analyze every single network packet or error type. This selective partial action provides sufficient application health measurement precision for reliability assessment without the excessive overhead of comprehensive packet-level analysis. The system captures error counts at strategic points in the network stack, balancing measurement needs with operational productivity.
Data Source
AI summary
Disclosed are systems, methods, and non-transitory computer-readable storage media for monitoring application health via correctable errors. The method includes identifying, by a network device, a network packet associated with an application and detecting an error associated with the network packet. In response to detecting the error, the network device increments a counter associated with the application, determines an application score based at least in part on the counter, and telemeters the application score to a controller. The controller can generate a graphical interface based at least in part on the application score and a timestamp associated with the application score, wherein the graphical interface depicts a trend in correctable errors experienced by the application over a network.


