Multi-domain Network Health Analysis for Root Cause Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network management solutions fail to accurately pinpoint the origin of performance problems in IT Data Center infrastructure, leading to inefficient resource allocation and increased maintenance costs, as they often alert on performance issues but do not identify the root cause, especially during data bursts or equipment overload.
Innovation Solution
A method and system that analyze network health by examining the performance metrics of nodes and links across multiple domains, using traffic behavior and resource utilization metrics to determine the health of network elements and identify the specific nodes or links causing infrastructure problems, allowing for targeted corrections before performance degradation occurs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If network management systems monitor performance metrics across multiple domains, then the ability to detect performance issues is improved, but the complexity of analyzing and pinpointing the root cause increases
Solution Approach 1:
The system segments the network into multiple domains (data center, wide area, access, core) and analyzes each domain separately using domain-specific metrics and algorithms. This segmentation allows the complex multi-domain network to be broken down into manageable parts, improving root cause identification while controlling analysis complexity through structured domain decomposition.
Solution Approach 2:
The system introduces a hierarchical dimension by implementing multi-level analysis from individual network elements up to entire domains. By adding this dimensional structure with intermediate aggregation levels, the system manages the complexity of analyzing all network elements simultaneously while maintaining comprehensive monitoring coverage across all domains.
2Reliability
If additional supporting hardware or software is introduced to correct bottlenecks, then network performance is improved, but the cost of the network infrastructure increases
Solution Approach 1:
The system performs preliminary identification and analysis of potential bottlenecks before they cause significant performance degradation. By detecting early signs of infrastructure problems through multi-domain metric analysis and calculating domain health scores, the system enables proactive remediation planning that can address issues with existing resources before additional hardware or software investments are required.
Solution Approach 2:
The system changes operational parameters of existing network infrastructure rather than immediately adding new hardware. By adjusting traffic routing, load balancing, and resource allocation based on real-time health assessments and bottleneck identification, the system optimizes performance using existing infrastructure, thereby avoiding unnecessary capital expenditures on additional network equipment.
3Stability of the object's composition
If access to the network is reduced to alleviate bottlenecks, then existing network performance is stabilized, but the delay perceived by new users increases
Solution Approach 1:
The system applies local quality control by identifying specific domains or network segments experiencing bottlenecks and applying targeted remediation actions only to those affected areas. Rather than reducing overall network access, the system locally optimizes problematic segments through selective traffic management and resource allocation adjustments, maintaining stable performance in affected areas while preserving access for new users in healthy domains.
Data Source
AI summary
A method, system, and program, product for analyzing a computer network comprising one or more domains, each of the one or more domains comprising a plurality of nodes and one or more links, the method comprising calculating the health of the computer network, determining, based on the computer network health, if an infrastructure problem exists, identifying, based on the determination, a domain of the one or more domains of the computer network, further identifying, based on the identified domain, an infrastructure problem selected from the group comprising the plurality of nodes and the one or more links of the identified domain, determining an origin of the cause of the infrastructure problem based on the identified infrastructure problem.


