Network Service Root Cause Mapping for SLA Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network management faces challenges in determining the optimal balance of hardware and software resources to meet service level agreements (SLAs) without incurring excessive costs or insufficient capabilities, as network architects must balance equipment deployment to satisfy performance conditions while managing potential failures and overloads.
Innovation Solution
A method and apparatus for mapping and identifying root causes of performance issues in network-based services by establishing performance objective values, monitoring thresholds, and correlating transaction performance with executing elements to determine the cause of degradation, allowing for proactive resource allocation and remedial actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If more hardware and software equipment is deployed to satisfy performance conditions, then network performance and reliability are improved, but cost for purchase and maintenance increases
Solution Approach 1:
The system performs preliminary mapping between SLOs and SLAs, and establishes performance thresholds before actual failures occur. By pre-configuring the relationship between infrastructure metrics and service-level agreements, the system can proactively identify root causes of performance degradation and take corrective actions before SLA violations happen, avoiding the need for over-provisioning of equipment.
Solution Approach 2:
The system continuously monitors actual performance metrics, compares them against mapped SLOs and SLAs, and provides feedback about root causes of deviations. This closed-loop feedback mechanism enables dynamic adjustment of resource allocation based on actual performance needs rather than static over-provisioning, optimizing the balance between reliability and equipment quantity.
2Quantity of substance
If fewer hardware and software equipment is deployed to reduce costs, then maintenance cost decreases, but network performance may fail to satisfy SLAs
Solution Approach 1:
The system introduces performance thresholds and root cause analysis as intermediary layers between infrastructure monitoring and SLA compliance verification. By mapping SLOs to SLAs and identifying root causes of performance deviations, the system can optimize equipment deployment to the minimum necessary level while ensuring SLA compliance, avoiding both over-provisioning and under-provisioning.
Solution Approach 2:
The system changes the approach from monitoring only final SLA metrics to monitoring underlying SLO parameters and their mapped relationships. By analyzing performance at the infrastructure level through the mapped SLOs and identifying root causes, the system can optimize equipment quantity to satisfy SLAs with minimal resources rather than relying on excessive equipment deployment.
3Reliability
If network equipment is optimized to satisfy performance conditions under normal operation, then service level agreements are met, but capability to handle failures and overloads is insufficient
Solution Approach 1:
The system performs preliminary mapping of SLOs to SLAs and establishes performance thresholds that account for failure scenarios. By pre-configuring the relationship between infrastructure metrics and service-level agreements including failure conditions, the system can proactively identify root causes of performance degradation and take corrective actions before SLA violations happen, ensuring both normal operation compliance and failure handling capability.
Solution Approach 2:
The system builds in performance margins and thresholds during the mapping phase that provide a cushion against failures and overloads. By establishing thresholds that account for variability and failure scenarios, the system creates a buffer that maintains SLA compliance even when equipment fails or operates at elevated levels, enhancing adaptability without requiring excessive equipment deployment.
Data Source
AI summary
A method, apparatus and computer-program product for mapping and identifying root causes of performance problems in network based services, wherein the service is composed of applications and transactions, is disclosed. The method comprises the steps of establishing a performance objective value, and a threshold value therefrom, for selected ones of the transactions for each of the applications, wherein the aggregate of the performance objective values insures a known service performance, monitoring a measure of performance for each of the selected transactions, generating an indication for each of the performance measures that exceeds a corresponding threshold value and determining the cause of the degradation by correlating the transactions generating the indication with the elements executing the transaction.


