MTBO Metric for Network Equipment Reliability Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The Mean Time Between Failures (MTBF) metric does not accurately reflect the reliability of communication network equipment, as it only counts hardware failures and ignores automatic recoveries and software failures that impact customers, leading to incomplete assessment of network service reliability.
Innovation Solution
Introducing the Mean Time Between Outages (MTBO) metric, which tracks customer impacting failures caused by both hardware and software, and compares it with a goal metric calculated based on vendor predictions, to provide a comprehensive reliability assessment of network equipment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the field MTBF metric is used to measure component reliability, then hardware failure tracking is simplified, but the metric fails to capture software failures and automatic recoveries that impact customers
Solution Approach 1:
The patent changes the measurement parameter from traditional MTBF (Mean Time Between Failures) to MTBO (Mean Time Between Outages). This parameter change enables the metric to capture both hardware failures and software failures, as well as distinguish between automatic recoveries and manual interventions, thereby improving reliability measurement accuracy without significantly increasing system complexity
Solution Approach 2:
The patent introduces an intermediary classification system that categorizes failures into different types (hardware vs. software) and recovery methods (automatic vs. manual). This intermediary layer allows the MTBO metric to comprehensively track customer-impacting events while maintaining a structured approach that doesn't overly complicate the measurement system
2Reliability
If automatic recovery events are excluded from failure counting, then the field MTBF metric remains simple to calculate, but it overstates the actual reliability experienced by customers
Solution Approach 1:
The patent transitions from counting only hardware failures to counting customer-impacting outages, where an outage is defined as a failure requiring manual intervention. This parameter change ensures that automatic recoveries are properly distinguished from true outages, improving both customer-experienced reliability measurement and measurement precision simultaneously
Solution Approach 2:
The patent segments failure events into distinct categories: hardware failures, software failures, automatic recoveries, and manual interventions. This segmentation allows for precise tracking of only those events that truly impact customers (manual interventions), while still capturing the full spectrum of failure modes for comprehensive analysis
Data Source
AI summary
A method and system for measuring a customer impacting failure rate in a communication network are disclosed. For example, the method collects a plurality of customer impacting network failure events, where the plurality of customer impacting network failure events comprises both hardware failure events and software failure events associated with a particular type of router or switch, or a particular type of component of the router or the switch. The method computes a Mean Time Between Outage (MTBO) metric from the plurality of customer impacting network failure events and compares the MTBO metric with a MTBO goal metric, wherein the MTBO goal metric is calculated in accordance with a predicted Mean Time Between Failure (MTBF) metric.


