Network Capacity Planning Through Probabilistic Failure Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional network planning tools struggle with inefficient capacity allocation due to manual hardcoding of failure events, leading to insufficient estimation of expected flow availability Service Level Objectives (SLO) and increased computational complexity, resulting in suboptimal network designs that fail to meet demands and SLO types.

Innovation Solution

An automated network planning and optimization tool that generates a dynamic set of probabilistic failures based on failure rates and probabilities, using a risk framework to identify key network failures, enabling optimized network designs that meet SLOs at a lower cost by focusing on failures with significant impacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual hardcoding of failure events is used in conventional network planning tools, then network capacity plans can be generated, but the process becomes inefficient and misses critical failure events as network size grows

Engineering Contradiction:
Improvenetwork capacity plan reliabilityVSAvoidnetwork planning efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system automatically generates failure events using a probabilistic framework without requiring manual hardcoding by network planners. The failure policy automatically samples failures from the failure space based on probabilistic models, enabling the system to self-generate comprehensive failure scenarios including critical ones that manual processes would miss.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transitions from static manual failure event specification to dynamic probabilistic failure generation. By changing the parameter representation from fixed hardcoded values to probabilistic distributions, the system can efficiently generate appropriate failure scenarios for networks of any size without manual intervention.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If comprehensive failure events are considered to improve SLO estimation accuracy, then expected flow availability SLO becomes more accurate, but computational complexity grows enormously large

Engineering Contradiction:
Improveexpected flow availability SLO accuracyVSAvoidsolver constraint complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts and focuses only on the most impactful failure events by using probabilistic sampling to identify key failures that contribute most to SLO violations. This extraction approach allows accurate SLO estimation without considering all possible failures, thereby reducing computational complexity while maintaining precision.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses probabilistic sampling to generate a representative subset of failure events that is sufficient for accurate SLO estimation without being exhaustive. This partial action approach provides the necessary accuracy for capacity planning without the overwhelming computational burden of complete failure space enumeration.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If extra capacity is allocated to protect against unexpected network failure events, then business continuity is maintained, but network costs grow inefficiently

Engineering Contradiction:
Improvebusiness continuityVSAvoidnetwork capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system uses expected flow availability SLO as feedback to optimize capacity allocation. By continuously evaluating SLO metrics against failure scenarios, the system identifies the minimum necessary capacity required to meet reliability targets, eliminating excessive capacity allocation while maintaining business continuity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system optimizes capacity parameters based on probabilistic failure analysis and SLO requirements. By changing from static over-provisioning to dynamic capacity optimization driven by failure probability and impact analysis, the system achieves the required reliability with efficient capacity utilization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4167540B1Availability SLO-aware network optimization
Publication Date: 2025.10.22 GOOGLE LLC
  • EP4167540B1 patent drawingFigure 1
  • EP4167540B1 patent drawingFigure 2
  • EP4167540B1 patent drawingFigure 3

AI summary

The subject matter described herein provides systems and techniques for a network planning and optimization tool that may allow for network capacity planning using key network failures for an arbitrary pair of network topology and demands. Performing network capacity planning with key network failures, instead of using other techniques, may avoid over-building the topology of a network. In particular, key network failures may be generated from the probabilistic failures, and the impact of these failures on a network may be computed. Expected flow availability SLO or a function thereof may be computed, using this information, and used by the tool to design a robust network. With an embedded flow availability calculation and updated risk framework, the capacitated cross-layer network topologies output by the tool may meet network demands/flows with their respective SLO type at the lowest cost.