Automated Fault Isolation in Managed Services Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern communication networks face challenges in quickly identifying and resolving faults, leading to network downtime that incurs significant costs, as existing monitoring systems overwhelm engineers with a high volume of alarms and require manual processing, lacking automation and end-to-end visibility.

Innovation Solution

A managed services system with automated fault isolation and resolution capabilities, utilizing a fault handling module, analysis module, and automation engine to determine fault causes, generate alarms, and initiate automated processes, integrating with maintenance and trouble ticketing systems to streamline fault detection and recovery.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual alarm processing is used, then network monitoring coverage can be maintained, but fault isolation time increases and network downtime increases

Engineering Contradiction:
Improvenetwork availabilityVSAvoidfault isolation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables self-service fault isolation by automatically analyzing alarms, determining root causes, and executing resolution actions without requiring manual intervention from network engineers. The automated fault isolation system processes alarms independently, making decisions about fault location and resolution based on predefined policies and network state information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by pre-configuring fault isolation policies, acceptable resolutions, and escalation procedures before faults occur. When alarms are received, the system has already prepared the framework for rapid automated response, enabling quick decision-making about which actions to take and in what sequence.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If automated fault isolation is implemented, then fault isolation time is reduced, but system complexity increases

Engineering Contradiction:
Improvefault isolation timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The automated fault isolation system is segmented into distinct functional modules: alarm reception module, root cause analysis module, fault isolation module, and resolution execution module. Each module handles a specific aspect of fault processing independently, making the overall complex system manageable through modular architecture where each component has well-defined inputs and outputs.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If manual alarm processing is used, then operational simplicity is maintained, but network restoration time increases

Engineering Contradiction:
Improveoperational simplicityVSAvoidnetwork restoration time
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system incorporates feedback mechanisms that continuously monitor network state, alarm conditions, and resolution progress. Based on this feedback, the automated system dynamically adjusts its fault isolation approach, learns from previous resolutions, and optimizes its actions to achieve faster restoration times while maintaining operational effectiveness.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8924533B2Method and system for providing automated fault isolation in a managed services network
Publication Date: 2014.12.30 ATLASSIAN US INC
  • US8924533B2 patent drawing
  • US8924533B2 patent drawing
  • US8924533B2 patent drawing

AI summary

An approach provides support for automatic fault isolation in a managed services system. A fault indication corresponding to a communications network of a customer is determined. An analysis is performed to determine a root cause of the fault indication using information associated with the determined fault to output an alarm. A determination is made whether the alarm is associated with a maintenance event to update the alarm. Further, a workflow event corresponding to the alarm is created.