Automatic Network Failure Recovery via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As virtual networks built using SDN and NFV diversify, it becomes challenging to create and manage operation flows for automatic recovery from network failures, and existing technologies do not support the correction of operation flows themselves.

Innovation Solution

An automatic failure recovery system utilizing machine learning to create and correct operation procedures, comprising a recovery execution unit, parameter creation unit, learning unit, procedure execution unit, success determination unit, and procedure correction unit, which outputs failure data, selects and executes recovery tasks, determines recovery success, and corrects procedures based on network failure data and models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If virtual networks diversify using SDN and NFV, then network functionality and flexibility improve, but the complexity of creating and managing operation flows for automatic recovery increases

Engineering Contradiction:
Improvenetwork functionalityVSAvoidoperation flow management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system enables automatic creation and correction of operation flows through self-service mechanisms. The learning unit automatically generates recovery procedures based on failure data, and the procedure correction unit autonomously refines these procedures based on execution results, eliminating the need for manual creation and management of operation flows despite network diversification

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the operational parameters from manual procedure creation to automated machine learning-based generation. By transforming the parameter creation process into an automated computational task using failure data and recovery models, the system adapts to diverse network configurations without increasing management complexity

Inventive Principle:
Principle #35Parameter changes

2Reliability

If existing technology automatically corrects parameter values in tasks, then task execution reliability improves, but the ability to correct operation flows themselves is lost

Engineering Contradiction:
Improvetask execution reliabilityVSAvoidprocedure correction capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system introduces dynamic correction capabilities at multiple levels. While existing technology statically corrects parameter values, this system dynamically corrects entire operation flows based on failure patterns and execution results. The procedure correction unit adaptively modifies operation flows based on recovery levels and determination results, enabling both parameter-level and flow-level corrections

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements comprehensive feedback loops. The success determination unit provides feedback on whether recovery was achieved, and the procedure correction unit uses this feedback along with recovery level information to automatically correct operation flows. This feedback mechanism enables continuous improvement of both task execution and overall procedure effectiveness

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If manual creation and correction of operation flows is performed, then procedure accuracy can be ensured, but workload and time consumption increase

Engineering Contradiction:
Improveprocedure accuracyVSAvoidrecovery procedure creation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system replaces the mechanical process of manual procedure creation with an automated machine learning system. The learning unit uses failure data and recovery models to automatically generate operation flows, substituting human manual work with computational processes that maintain accuracy while dramatically improving productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary learning and model acquisition before actual recovery operations. By pre-training the learning unit with failure data and recovery models, the system prepares accurate operation flow templates in advance, enabling quick and accurate response when failures occur without requiring manual procedure creation at the time of incident

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11080128B2Automatic failure recovery system, control device, procedure creation device, and computer-readable storage medium
Publication Date: 2021.08.03 KDDI CORP
  • US11080128B2 patent drawing
  • US11080128B2 patent drawing
  • US11080128B2 patent drawing

AI summary

An automatic failure recovery system that, using machine learning, creates an operation procedure for recovering from a network failure or corrects the created operation procedure has a plurality of recovery tasks for recovering from the network failure; outputs failure data indicating network configuration information and failure information acquired upon occurrence of the network failure; selects an execution procedure of the recovery tasks, based on the failure data and a recovery model acquired in advance; executes the selected execution procedure when the network failure occurs; determines whether or not recovery from the network failure was achieved by the execution procedure; and notifies that the procedure is to be corrected, depending upon the result of the determination and a recovery level of the network failure.