Automatic Network Failure Recovery via Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As virtual networks built using SDN and NFV diversify, it becomes challenging to create and manage operation flows for automatic recovery from network failures, and existing technologies do not support the correction of operation flows themselves.
Innovation Solution
An automatic failure recovery system utilizing machine learning to create and correct operation procedures, comprising a recovery execution unit, parameter creation unit, learning unit, procedure execution unit, success determination unit, and procedure correction unit, which outputs failure data, selects and executes recovery tasks, determines recovery success, and corrects procedures based on network failure data and models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If virtual networks diversify using SDN and NFV, then network functionality and flexibility improve, but the complexity of creating and managing operation flows for automatic recovery increases
Solution Approach 1:
The system enables automatic creation and correction of operation flows through self-service mechanisms. The learning unit automatically generates recovery procedures based on failure data, and the procedure correction unit autonomously refines these procedures based on execution results, eliminating the need for manual creation and management of operation flows despite network diversification
Solution Approach 2:
The system changes the operational parameters from manual procedure creation to automated machine learning-based generation. By transforming the parameter creation process into an automated computational task using failure data and recovery models, the system adapts to diverse network configurations without increasing management complexity
2Reliability
If existing technology automatically corrects parameter values in tasks, then task execution reliability improves, but the ability to correct operation flows themselves is lost
Solution Approach 1:
The system introduces dynamic correction capabilities at multiple levels. While existing technology statically corrects parameter values, this system dynamically corrects entire operation flows based on failure patterns and execution results. The procedure correction unit adaptively modifies operation flows based on recovery levels and determination results, enabling both parameter-level and flow-level corrections
Solution Approach 2:
The system implements comprehensive feedback loops. The success determination unit provides feedback on whether recovery was achieved, and the procedure correction unit uses this feedback along with recovery level information to automatically correct operation flows. This feedback mechanism enables continuous improvement of both task execution and overall procedure effectiveness
3Manufacturing precision
If manual creation and correction of operation flows is performed, then procedure accuracy can be ensured, but workload and time consumption increase
Solution Approach 1:
The system replaces the mechanical process of manual procedure creation with an automated machine learning system. The learning unit uses failure data and recovery models to automatically generate operation flows, substituting human manual work with computational processes that maintain accuracy while dramatically improving productivity
Solution Approach 2:
The system performs preliminary learning and model acquisition before actual recovery operations. By pre-training the learning unit with failure data and recovery models, the system prepares accurate operation flow templates in advance, enabling quick and accurate response when failures occur without requiring manual procedure creation at the time of incident
Data Source
AI summary
An automatic failure recovery system that, using machine learning, creates an operation procedure for recovering from a network failure or corrects the created operation procedure has a plurality of recovery tasks for recovering from the network failure; outputs failure data indicating network configuration information and failure information acquired upon occurrence of the network failure; selects an execution procedure of the recovery tasks, based on the failure data and a recovery model acquired in advance; executes the selected execution procedure when the network failure occurs; determines whether or not recovery from the network failure was achieved by the execution procedure; and notifies that the procedure is to be corrected, depending upon the result of the determination and a recovery level of the network failure.


