Cloud Network Fault Management via Performance-Based Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current network fault management systems in cloud environments face limitations in increasing availability due to thresholds set irrelevant to the performance of target servers, hindering the system's overall performance and responsiveness.

Innovation Solution

A network fault management system that employs testing tools to measure server performance, a fault management unit to determine and analyze thresholds and policies, and monitoring tools to set and update monitoring policies based on real-time data, utilizing both rule-based and machine learning-based methods for accurate threshold calculation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a threshold is set for generating alarm events in a network fault management system, then the system can notify administrators of faults, but the availability cannot be increased further if the threshold is irrelevant to the performance of the target server

Engineering Contradiction:
ImproveavailabilityVSAvoidthreshold accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary performance testing of the target server before setting monitoring thresholds. Testing tools measure the actual performance characteristics of the server, and these measurements are used to establish accurate thresholds for fault detection. This preliminary action ensures that thresholds are relevant to the specific server being monitored, thereby improving both availability and measurement precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where performance measurement results from testing tools are fed back to the fault management unit. This feedback loop allows the system to continuously refine and adjust thresholds based on actual server performance data, ensuring that thresholds remain accurate and relevant over time, thus improving availability without sacrificing measurement precision.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If testing tools are used to measure server performance and determine thresholds, then accurate thresholds can be set, but the system complexity increases due to multiple components including testing tools, fault management unit, and monitoring tools

Engineering Contradiction:
Improvethreshold accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system is divided into distinct functional modules: testing tools for performance measurement, a fault management unit for analyzing results and determining thresholds, and monitoring tools for applying thresholds. This segmentation allows each component to perform its specific function efficiently, managing complexity through modular design while maintaining high measurement precision for threshold setting.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The fault management unit acts as an intermediary between the testing tools and the monitoring tools. It receives performance measurement results from testing tools, processes this information to determine appropriate thresholds, and then provides these thresholds to the monitoring tools. This intermediary role simplifies the overall system architecture by creating a clear interface between measurement and monitoring functions, reducing complexity while preserving threshold accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12015537B2Method for managing network failure in cloud environment and network failure management system
Publication Date: 2024.06.18 FOUND OF SOONGSIL UNIV IND COOP
  • US12015537B2 patent drawing
  • US12015537B2 patent drawing
  • US12015537B2 patent drawing

AI summary

Disclosed are a method for performing network fault management in a cloud environment and a network fault management system. A method for performing network fault management in a cloud environment according to another exemplary embodiment of the present invention includes steps of measuring, by testing tools, the performance of a target server and transmitting a measurement result to a fault management unit, determining, by the fault management unit, a threshold and a policy for a target host based on the transmitted measurement result, generating, by the fault management unit, templated information including the determined threshold and policy, transmitting, by the fault management unit, the templated information to monitoring tools, and setting, by the monitoring tools, a monitoring policy of the target host based on the transmitted information.