Topology-Dependent ML Engine Selection for Network Fault Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning (ML)-based approaches for network fault analysis and diagnostics are less than optimal in dynamically changing network environments, particularly in Next Generation mobile networks and large-scale data centers, due to the challenges of varying topologies and Key Performance Indicator (KPI) requirements.
Innovation Solution
A system and method for selecting a topology-dependent ML engine that adapts to changes in network topological configurations and KPI requirements, using a rule-based or built-in ML-based selector to facilitate root cause determination of faults, with data collection from SDN and non-SDN infrastructures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single ML engine is used for fault analysis, then the system structure is simple, but the system cannot adapt to dynamically changing network topologies and KPI requirements
Solution Approach 1:
The system dynamically selects ML engines based on current network topology and KPI requirements. Instead of using a static single ML engine, the system evaluates topology change indicators and KPI variations to determine when to switch between different ML engines, enabling adaptation to changing conditions while maintaining operational simplicity through automated selection
Solution Approach 2:
The system implements a multi-engine architecture where multiple ML engines are maintained, each optimized for specific topology types or KPI scenarios. A selection mechanism universally manages these engines, choosing the most appropriate one based on current network conditions, thereby providing versatility without requiring complete redesign for each scenario
2Adaptability or versatility
If multiple ML engines are maintained for different topologies, then adaptability improves, but system complexity and resource requirements increase
Solution Approach 1:
The system changes parameters such as topology indicators and KPI thresholds to determine when to switch between ML engines. By monitoring these parameters and comparing them against predefined thresholds or ranges, the system can selectively activate appropriate ML engines without maintaining all engines simultaneously, reducing resource requirements while preserving adaptability
Solution Approach 2:
The system segments the network operation space into different regions based on topology types and KPI ranges. Each segment is associated with a specific ML engine that is optimized for that segment's characteristics. This segmentation allows the system to use only the necessary ML engine for the current operating conditions rather than maintaining a full set of engines for all possible scenarios
3Measurement precision
If ML engine selection is performed frequently to adapt to changes, then fault analysis accuracy improves, but processing time and computational overhead increase
Solution Approach 1:
The system performs ML engine selection periodically based on monitored topology changes and KPI variations rather than continuously. By establishing thresholds for topology change indicators and KPI deviations, the system triggers engine selection only when significant changes occur, maintaining high fault analysis accuracy while minimizing unnecessary processing time and computational overhead
Solution Approach 2:
The system implements feedback mechanisms that monitor the performance and applicability of the current ML engine. When feedback indicates that the current engine is no longer suitable for the network conditions, the system triggers a reselection process. This feedback-driven approach ensures accurate fault analysis by selecting appropriate engines while reducing processing time by avoiding unnecessary frequent selections
Data Source
AI summary
A system, method and non-transitory computer readable media for effectuating ML-based fault analysis in a network (102A, 102B) comprising a plurality of nodes (104-N, 120-M). An example method (200A) comprises determining (202) that at least one of a topological configuration of the network (102 A, 102B) and one or more key performance indicator (KPI) requirements associated with the network (102A, 102B) have changed; and responsive to the determining, selecting (204) a machine language (ML) engine optimally adapted to facilitate root cause determination of any faults detected in the network (102 A, 102B) after the topological configuration or any KPI requirements of the network (102A, 102B) have changed.


